Realtime-2 for Builders: Speech-to-Speech Agents With Flagship-Class Reasoning
OpenAI’s Realtime API adds GPT-Realtime-2 (GPT-5-class reasoning for live voice agents), Realtime-Translate (70+→13 languages), and Realtime-Whisper (streaming STT) — API-first infrastructure for phone agents and multilanguage apps, distinct from ChatGPT consumer Voice. Pricing: Realtime-2 $32/$64 per M audio tokens.

May 5 made GPT-5.5 Instant the ChatGPT default. May 7 is the parallel story for builders: speech-to-speech agents with flagship-class reasoning — infrastructure for phone agents and multilanguage apps, not a consumer Voice rebrand.
OpenAI (May 7) released three Realtime API voice models: GPT-Realtime-2 (most intelligent speech-to-speech path yet, GPT-5-class reasoning for harder multi-turn voice agents), GPT-Realtime-Translate (live speech translation: 70+ input → 13 output languages), and GPT-Realtime-Whisper (streaming speech-to-text while the user is speaking). Playground available.
What shipped
| Model | Role | Pricing notes (OpenAI) |
|---|---|---|
| GPT-Realtime-2 | Live conversational voice agents | $32 / $64 per M audio in/out; $0.40 cached input |
| GPT-Realtime-Translate | Live translation | $0.034 / min |
| GPT-Realtime-Whisper | Streaming STT | $0.017 / min |
Surface: Realtime API (WebRTC/WebSocket/SIP-style stacks) — API-first.
Claims vs checks
Model roles and prices are OpenAI primary. “GPT-5-class reasoning” and “most intelligent yet” are vendor positioning — check independent voice-agent benches when published; do not invent Elo. TechCrunch same-day coverage confirms the trio.
Limits
- Latency and interruption quality are workload-dependent; no independent MOS suite in the launch post.
- Translate language lists are OpenAI-stated.
- Distinct from ChatGPT Voice feature set and pricing.
Sources
- OpenAI: “Advancing voice intelligence with new models in the API” (May 7, 2026).
- TechCrunch same-day coverage; OpenAI community announcement thread (May 7, 2026).