OpenAI Previews Ultrafast: GPT-5.6 Sol at Up to 14× Standard Speed
On August 13, 2026, OpenAI previewed Ultrafast, a new API service tier that runs GPT-5.6 Sol up to 14× faster than Standard processing—up to 750 output tokens per second on Cerebras hardware—in a limited customer preview for incident response, voice, finance, and other latency-sensitive work.
TLDR
OpenAI on August 13, 2026 published Previewing Ultrafast mode: a new API service tier that runs GPT-5.6 Sol up to 14× faster than Standard processing, generating up to 750 output tokens per second. The stack is powered by Cerebras. Access is a limited preview for a select group of customers (Jane Street, Podium, Basis, and Rogo are quoted); a waitlist form is open as capacity grows. This is a speed class, not a new base model—Sol’s intelligence with a real-time serving path.
What shipped
| Item | Detail (openai.com/index/previewing-ultrafast) |
|---|---|
| Date | August 13, 2026 (Product) |
| Model | GPT-5.6 Sol on Ultrafast |
| Speed claim | Up to 14× Standard; up to 750 output tok/s |
| Hardware | Cerebras partnership for ultra-low-latency inference |
| Surface | OpenAI API first |
| Access | Limited preview; expand as capacity grows |
| Use cases cited | Incident response, financial research/security, customer support/voice, commerce, live research loops |
OpenAI says internal teams already use Ultrafast to compress overnight experiment batches into same-day iteration and to read logs/traces while an outage is still unfolding. Early customer quotes: Jane Street’s John Crepezzi on developer focus; Podium on voice-stack latency; Basis on synchronous UX that previously required dumber models; Rogo on real-time financial research.
Independent rankings
Ultrafast is a serving tier, not a new checkpoint. Rankings for GPT-5.6 Sol already live in the July 9 / Aug 6 corpus. Do not invent a separate Arena/AA entry for “Sol Ultrafast.” The news is tokens per second + Cerebras, not a new Elo.
Why this story matters
Labs have spent 2026 selling intelligence; Ultrafast is OpenAI arguing that frontier Sol at 750 tok/s unlocks work that cheaper-faster-smaller models could not (voice, incident response, interactive research). It also deepens the Cerebras inference partnership as NVIDIA remains the training default. Watch: list price vs Standard, how wide the preview goes, and whether Anthropic/Google answer with a named speed SKU.