Sunday, Aug 23 | --:--
Back to home

OpenAI Previews Ultrafast: GPT-5.6 Sol at Up to 14× Standard Speed

On August 13, 2026, OpenAI previewed Ultrafast, a new API service tier that runs GPT-5.6 Sol up to 14× faster than Standard processing—up to 750 output tokens per second on Cerebras hardware—in a limited customer preview for incident response, voice, finance, and other latency-sensitive work.

Tech Insights Reporter 4 min read San Francisco, CA
Cover illustration for OpenAI Previews Ultrafast: GPT-5.6 Sol at Up to 14× Standard Speed

TLDR

OpenAI on August 13, 2026 published Previewing Ultrafast mode: a new API service tier that runs GPT-5.6 Sol up to 14× faster than Standard processing, generating up to 750 output tokens per second. The stack is powered by Cerebras. Access is a limited preview for a select group of customers (Jane Street, Podium, Basis, and Rogo are quoted); a waitlist form is open as capacity grows. This is a speed class, not a new base model—Sol’s intelligence with a real-time serving path.

What shipped

Item Detail (openai.com/index/previewing-ultrafast)
Date August 13, 2026 (Product)
Model GPT-5.6 Sol on Ultrafast
Speed claim Up to 14× Standard; up to 750 output tok/s
Hardware Cerebras partnership for ultra-low-latency inference
Surface OpenAI API first
Access Limited preview; expand as capacity grows
Use cases cited Incident response, financial research/security, customer support/voice, commerce, live research loops

OpenAI says internal teams already use Ultrafast to compress overnight experiment batches into same-day iteration and to read logs/traces while an outage is still unfolding. Early customer quotes: Jane Street’s John Crepezzi on developer focus; Podium on voice-stack latency; Basis on synchronous UX that previously required dumber models; Rogo on real-time financial research.

Independent rankings

Ultrafast is a serving tier, not a new checkpoint. Rankings for GPT-5.6 Sol already live in the July 9 / Aug 6 corpus. Do not invent a separate Arena/AA entry for “Sol Ultrafast.” The news is tokens per second + Cerebras, not a new Elo.

Why this story matters

Labs have spent 2026 selling intelligence; Ultrafast is OpenAI arguing that frontier Sol at 750 tok/s unlocks work that cheaper-faster-smaller models could not (voice, incident response, interactive research). It also deepens the Cerebras inference partnership as NVIDIA remains the training default. Watch: list price vs Standard, how wide the preview goes, and whether Anthropic/Google answer with a named speed SKU.

Sources

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading