Wednesday, Oct 7 | --:--
Back to home

Ultrafast Puts Sol at 750 tok/s — a Speed Class, Not a New Checkpoint

OpenAI previewed Ultrafast, a new API service tier that runs GPT-5.6 Sol up to 14× faster than Standard processing — up to 750 output tokens per second on Cerebras hardware — in a limited customer preview for incident response, voice, finance, and other latency-sensitive work.

Times of AI Desk 4 min read San Francisco, CA View as Markdown
Cover illustration for Ultrafast Puts Sol at 750 tok/s — a Speed Class, Not a New Checkpoint

Labs have spent 2026 selling intelligence; Ultrafast is OpenAI arguing that frontier Sol at 750 tok/s unlocks work that cheaper-faster-smaller models could not (voice, incident response, interactive research). It also deepens the Cerebras inference partnership as NVIDIA remains the training default.

OpenAI published Previewing Ultrafast mode: a new API service tier that runs GPT-5.6 Sol up to 14× faster than Standard processing, generating up to 750 output tokens per second. The stack is powered by Cerebras. Access is a limited preview for a select group of customers (Jane Street, Podium, Basis, and Rogo are quoted); a waitlist form is open as capacity grows. This is a speed class, not a new base model — Sol’s intelligence with a real-time serving path.

What shipped

Item Detail (openai.com/index/previewing-ultrafast)
Model GPT-5.6 Sol on Ultrafast
Speed claim Up to 14× Standard; up to 750 output tok/s
Hardware Cerebras partnership for ultra-low-latency inference
Surface OpenAI API first
Access Limited preview; expand as capacity grows
Use cases cited Incident response, financial research/security, customer support/voice, commerce, live research loops

OpenAI says internal teams already use Ultrafast to compress overnight experiment batches into same-day iteration and to read logs/traces while an outage is still unfolding. Early customer quotes: Jane Street’s John Crepezzi on developer focus; Podium on voice-stack latency; Basis on synchronous UX that previously required dumber models; Rogo on real-time financial research.

Independent rankings

Ultrafast is a serving tier, not a new checkpoint. Rankings for GPT-5.6 Sol already live in the July 9 / Aug 6 corpus. Do not invent a separate Arena/AA entry for “Sol Ultrafast.” The news is tokens per second + Cerebras, not a new Elo.

Limits

  • Limited preview; list price vs Standard not published here.
  • How wide the preview goes is capacity-gated.
  • No independent tok/s board cited in the primary.

Sources

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading