# Ultrafast Puts Sol at 750 tok/s — a Speed Class, Not a New Checkpoint

Times of AI Desk · 2026-08-13 · Models

[https://timesof.ai/2026/08/openai-ultrafast-gpt-5-6-sol-cerebras](https://timesof.ai/2026/08/openai-ultrafast-gpt-5-6-sol-cerebras)

> OpenAI previewed Ultrafast, a new API service tier that runs GPT-5.6 Sol up to 14× faster than Standard processing — up to 750 output tokens per second on Cerebras hardware — in a limited customer preview for incident response, voice, finance, and other latency-sensitive work.

Labs have spent 2026 selling intelligence; Ultrafast is OpenAI arguing that **frontier Sol at 750 tok/s** unlocks work that cheaper-faster-smaller models could not (voice, incident response, interactive research). It also deepens the **Cerebras** inference partnership as NVIDIA remains the training default.

**OpenAI** published **Previewing Ultrafast mode**: a new **API service tier** that runs **GPT-5.6 Sol** up to **14× faster** than Standard processing, generating up to **750 output tokens per second**. The stack is **powered by Cerebras**. Access is a **limited preview** for a select group of customers (Jane Street, Podium, Basis, and Rogo are quoted); a waitlist form is open as capacity grows. This is a **speed class**, not a new base model — Sol’s intelligence with a real-time serving path.

## What shipped

| Item | Detail (openai.com/index/previewing-ultrafast) |
|------|------------------------------------------------|
| **Model** | **GPT-5.6 Sol** on Ultrafast |
| **Speed claim** | Up to **14×** Standard; up to **750** output tok/s |
| **Hardware** | **Cerebras** partnership for ultra-low-latency inference |
| **Surface** | OpenAI **API** first |
| **Access** | Limited preview; expand as capacity grows |
| **Use cases cited** | Incident response, financial research/security, customer support/voice, commerce, live research loops |

OpenAI says internal teams already use Ultrafast to compress overnight experiment batches into same-day iteration and to read logs/traces while an outage is still unfolding. Early customer quotes: Jane Street’s John Crepezzi on developer focus; Podium on voice-stack latency; Basis on synchronous UX that previously required dumber models; Rogo on real-time financial research.

## Independent rankings

Ultrafast is a **serving tier**, not a new checkpoint. Rankings for **GPT-5.6 Sol** already live in the July 9 / Aug 6 corpus. Do not invent a separate Arena/AA entry for “Sol Ultrafast.” The news is **tokens per second + Cerebras**, not a new Elo.

## Limits

- Limited preview; list price vs Standard not published here.
- How wide the preview goes is capacity-gated.
- No independent tok/s board cited in the primary.

## Sources

- [OpenAI: Previewing Ultrafast mode — GPT-5.6 Sol at up to 14X the speed (August 13, 2026)](https://openai.com/index/previewing-ultrafast/)
- [OpenAI Developer Community: Ultrafast preview note (August 13–14, 2026)](https://community.openai.com/)
