OpenAI Publishes Jalapeño Inference Numbers Against Blackwell
On August 25, 2026, OpenAI’s Engineering post “Jalapeño’s first results” reported the first measured InferenceX numbers for its Broadcom-co-developed inference ASIC: 1.5–1.9× more AI work per watt at peak throughput and 1.7–3.6× lower end-to-end latency versus GB200/GB300, with first-party deployment planned by year-end.
TLDR
OpenAI on Tuesday, August 25, 2026 (Engineering / Company) published “Jalapeño’s first results show industry-leading speed and efficiency in AI inference.” The inference-only ASIC, co-developed with Broadcom and first unveiled in June, now has SemiAnalysis InferenceX numbers against commercially available Nvidia GB200 / GB300 racks. Across GPT‑OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T: 1.5–1.9× more AI work per watt at peak throughput and 1.7–3.6× lower end-to-end latency. Interactive workloads: 2.1–4.1× higher performance. Package TDP 700 W; measured sustained draw ≤550 W. OpenAI will start deploying Jalapeño in its own compute by end of 2026. Gen 2 is “deep in development”; Gen 3 is “taking shape.” Nvidia GPUs remain in the mix for training and inference.
What OpenAI published
| Item | openai.com Engineering Aug 25 |
|---|---|
| Chip | Jalapeño — first custom inference ASIC |
| Benchmark | InferenceX (SemiAnalysis); public models, not only OpenAI weights |
| Power basis | Jalapeño 700 W TDP vs GB200 1,200 W / GB300 1,400 W |
| GPT‑OSS 120B | ≈1.9× peak mixed TPS/kW (85,448 vs 44,960); ≈1.7× lower E2E latency |
| DeepSeek R1 | ≈1.7× peak TPS/kW; ≈3.6× lower E2E (1.65 s vs 5.99 s); min TBT ≈4.1× |
| Kimi K2.5 1T | ≈1.5× peak TPS/kW; ≈3.4× lower E2E |
| Design loop | Tape-out in nine months; Codex + GPT‑Astra brought three open-weight models to high performance in two months |
| Caveat | Lab-reported vs commercially shipping Blackwell; not a customer SKU; not a training chip |
Hot Chips coverage the same week (Tom’s Hardware, Bloomberg, Forbes) treats the Tuesday post as the primary. Nvidia is not a co-author; the comparison is OpenAI’s, with SemiAnalysis verifying some runs on-site.
Product-line de-dupe: not June Jalapeño unveil (no named comparison chip), not Aug 17 PORTS-Pike (Nvidia/OpenAI Ohio lease), not Aug 24 Vera CPU / Starmind. First-party silicon results are a separate hardware lane from Nvidia factory deals.
Why this story matters
OpenAI is no longer only a Nvidia tenant. A 700 W inference part that posts Pareto numbers on open models, programmed in part by Astra, is the lab’s bid to make tokens per watt an in-house lever before an IPO window. Watch: whether end-2026 “very small volumes” (Richard Ho, press call) slip, and whether Vera Rubin / HBM4 comparisons replace today’s Blackwell / HBM3E slide.
Sources
- OpenAI: Jalapeño’s first results show industry-leading speed and efficiency in AI inference (August 25, 2026)
- Tom’s Hardware: OpenAI’s 700W Jalapeño ASIC outpaces 1,400W Nvidia flagship GPU (August 25, 2026)
- Bloomberg: OpenAI Claims Its New Chips Can Outperform Nvidia Processors in Tests (August 25, 2026)