Tuesday, Sep 15 | --:--
Back to home

OpenAI Publishes Jalapeño Inference Numbers Against Blackwell

On August 25, 2026, OpenAI’s Engineering post “Jalapeño’s first results” reported the first measured InferenceX numbers for its Broadcom-co-developed inference ASIC: 1.5–1.9× more AI work per watt at peak throughput and 1.7–3.6× lower end-to-end latency versus GB200/GB300, with first-party deployment planned by year-end.

Tech Insights Reporter 5 min read San Francisco, CA
Cover illustration for OpenAI Publishes Jalapeño Inference Numbers Against Blackwell

TLDR

OpenAI on Tuesday, August 25, 2026 (Engineering / Company) published “Jalapeño’s first results show industry-leading speed and efficiency in AI inference.” The inference-only ASIC, co-developed with Broadcom and first unveiled in June, now has SemiAnalysis InferenceX numbers against commercially available Nvidia GB200 / GB300 racks. Across GPT‑OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T: 1.5–1.9× more AI work per watt at peak throughput and 1.7–3.6× lower end-to-end latency. Interactive workloads: 2.1–4.1× higher performance. Package TDP 700 W; measured sustained draw ≤550 W. OpenAI will start deploying Jalapeño in its own compute by end of 2026. Gen 2 is “deep in development”; Gen 3 is “taking shape.” Nvidia GPUs remain in the mix for training and inference.

What OpenAI published

Item openai.com Engineering Aug 25
Chip Jalapeño — first custom inference ASIC
Benchmark InferenceX (SemiAnalysis); public models, not only OpenAI weights
Power basis Jalapeño 700 W TDP vs GB200 1,200 W / GB300 1,400 W
GPT‑OSS 120B 1.9× peak mixed TPS/kW (85,448 vs 44,960); ≈1.7× lower E2E latency
DeepSeek R1 1.7× peak TPS/kW; ≈3.6× lower E2E (1.65 s vs 5.99 s); min TBT ≈4.1×
Kimi K2.5 1T 1.5× peak TPS/kW; ≈3.4× lower E2E
Design loop Tape-out in nine months; Codex + GPT‑Astra brought three open-weight models to high performance in two months
Caveat Lab-reported vs commercially shipping Blackwell; not a customer SKU; not a training chip

Hot Chips coverage the same week (Tom’s Hardware, Bloomberg, Forbes) treats the Tuesday post as the primary. Nvidia is not a co-author; the comparison is OpenAI’s, with SemiAnalysis verifying some runs on-site.

Product-line de-dupe: not June Jalapeño unveil (no named comparison chip), not Aug 17 PORTS-Pike (Nvidia/OpenAI Ohio lease), not Aug 24 Vera CPU / Starmind. First-party silicon results are a separate hardware lane from Nvidia factory deals.

Why this story matters

OpenAI is no longer only a Nvidia tenant. A 700 W inference part that posts Pareto numbers on open models, programmed in part by Astra, is the lab’s bid to make tokens per watt an in-house lever before an IPO window. Watch: whether end-2026 “very small volumes” (Richard Ho, press call) slip, and whether Vera Rubin / HBM4 comparisons replace today’s Blackwell / HBM3E slide.

Sources

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading