Jalapeño: OpenAI’s Bid for Inference Economics—Not Another Nvidia Tenant Forever
OpenAI and Broadcom unveiled Jalapeño, OpenAI’s first Intelligence Processor for LLM inference—end-of-2026 initial deployment goal, vendor claims of better perf/watt. Custom silicon is now table stakes for labs serving billions of tokens a day.

Whoever controls inference economics controls how widely frontier models can be sold. A Broadcom-co-developed, OpenAI-designed inference chip is a structural bet that OpenAI will not remain a pure software tenant of Nvidia forever—and that custom silicon is now table stakes for any lab serving billions of tokens a day.
OpenAI and Broadcom unveiled Jalapeño, OpenAI’s first custom Intelligence Processor—an accelerator built specifically for LLM inference, not a general training GPU. It is the first chip in a multi-generation platform the companies are building together, with end-of-2026 initial deployment goals and OpenAI claims of substantially better performance per watt than current state-of-the-art in early testing.
What Jalapeño is
From OpenAI’s primary company announcement:
- Name / class: Jalapeño — OpenAI’s first Intelligence Processor.
- Workload: Architected around OpenAI’s vision for the future of LLM inference (serving ChatGPT, API, and related production inference).
- Partners:
- Broadcom — silicon implementation, networking/connectivity (e.g., interconnect/Tomahawk-class networking in secondary technical coverage).
- Celestica — board, rack, and system expertise for deployable racks (cited in multi-partner platform description).
- Roadmap: First step in a multi-generation compute platform; initial deployment by end of 2026, expanding thereafter; Broadcom leadership messaging has referenced gigawatt-scale data-center rollouts with Microsoft and other partners beginning in 2026.
- Ceremony note: Hardware delivery to OpenAI CEO Sam Altman and President Greg Brockman by Broadcom CEO Hock Tan and President Charlie Kawwas.
- Performance framing: Early testing claims performance per watt substantially better than current SOTA; full public FLOPs/spec sheets remained limited at unveil. Secondary reporting described ~nine-month design-to-silicon cycle and inference-only positioning (training still expected to lean on Nvidia-class GPUs).
Strategic read
Google (TPU), Amazon (Trainium/Inferentia), Microsoft, and Meta already field custom silicon. Jalapeño is OpenAI’s formal entry into first-party inference ASICs—aimed at cost, latency, and supply independence for the token volume of ChatGPT-scale products as the company moves toward public markets.
Limits
- Perf/watt claims are early testing / vendor framing—no public FLOPs sheet at unveil.
- Inference-only positioning: training still expected on Nvidia-class GPUs.
- End-of-2026 deployment is a target, not a shipping receipt.
Sources
- OpenAI: “OpenAI and Broadcom unveil LLM-optimized inference chip” (openai.com/index/openai-broadcom-jalapeno-inference-chip/, June 24, 2026). Primary announcement.
- Reuters / industry coverage of the Jalapeño unveil, inference focus, and partner deployment framing (June 24, 2026).