Nvidia Puts Groq 3 LPX Inference Racks Into Full Production
On August 24, 2026, Nvidia said Groq 3 LPX—the interactive inference accelerator that extends Vera Rubin—is in full production. Artificial Analysis measured 3,431 output tokens per second on Gemma 4 31B at 100K context. Nebius is the first named cloud; Groq will follow with Dell. Each LPX rack holds 256 LP30 accelerators.
TLDR
Monday, August 24, 2026 (Hot Chips): NVIDIA newsroom — NVIDIA Groq 3 LPX, the interactive AI inference accelerator that extends Vera Rubin NVL72, is in full production. Artificial Analysis ran its 100K-context suite on Gemma 4 31B against an LPX system in Nvidia’s own data centers and recorded 3,431 output tokens/second (newsroom rounds to 3,400; 4× the nearest alternative in Nvidia’s comparison). Nebius is the first AI cloud (Token Factory, same API). Groq (the inference-cloud company) said it will be among the first adopters, deploying with Dell. Product page: 256 interconnected LPU / LP30 accelerators per LPX rack.
What shipped
| Item | NVIDIA newsroom + product page + developer blog Aug 24 |
|---|---|
| SKU | Groq 3 LPX — inference accelerator, not a new GPU |
| Status | Full production (GTC preview was March; this is the GA clock) |
| Rack | 256 LP30 accelerators; 128 GB collective on-chip SRAM |
| Independent bench | AA 100K context, Gemma 4 31B: 3,431 tok/s output |
| First cloud | Nebius Token Factory — CTO Danila Shtan quoted |
| Next | Groq cloud + Dell; CNBC: racks online later this year |
| Huang | LPX is the ultrafast token-generation tier of Vera Rubin factories |
Same-day NVIDIA blogs (30× throughput/watt vs GB300 NVL72 on vendor AgentX traces) are supporting efficiency claims, not a second SKU. Folded here; not a separate article.
Caveat on the AA number: the developer blog is explicit that AA hit a private Nvidia-hosted endpoint. StorageReview and others note competing 870 tok/s-class figures are live public endpoints. Report the 3,431 as AA-on-Nvidia-gear, not a worldwide production-API record.
Product-line de-dupe: not GTC March LPX unveil, not Aug 11 Nemotron 3.5 Lightning, not Aug 17 Groq $350M Series A (that is the inference-cloud company’s round), not Aug 20 Poolside, not Aug 22 server ASP hikes. Vera CPU × SpaceXAI is a different SKU the same day.
Why this story matters
Nvidia is productizing the $20B Groq deal as a rack you can order, sitting next to Rubin GPUs rather than replacing them. Token generation—not prefill—is the agent bottleneck; 256 SRAM LPUs are the bet that GPUs stay busy while agents stay snappy. Watch: when Nebius Token Factory actually lights up, whether AA re-runs on a public endpoint, and Q2 language on August 26.
Sources
- NVIDIA: Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI (August 24, 2026)
- NVIDIA Groq 3 LPX product page
- NVIDIA Technical Blog: How Groq 3 LPX Unlocks Ultrafast Interactivity (August 24, 2026) — AA 3,431 tok/s
- Groq: Among the First to Bring NVIDIA Groq 3 LPX and Vera Rubin NVL72 to Market (August 24, 2026)
- CNBC: Nvidia says Groq racks will be online this year (August 24, 2026)