Tuesday, Aug 25 | --:--
Back to home

Nvidia Puts Groq 3 LPX Inference Racks Into Full Production

On August 24, 2026, Nvidia said Groq 3 LPX—the interactive inference accelerator that extends Vera Rubin—is in full production. Artificial Analysis measured 3,431 output tokens per second on Gemma 4 31B at 100K context. Nebius is the first named cloud; Groq will follow with Dell. Each LPX rack holds 256 LP30 accelerators.

Tech Insights Reporter 5 min read Santa Clara, CA
Cover illustration for Nvidia Puts Groq 3 LPX Inference Racks Into Full Production

TLDR

Monday, August 24, 2026 (Hot Chips): NVIDIA newsroom — NVIDIA Groq 3 LPX, the interactive AI inference accelerator that extends Vera Rubin NVL72, is in full production. Artificial Analysis ran its 100K-context suite on Gemma 4 31B against an LPX system in Nvidia’s own data centers and recorded 3,431 output tokens/second (newsroom rounds to 3,400; the nearest alternative in Nvidia’s comparison). Nebius is the first AI cloud (Token Factory, same API). Groq (the inference-cloud company) said it will be among the first adopters, deploying with Dell. Product page: 256 interconnected LPU / LP30 accelerators per LPX rack.

What shipped

Item NVIDIA newsroom + product page + developer blog Aug 24
SKU Groq 3 LPX — inference accelerator, not a new GPU
Status Full production (GTC preview was March; this is the GA clock)
Rack 256 LP30 accelerators; 128 GB collective on-chip SRAM
Independent bench AA 100K context, Gemma 4 31B: 3,431 tok/s output
First cloud Nebius Token Factory — CTO Danila Shtan quoted
Next Groq cloud + Dell; CNBC: racks online later this year
Huang LPX is the ultrafast token-generation tier of Vera Rubin factories

Same-day NVIDIA blogs (30× throughput/watt vs GB300 NVL72 on vendor AgentX traces) are supporting efficiency claims, not a second SKU. Folded here; not a separate article.

Caveat on the AA number: the developer blog is explicit that AA hit a private Nvidia-hosted endpoint. StorageReview and others note competing 870 tok/s-class figures are live public endpoints. Report the 3,431 as AA-on-Nvidia-gear, not a worldwide production-API record.

Product-line de-dupe: not GTC March LPX unveil, not Aug 11 Nemotron 3.5 Lightning, not Aug 17 Groq $350M Series A (that is the inference-cloud company’s round), not Aug 20 Poolside, not Aug 22 server ASP hikes. Vera CPU × SpaceXAI is a different SKU the same day.

Why this story matters

Nvidia is productizing the $20B Groq deal as a rack you can order, sitting next to Rubin GPUs rather than replacing them. Token generation—not prefill—is the agent bottleneck; 256 SRAM LPUs are the bet that GPUs stay busy while agents stay snappy. Watch: when Nebius Token Factory actually lights up, whether AA re-runs on a public endpoint, and Q2 language on August 26.

Sources

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading