Friday, Oct 9 | --:--
Back to home

Groq 3 LPX Is a Rack You Can Order — AA Hit 3,431 tok/s on Nvidia’s Own Gear

Nvidia put Groq 3 LPX — the interactive inference accelerator that extends Vera Rubin — into full production. Artificial Analysis measured 3,431 output tokens per second on Gemma 4 31B at 100K context on a private Nvidia-hosted endpoint. Nebius is the first named cloud; Groq will follow with Dell. Report 3,431 as AA-on-Nvidia-gear, not a worldwide production-API record.

Times of AI Desk 5 min read Santa Clara, CA View as Markdown
Cover illustration for Groq 3 LPX Is a Rack You Can Order — AA Hit 3,431 tok/s on Nvidia’s Own Gear

Nvidia is productizing the $20B Groq deal as a rack you can order, sitting next to Rubin GPUs rather than replacing them. Token generation — not prefill — is the agent bottleneck; 256 SRAM LPUs are the bet that GPUs stay busy while agents stay snappy.

NVIDIA newsroom: NVIDIA Groq 3 LPX, the interactive AI inference accelerator that extends Vera Rubin NVL72, is in full production. Artificial Analysis ran its 100K-context suite on Gemma 4 31B against an LPX system in Nvidia’s own data centers and recorded 3,431 output tokens/second (newsroom rounds to 3,400; 4× the nearest alternative in Nvidia’s comparison). Nebius is the first AI cloud (Token Factory, same API). Groq (the inference-cloud company) said it will be among the first adopters, deploying with Dell. Product page: 256 interconnected LPU / LP30 accelerators per LPX rack.

What shipped

Item NVIDIA newsroom + product page + developer blog Aug 24
SKU Groq 3 LPX — inference accelerator, not a new GPU
Status Full production (GTC preview was March; this is the GA clock)
Rack 256 LP30 accelerators; 128 GB collective on-chip SRAM
Independent bench AA 100K context, Gemma 4 31B: 3,431 tok/s output
First cloud Nebius Token Factory — CTO Danila Shtan quoted
Next Groq cloud + Dell; CNBC: racks online later this year
Huang LPX is the ultrafast token-generation tier of Vera Rubin factories

Same-day NVIDIA blogs (30× throughput/watt vs GB300 NVL72 on vendor AgentX traces) are supporting efficiency claims, not a second SKU.

Caveat on the AA number: the developer blog is explicit that AA hit a private Nvidia-hosted endpoint. StorageReview and others note competing 870 tok/s-class figures are live public endpoints. Report the 3,431 as AA-on-Nvidia-gear, not a worldwide production-API record.

Not GTC March LPX unveil, not August 11 Nemotron 3.5 Lightning, not August 17 Groq $350M Series A (the inference-cloud company’s round), not August 20 Poolside, not August 22 server ASP hikes. Vera CPU × xAI is a different SKU the same day.

Limits

  • 3,431 tok/s is AA on a private Nvidia-hosted endpoint — not a public production-API record.
  • Nebius Token Factory lighting schedule and whether AA re-runs on a public endpoint are open.
  • Vendor AgentX 30× claims are Nvidia’s, not independent.

Sources

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading