Wednesday, Oct 7 | --:--
Back to home

Nemotron Lightning Is a Cheap Worker Node — Switchyard Is the Real Product

NVIDIA shipped Nemotron 3.5 Lightning, a 30B MoE with ~3B active parameters for high-volume agent execution, plus NeMo Switchyard, an open router that steers each step to the cheapest capable model. Artificial Analysis puts Lightning near ~24 on intelligence — below Nemotron 3 Super — consistent with NVIDIA’s thesis that most agent tokens are execution, not reasoning.

Times of AI Desk 5 min read Santa Clara, CA View as Markdown
Cover illustration for Nemotron Lightning Is a Cheap Worker Node — Switchyard Is the Real Product

If agents spend most tokens on boring steps, the winning stack is frontier planner + fast open worker + router. Switchyard is strategically heavier than another small open model: it puts NVIDIA inside the decision of which model runs.

NVIDIA expanded the Nemotron 3 open family with Nemotron 3.5 Lightning — a 30-billion-parameter mixture-of-experts model with roughly 3 billion active parameters, distilled from Nemotron 3 Ultra for always-on agent execution — and NeMo Switchyard, an open-source model routing library that steers each request to the right model in a developer-defined pool. NVIDIA claims ~4× throughput versus comparable in-class models at similar intelligence and up to ~30% faster completion of agentic benchmark workloads. Weights land on Hugging Face / ModelScope / OpenRouter and as NIM microservices; Switchyard is on GitHub with gateway partners including OpenRouter, LiteLLM, and Kong.

What shipped

Piece Detail (NVIDIA blog / analyst notes)
Nemotron 3.5 Lightning 30B MoE, ~3B active; hybrid Mamba-Transformer; multi-token prediction; latent MoE; speculative decoding
Role High-volume specialized calls in long-running agents — not frontier planning
Open stack Open weights + post-training datasets + training recipes (Nemotron Coalition)
Hardware targets Jetson, GeForce RTX, DGX Spark, DGX Station; single-GPU laptop/desktop class claims
NeMo Switchyard Open router: custom algorithms/policies across a model pool; mid-task reshuffle
Partner claim LangChain-cited path: large cost cut routing Lightning vs Claude Opus-class with small accuracy tradeoff (vendor/partner measured — not independent AA)
Availability HF, ModelScope, OpenRouter, build.nvidia.com NIM; Switchyard on GitHub

Claims vs checks

Lab / partner: Throughput and agent-workload latency claims are NVIDIA’s; treat as vendor benches with disclosed architecture advantages (MoE + MTP + speculative decoding).

Independent:

  • Artificial Analysis (cited in contemporaneous analysis): Lightning around ~24 on AA’s index — near the bottom of the frontier chart and below NVIDIA’s own Nemotron 3 Super (~26), trailing Gemma 4 31B. Consistent with the pitch that most agent tokens are execution, not reasoning.
  • Do not present Lightning as a Sol/Opus peer; the product is cheap reliable worker nodes behind a router.

Distinct from the $500B Wall Street financing MOUs (Aug 10) and Firebird factory open (Aug 8). Lightning + Switchyard is open model + agent infrastructure software — NVIDIA selling the orchestration layer, not only GPUs.

Limits

  • AA intelligence ~24 is a size-class worker score, not a frontier claim.
  • LangChain cost-reduction path is partner-measured, not an independent composite.
  • Whether Switchyard becomes default in LiteLLM/Kong gateways is still open.

Sources

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading