Mistral Small 4 — One MoE Replaces Magistral, Pixtral, Devstral
Mistral Small 4 (March 16): 119B MoE (6B active/token, 8B with emb/out), 256k context, Apache 2.0 — unifies reasoning (Magistral), vision (Pixtral), agentic coding (Devstral) with configurable reasoning_effort. Claims 40% lower latency / 3× throughput vs Small 3; AA LCR/LiveCodeBench vs GPT-OSS 120B are lab/partner — treat as vendor.

Open-weight shops usually ship a zoo. Mistral’s proprietary move with Small 4: one Apache 2.0 MoE that absorbs Magistral reasoning, Pixtral vision, and Devstral coding — plus a reasoning_effort dial instead of three endpoints.
Mistral AI (March 16) released Small 4: hybrid MoE, 119B total parameters, 6B active per token (8B including embeddings and output), 256k context, text+image. reasoning_effort: none (fast, prior Small-like) vs high (Magistral-class step-by-step). Company: 40% lower end-to-end latency in optimized setups and 3× throughput vs Mistral Small 3; on AA LCR and LiveCodeBench with reasoning, matches/exceeds larger opens such as GPT-OSS 120B while producing up to 3.5–4× fewer characters in some cases. Available via Mistral API, AI Studio, Hugging Face; NVIDIA NIM; optimized for vLLM, SGLang, Transformers, llama.cpp. Mistral joined NVIDIA Nemotron Coalition as founding member.
Claims vs checks
Architecture, license, and availability are Mistral primary. Latency/throughput and bench comparisons are vendor/partner. Independent Artificial Analysis index should be checked before freezing open-weight ranks.
Limits
- “Unified” still means MoE routing quality varies by task.
- Shorter outputs can help cost but are not automatically better answers.
- Coalition membership ≠ shared training compute guarantees.
Sources
- Mistral AI: “Introducing Mistral Small 4” (March 16, 2026).
- NVIDIA NIM model card for mistral-small-4-119b-2603.
- Simon Willison’s blog and technical roundups confirming release details.