Thinking Machines Releases Inkling-Small Open Weights
On July 30, 2026, Thinking Machines Lab released Inkling-Small—a 276B-total / 12B-active open-weight MoE that beats full Inkling (975B/41B) on SWE-Bench Verified (80.2% vs 77.6%) and ARC-AGI-2 (40.1% vs 36.5%) at roughly one-quarter the size and $0.30/$1.20 per million tokens.
TLDR
Thinking Machines Lab on July 30, 2026 released Inkling-Small as full open weights—the efficient sibling previewed with Inkling on July 15. Specs: 276B total / 12B active MoE, native multimodal (text, images, audio), up to 1M context, variable thinking effort. Lab claims it beats full Inkling on agentic coding (SWE-Bench Verified 80.2% vs 77.6%) and ARC-AGI-2 (40.1% vs 36.5%) while lagging on factual recall (SimpleQA 20.6% vs 43.9%). List pricing cited: $0.30 / $1.20 per million input/output tokens (~3–4× cheaper than Inkling). Weights on Hugging Face; fine-tune/chat on Tinker.
What Inkling-Small is
| Spec | Detail (Thinking Machines primary) |
|---|---|
| Architecture | MoE transformer; encoder-free multimodal (images via hierarchical patch encoding; audio via dMel spectrograms) |
| Scale | 276B total / 12B active (Inkling: 975B / 41B) |
| Context | Up to 1M tokens |
| Training note | On-policy distillation from Inkling + ~two weeks extra agentic-coding RL after the Jul 15 preview checkpoint |
| Hardware | Trained on NVIDIA GB300 NVL72 |
| Access | Hugging Face weights; Tinker playground + fine-tuning |
| Day-zero infra | SGLang, Unsloth GGUF, Baseten called out in secondary roundups |
Selected lab benchmarks (effort high; primary tables)
| Benchmark | Inkling-Small | Inkling |
|---|---|---|
| SWE-Bench Verified | 80.2% | 77.6% |
| ARC-AGI-2 | 40.1% | 36.5% |
| HLE text-only | 31.6% | 29.7% |
| SimpleQA Verified | 20.6% | 43.9% |
| AA Index v4.1 (lab table) | 40.0% | (Inkling previously ~41 on AA at Jul 15 debut) |
| Terminal Bench 2.1 (internal harness) | 64.7% | — |
Safety: StrongREJECT 98.4%; FORTRESS adversarial 71.6% / benign 96.9% at effort 0.99 (primary table).
Independent rankings
At launch, Artificial Analysis and Arena listings for Inkling-Small were still settling; the company post heavily cites AA evaluation methodology and public AA numbers for peers. Treat SWE-Bench / ARC figures as lab-reported unless/until AA or Arena publish independent Inkling-Small cards. Full Inkling held AA Intelligence ~41 as leading U.S. open weights at Jul 15 debut—Small is the efficiency play, not a closed-frontier rival to Fable/Sol/Kimi K3.
Product-line placement
Distinct slug from thinking-machines-inkling-open-weights (Jul 15 full Inkling). Same lab, new product line tier: small open MoE optimized for coding/agent cost curves. Thesis unchanged: customize on Tinker, compete on efficiency and post-trainability.
Why this story matters
Open-weight labs are racing on performance-per-active-parameter, not only total parameter headlines. Beating a 975B sibling on coding with 12B active is a sharp efficiency signal—and a direct answer to Chinese open MoE cost pressure. Watch community fine-tunes, serving cost on GB300-class clusters, factuality regression in production RAG, and whether Inkling-Small climbs independent AA/Arena boards in the next two weeks.