Sonnet 5.5 Matches Opus on the Index — and Burns the Most Tokens AA Has Measured
Claude Sonnet 5.5 keeps Sonnet 5’s $2/$10 sticker and, at max effort, scores 56 on Artificial Analysis’s Intelligence Index, two points behind Opus 5.5. Anthropic says it costs up to 30% less per task. AA measured about $7.60 a task at max, roughly 50% more than Sonnet 5, because output ran to about 193,000 tokens. Sticker price is not cost per task.

The workhorse tier just posted a near-Opus score at a Sonnet price. It also posted the heaviest token bill Artificial Analysis says it has ever measured. Both facts are the launch. The $2 / $10 sticker, unchanged from Sonnet 5, does not decide which one a buyer feels.
Claude Sonnet 5.5 is the second model in the Claude 5.5 family, a week after Opus 5.5. Anthropic prices it at $2 per million input tokens, $10 per million output, $0.20 cache reads, and $2.50 cache writes — half of Opus on input, output, and cache writes, tied with Opus on cache reads. API slug: claude-sonnet-5-5. Anthropic says it is on the Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure, with zero data retention available as on Opus 5.5 and Sonnet 5. Haiku 5.5 is still “coming weeks,” not this release.
Lab claims
Anthropic’s post (September 28) says generation is 30%+ faster than Sonnet 5, and that fewer tokens mean up to 30% lower cost per task than that predecessor. It also says Opus 5.5 remains clearly stronger at complex, open-ended work. Selected lab scores, at the effort Anthropic prints for Sonnet 5.5:
| Bench | Sonnet 5.5 | Sonnet 5 | Opus 5.5 |
|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4% (xhigh, its highest) |
| FrontierCode 1.1 Main | 46.2% max / 52.1% xhigh | 42.4% | 54.4% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% |
| GDPval-AA v2.1 | 1844 | 1449 | 1846 |
| AA-Briefcase v1.1 | 1811 | 1359 | 1822 |
| Humanity’s Last Exam, with tools | 64.5% | 54.9% | 67.7% |
| OSWorld 2.1, partial | 80.1% | 57.0% | 81.8% |
FrontierCode is lower at max than at xhigh. Anthropic’s footnote says max effort more often ran a wide code-review skill and, in cases Cognition looked at, timed out or edited past the task. GDPval-AA and AA-Briefcase were run by Artificial Analysis on a pre-release deployment Anthropic says had a structured-output bug. Anthropic expects any effect to be small and to understate the model.
Cyber safeguards are new at this tier. Anthropic says Sonnet 5.5’s cyber capability is in Opus 5’s range, so higher-risk cyber requests fall back to Sonnet 5. Biology safeguards match Sonnet 5. An automated audit of roughly 1,850 scenarios is lab-reported; Anthropic says no eval set catches every failure. Default effort is Medium in Claude apps and Claude Code, High on the platform. “Up to 30% cheaper per task” is Anthropic’s testing claim, not a guarantee at max effort.
Independent checks
Artificial Analysis, September 28, is the check that changes the invoice:
- Intelligence Index 56 at max effort with Anthropic’s default fallback — second, two points behind Opus 5.5 max at 58, and 18 points above Sonnet 5 max.
- About 193,000 output tokens per Index task, the highest AA says it has measured: about 60% above Opus 5.5 max or Sonnet 5 max, and about 7× GPT-6 Astra max.
- Cost per task about $7.60 at that max setting, which AA says is roughly 50% higher than Sonnet 5’s cost per task. Same sticker, higher bill, because the model writes more.
- AA’s own Terminal-Bench 4.0 run is 64%, against 60% for Opus 5.5 and GPT-6 Astra on that page — not Anthropic’s 70.6%. The gap is the point. The writeups do not share one harness.
- Near-parity with Opus on AA-Briefcase (1811 vs 1822), GDPval-AA (1844 vs 1846), and AutomationBench-AA (71% vs 70%), with the token burn attached.
- Weaker than Opus on AA-Omniscience factual accuracy (54% vs 66%) with a lower hallucination rate (47% vs 59%), and about six points lower on Humanity’s Last Exam and SciCode in AA’s comparison.
- Fallback fired on about 0.1% of Index tasks, to Sonnet 5, mostly inside Terminal-Bench.
- AA’s cost curve: at high effort the model sits just behind GPT-6 Sol on intelligence at essentially the same cost per task. At max it is near Opus and off the frontier on cost. At lower efforts, AA says GPT-6 Sol or Astra configurations deliver the intelligence with fewer output tokens.
AA ran the pre-release build with the structured-output bug and says scores may be slightly low. It plans a rerun. No Arena Elo for Sonnet 5.5 is in these sources. None is claimed here.
The OpenAI workhorse comparison is the sticker. $2 / $10 is also GPT-6 Sol’s price, and GPT-6.1 Sol launched the next day at the same rates. An Intelligence Index rank is not a coding-agent bakeoff. AA’s curve says max-effort Sonnet is an expensive way to buy the two points under Opus.
Limits
- Lab “up to 30% lower cost per task” and AA’s “about 50% higher at max” are different effort points and different task baskets. Quote both. Do not average them.
- Terminal-Bench 70.6 (Anthropic) versus 64 (AA) is unresolved. Opus’s lab number is at xhigh.
- The pre-release structured-output bug may understate AA’s index and the lab’s AA-run knowledge-work scores.
- Cyber fallback to Sonnet 5 will move some real security tasks off this model. Arena has not ranked Sonnet 5.5.