Thursday, Oct 8 | --:--
Back to home

Opus 4.6 — 1M Context, Agent Teams, and an Independent GDPval Lead

Anthropic launched Claude Opus 4.6 (Feb 5): 1M-token context beta, native agent teams, context compaction, adaptive thinking; $5/$25 MTok. Vendor Terminal-Bench 2.0 / HLE leads; Artificial Analysis GDPval-AA ~144 Elo over next-best (GPT-5.2 in Anthropic’s citation). Long-running collaborators, not single-turn chat.

Times of AI Desk 5 min read San Francisco View as Markdown
Cover illustration for Opus 4.6 — 1M Context, Agent Teams, and an Independent GDPval Lead

Frontier models are shifting from single-turn chat to long-running collaborators. Anthropic’s proprietary package on Opus 4.6: context scale + agent orchestration + compaction aimed at “context rot”—and a named Artificial Analysis board in the footnotes, not only vendor tables.

Anthropic released Claude Opus 4.6 on February 5. New capabilities: 1M token context (beta) with better long-context retrieval (company: 76% on a 1M needle-in-haystack variant vs much lower priors); agent teams in Claude Code for parallel subtasks; context compaction on the API; adaptive thinking / effort controls; expanded finance/docs/spreadsheet/presentation/research workflows. Pricing stays $5/$25 per million tokens. Vendor: highest Terminal-Bench 2.0; leads Humanity’s Last Exam for complex reasoning (company tables). On GDPval-AA (run independently by Artificial Analysis), Anthropic cites ~144 Elo over next-best (GPT‑5.2 in footnotes)—roughly 70% head-to-head framing. Contemporaneous Arena/community snapshots placed Opus 4.6 thinking variants at or near #1 across some text/coding/expert categories (secondary). Early partners: gains in multi-step autonomy, large codebase navigation, legal/finance reasoning.

Claims vs checks

Artificial Analysis GDPval-AA is the named third-party board in Anthropic’s primary post—stronger than pure vendor wallpaper. Terminal-Bench 2.0 / HLE / needle scores remain lab-reported unless boards replicate. Arena #1 snapshots are time-slice secondary—later 2026 boards still showed Opus 4.6 thinking top-tier after Opus 4.8 / Fable 5, but do not invent Elo here.

Limits

  • 1M context beta: quality and cost at extreme length vary by workload.
  • Agent teams add orchestration failure modes (handoffs, duplicated work).
  • Same-day GPT-5.3-Codex comparisons need matched harnesses.

Sources

  • Anthropic: “Introducing Claude Opus 4.6” (February 5, 2026). Primary.
  • Artificial Analysis GDPval-AA methodology and scores as cited in Opus 4.6 footnotes.
  • Terminal-Bench 2.0 and Humanity’s Last Exam leaderboards referenced in the release.
  • Anthropic docs on agent teams, compaction, adaptive thinking (Feb 2026).

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading