# Opus 4.6 — 1M Context, Agent Teams, and an Independent GDPval Lead

Times of AI Desk · 2026-02-05 · Models

[https://timesof.ai/2026/02/anthropic-releases-claude-opus-4-6-1m-context-agent-teams](https://timesof.ai/2026/02/anthropic-releases-claude-opus-4-6-1m-context-agent-teams)

> Anthropic launched Claude Opus 4.6 (Feb 5): 1M-token context beta, native agent teams, context compaction, adaptive thinking; $5/$25 MTok. Vendor Terminal-Bench 2.0 / HLE leads; Artificial Analysis GDPval-AA ~144 Elo over next-best (GPT-5.2 in Anthropic’s citation). Long-running collaborators, not single-turn chat.

Frontier models are shifting from single-turn chat to long-running collaborators. Anthropic’s proprietary package on Opus 4.6: **context scale + agent orchestration + compaction** aimed at “context rot”—and a named **Artificial Analysis** board in the footnotes, not only vendor tables.

**Anthropic** released **Claude Opus 4.6** on February 5. New capabilities: **1M token** context (beta) with better long-context retrieval (company: **76%** on a 1M needle-in-haystack variant vs much lower priors); **agent teams** in Claude Code for parallel subtasks; **context compaction** on the API; adaptive thinking / effort controls; expanded finance/docs/spreadsheet/presentation/research workflows. Pricing stays **$5/$25** per million tokens. Vendor: highest **Terminal-Bench 2.0**; leads **Humanity’s Last Exam** for complex reasoning (company tables). On **GDPval-AA** (run independently by **Artificial Analysis**), Anthropic cites ~**144 Elo** over next-best (**GPT‑5.2** in footnotes)—roughly **70%** head-to-head framing. Contemporaneous Arena/community snapshots placed Opus 4.6 thinking variants at or near #1 across some text/coding/expert categories (secondary). Early partners: gains in multi-step autonomy, large codebase navigation, legal/finance reasoning.

## Claims vs checks

**Artificial Analysis GDPval-AA** is the named third-party board in Anthropic’s primary post—stronger than pure vendor wallpaper. Terminal-Bench 2.0 / HLE / needle scores remain **lab-reported** unless boards replicate. Arena #1 snapshots are time-slice secondary—later 2026 boards still showed Opus 4.6 thinking top-tier after Opus 4.8 / Fable 5, but do not invent Elo here.

## Limits

- 1M context beta: quality and cost at extreme length vary by workload.
- Agent teams add orchestration failure modes (handoffs, duplicated work).
- Same-day GPT-5.3-Codex comparisons need matched harnesses.

## Sources

- [Anthropic: “Introducing Claude Opus 4.6”](https://www.anthropic.com/news/claude-opus-4-6) (February 5, 2026). Primary.
- Artificial Analysis GDPval-AA methodology and scores as cited in Opus 4.6 footnotes.
- Terminal-Bench 2.0 and Humanity’s Last Exam leaderboards referenced in the release.
- Anthropic docs on agent teams, compaction, adaptive thinking (Feb 2026).
