Thursday, Oct 8 | --:--
Back to home

Opus 4.7: Same Price, Harder SWE — Plus Cyber Blocks Before Mythos

Claude Opus 4.7 GA: Anthropic claims big lifts on hardest software-engineering and long-horizon autonomy, vision to 2,576px long edge, tighter instruction following — still $5/$25 MTok — while shipping automatic high-risk cyber blocks as Mythos-class mitigation rehearsal. Lab + early-tester benches; less capable than gated Mythos Preview by Anthropic’s own framing.

Times of AI Desk 7 min read San Francisco View as Markdown
Cover illustration for Opus 4.7: Same Price, Harder SWE — Plus Cyber Blocks Before Mythos

Anthropic’s dual track is explicit: broad GA improvements on Opus, gated power on Mythos. Opus 4.7’s proprietary frame is not another chat polish note — it is price-flat SWE/vision lift plus production cyber refusal machinery rehearsed before wider Mythos-class release.

Anthropic (April 16) made Claude Opus 4.7 generally available. Claims vs Opus 4.6: notable gains on hardest/long-running software engineering; substantially better vision (images up to 2,576 pixels long edge, ~3.75 MP, >3× prior Claude); more precise instruction following; more “tasteful” creative artifacts (UI, slides, docs). Pricing unchanged: $5 / $25 per million input/output tokens. Available on Claude products, API (claude-opus-4-7), Bedrock, Vertex AI, Microsoft Foundry; Claude Code, Cursor, GitHub Copilot integrations. API breaking changes vs 4.6; migration guidance provided. Anthropic positions 4.7 as less broadly capable than Mythos Preview.

Early checks (lab + testers)

Signal Reported lift (Anthropic / named early testers)
93-task coding bench +13% resolution vs 4.6; four tasks neither prior Opus nor Sonnet 4.6 solved
CursorBench 70% vs 58% (4.6)
Rakuten-SWE-Bench ~3× more production tasks resolved
Vision / multimodal Dense screenshots, diagrams, chemistry; one pentest visual eval 98.5% vs 54.5% cited

Themes from Hex, Replit, Cursor, Notion, Vercel, Databricks, Ramp, and others: autonomy, instruction fidelity, design taste, fewer corrections — selected partner quotes, not a blind third-party index crown in this release note.

Safety surface

Post–Project Glasswing: automatic detection/blocking of prohibited or high-risk cybersecurity requests; Cyber Verification Program for legitimate security pros (vuln research, pentest, red-team). Framed as learning mitigations before broader Mythos release.

Claims vs checks

GA availability, price, and cyber features are Anthropic primary. Benchmark deltas are lab-selected / early-tester — treat as directional, not Artificial Analysis Elo. Mythos comparison is Anthropic’s own capability hierarchy.

Limits

  • No AA Intelligence Index snapshot bundled in the launch post.
  • Partner benches vary in methodology and can be effort-skewed.
  • Cyber blocks may false-positive legitimate security work without verification enrollment.
  • “Tasteful” outputs remain subjective.

Sources

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading