Wednesday, Oct 7 | --:--
Back to home

Anthropic Opens Three Cyber Tiers — Red Team Hits Near-Unguarded Opus Rates

Anthropic folded Project Glasswing into an expanded Cyber Verification Program with Defense, Red Team and Specialized tiers covering Opus 5.5, Sonnet 5.5, Mythos 5.1 and future models. On Anthropic's CyScenarioBench, Defense blocked 46 of 50 Opus 5.5 runs; Red Team blocked none and completed 34 of 50 — which Anthropic equates to the model's 67.6% unguarded rate. Red Team still blocks physical-harm and mass-disruption actions. Partner-survey vulnerability counts (129k + 5.5k) are Anthropic-reported; the '5× higher' figure is an expectation, not a count.

Times of AI Desk 6 min read San Francisco, CA View as Markdown
Cover illustration for Anthropic Opens Three Cyber Tiers — Red Team Hits Near-Unguarded Opus Rates

Frontier labs are converging on the same cyber-release design: conservative public models, plus a much less filtered build for vetted defenders. Anthropic just widened its version of that lane.

On October 6, Anthropic launched an expanded Cyber Verification Program (CVP) that merges two six-month-old schemes — Project Glasswing (Mythos for a small critical-software cohort) and the original CVP (lowered safeguards on Opus and Sonnet for vetted security teams). All three new tiers include Claude Opus 5.5, Sonnet 5.5, Mythos 5.1 and future models. Generally available Claude keeps "conservative cyber safeguards that block most cyber work."

Three tiers

Tier Who What changes
Defense Access Company, non-profit, university and government security teams; critical-infrastructure operators of any size; smaller security firms; open-source maintainers; individual researchers with a record of reported vulnerabilities SOC / IR, malware RE, vulnerability validation. Anthropic aims to answer within days.
Red Team Access Organisations only Authorised penetration testing. Still blocks actions that could cause physical harm or mass disruption (e.g. deploying ransomware).
Specialized Access Orgs authorised to test flight OS, power grids, telecom, interbank transfer, government networks Fewest blocks. Every member reviewed in depth with the US government. Glasswing members move here.

Enrolled organisations must retain data so Anthropic can monitor for misuse until a planned Enterprise Frontier Safeguards option arrives later this fall. CVP runs on the Claude Platform, Vertex AI and Microsoft Foundry; on Bedrock only for customers eligible for the planned safeguards.

What the efficacy numbers say

On Anthropic's CyScenarioBench test of Opus 5.5 (10 challenges × 5 attempts):

  • No CVP: every task blocked on the first prompt.
  • Defense: blocked 46 of 50 runs.
  • Red Team: 0 blocks; completed 34 of 50, which Anthropic says is effectively the model's 67.6% unguarded success rate (34/50 is 68% — Anthropic's framing).

That is the point of the program: in the Red Team tier, offensive multi-stage completion approaches the unguarded model, and safety rests on vetting, retention and monitoring rather than refusal. Red Team is not ungated: physical-harm and mass-disruption blocks remain.

Vulnerability counts (Anthropic-reported)

Anthropic says Glasswing partners found at least 129,000 verified vulnerabilities from April to July 2026; its own open-source scanning found 5,500 more from April to October; more than 33,000 rated critical or high. These are drawn from 33 partner reports and called an undercount; Anthropic expects the true figure to be at least five times higher. That "5×" is an expectation, not a count. Reuters' lead rounded to "more than 100,000."

Same design, different labs

Same-day, Mistral Large 4 put a reduced-moderation partner build ahead of open weights. Gemini 4 Argon is rolling first through Google's Fairwind cyber-defender programme. Anthropic is widening from a Glasswing club (initial update, Mythos 5 defender fund) toward thousands of organisations and individual researchers.

Limits

  • CyScenarioBench and vulnerability totals are Anthropic's own.
  • Partner vulnerability figures are survey-based, not independently audited.
  • "5× higher" = Anthropic expectation, not measured count.
  • Red Team ≠ unguarded: physical-harm / mass-disruption blocks remain.

Sources

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading