Thursday, Oct 8 | --:--
Back to home

Claude Containment: 93% Approvals Were the Failure Mode — Sandboxes Cut Prompts 84%

Anthropic’s engineering post on containing Claude across claude.ai, Code, and Cowork centers environment isolation over approval dialogs — telemetry showed ~93% prompt approvals (fatigue), OS sandboxes cut prompts 84%, Gray Swan single-attempt ASR ~0.1% on Opus 4.7. Rare public postmortems of pre-trust hooks and egress via attacker API keys.

Times of AI Desk 7 min read San Francisco, CA View as Markdown
Cover illustration for Claude Containment: 93% Approvals Were the Failure Mode — Sandboxes Cut Prompts 84%

As coding and desktop agents gain shell, filesystem, and network access, containment engineering becomes as important as model capability. Anthropic’s rare public post is valuable less for the architecture diagram than for the measurable failure modes it admits.

Anthropic (May 25) published “How we contain Claude across products,” a long-form engineering account of blast-radius bounds across claude.ai, Claude Code, and Claude Cowork. Rather than relying only on human approvals or model-layer classifiers, the post centers environment isolation — gVisor containers, OS sandboxes, local VMs — backed by telemetry, red-team findings, and postmortems.

Three risks, three defense layers

Anthropic groups agent risk into user misuse, model misbehavior, and external attackers (including prompt injection).

  • Environment: process sandboxes, VMs, filesystem mounts, egress controls.
  • Model: system prompts, classifiers, probes, training — never 100% guarantees. On Gray Swan Agent Red Teaming, Claude Opus 4.7 held single-attempt attack success to roughly 0.1% (~5–6% after 100 adaptive attempts). Claude Code auto mode catches roughly 83% of overeager behaviors before execution.
  • External content: MCP servers, plugins, and web tools can inject untrusted text even when the connector itself is audited.

Product-specific patterns

claude.ai — Ephemeral gVisor containers on isolated infrastructure; no host filesystem — minimal blast radius, limited local capability.

Claude Code — Runs on the user machine. Early HITL approvals suffered approval fatigue: telemetry showed users approved about 93% of prompts. An OS-level sandbox (Seatbelt on macOS, bubblewrap on Linux) cut permission prompts by 84% and was open-sourced as sandbox-runtime. Anthropic also disclosed pre-trust vulnerabilities where project hooks ran before the “trust this folder” dialog — fixed by deferring project config until after consent.

Claude Cowork — Full local VM (Apple Virtualization on macOS, HCS on Windows) with only the selected workspace mounted; credentials stay on the host. A third-party disclosure showed attackers could still exfiltrate workspace files through api.anthropic.com via an attacker-controlled API key; fixed with an in-VM MITM proxy that only allows the VM’s provisioned session token.

Claims vs checks

Telemetry (93%, 84%), Gray Swan rates, and auto-mode catch rate are Anthropic-reported (Gray Swan bench is independent methodology; success rates as cited by Anthropic). Incident write-ups are admissions, not third-party audits of residual risk. Treat numbers as company telemetry snapshots.

Limits

  • Gray Swan adaptive ASR still ~5–6% after 100 attempts — not zero.
  • Model-layer defenses “never guarantee 100% safety” (Anthropic’s own framing).
  • MCP/plugin injection surface remains even with audited connectors.

Sources

  • Anthropic Engineering: “How we contain Claude across products” (May 25, 2026).
  • Related Anthropic posts linked from the article (Claude Code auto mode; sandbox-runtime; Mythos Preview sandbox escape; Glasswing).

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading