# Claude Containment: 93% Approvals Were the Failure Mode — Sandboxes Cut Prompts 84%

Times of AI Desk · 2026-05-25 · Research

[https://timesof.ai/2026/05/anthropic-how-we-contain-claude-agent-security](https://timesof.ai/2026/05/anthropic-how-we-contain-claude-agent-security)

> Anthropic’s engineering post on containing Claude across claude.ai, Code, and Cowork centers environment isolation over approval dialogs — telemetry showed ~93% prompt approvals (fatigue), OS sandboxes cut prompts 84%, Gray Swan single-attempt ASR ~0.1% on Opus 4.7. Rare public postmortems of pre-trust hooks and egress via attacker API keys.

As coding and desktop agents gain shell, filesystem, and network access, **containment engineering** becomes as important as model capability. Anthropic’s rare public post is valuable less for the architecture diagram than for the **measurable failure modes** it admits.

**Anthropic** (May 25) published “How we contain Claude across products,” a long-form engineering account of blast-radius bounds across **claude.ai**, **Claude Code**, and **Claude Cowork**. Rather than relying only on human approvals or model-layer classifiers, the post centers **environment isolation** — gVisor containers, OS sandboxes, local VMs — backed by telemetry, red-team findings, and postmortems.

## Three risks, three defense layers

Anthropic groups agent risk into user misuse, model misbehavior, and external attackers (including prompt injection).

- **Environment**: process sandboxes, VMs, filesystem mounts, egress controls.
- **Model**: system prompts, classifiers, probes, training — never 100% guarantees. On **Gray Swan Agent Red Teaming**, Claude Opus 4.7 held single-attempt attack success to roughly **0.1%** (~5–6% after 100 adaptive attempts). Claude Code auto mode catches roughly **83%** of overeager behaviors before execution.
- **External content**: MCP servers, plugins, and web tools can inject untrusted text even when the connector itself is audited.

## Product-specific patterns

**claude.ai** — Ephemeral **gVisor** containers on isolated infrastructure; no host filesystem — minimal blast radius, limited local capability.

**Claude Code** — Runs on the user machine. Early HITL approvals suffered **approval fatigue**: telemetry showed users approved about **93%** of prompts. An OS-level sandbox (Seatbelt on macOS, bubblewrap on Linux) cut permission prompts by **84%** and was open-sourced as **sandbox-runtime**. Anthropic also disclosed pre-trust vulnerabilities where project hooks ran before the “trust this folder” dialog — fixed by deferring project config until after consent.

**Claude Cowork** — Full local VM (Apple Virtualization on macOS, HCS on Windows) with only the selected workspace mounted; credentials stay on the host. A third-party disclosure showed attackers could still exfiltrate workspace files through api.anthropic.com via an attacker-controlled API key; fixed with an in-VM MITM proxy that only allows the VM’s provisioned session token.

## Claims vs checks

Telemetry (93%, 84%), Gray Swan rates, and auto-mode catch rate are **Anthropic-reported** (Gray Swan bench is independent methodology; success rates as cited by Anthropic). Incident write-ups are admissions, not third-party audits of residual risk. Treat numbers as company telemetry snapshots.

## Limits

- Gray Swan adaptive ASR still ~5–6% after 100 attempts — not zero.
- Model-layer defenses “never guarantee 100% safety” (Anthropic’s own framing).
- MCP/plugin injection surface remains even with audited connectors.

## Sources

- Anthropic Engineering: [“How we contain Claude across products”](https://anthropic.com/engineering/how-we-contain-claude) (May 25, 2026).
- Related Anthropic posts linked from the article (Claude Code auto mode; sandbox-runtime; Mythos Preview sandbox escape; Glasswing).
