Guidelight Grades Frontier Labs on Control — Nobody Above a C+
Guidelight AI Standards graded Anthropic, OpenAI, Google, xAI, and Meta on six rogue-model control practices. No lab scored above 3/5 on any practice. Anthropic and OpenAI tied at C+ (2.50); Google D+; xAI D−; Meta F. Anthropic got a zero on containment despite the safety brand.

The industry is better at spotting misbehavior than at stopping it. After a month of eval-breakout disclosures, an independent grader put containment plans at the bottom of the sheet — and gave Anthropic a zero on that cell despite the safety brand.
Guidelight AI Standards, a nonprofit founded by former OpenAI safety leads Page Hedley and Steven Adler, published its first Control scorecard (Adler: “launching today”). It grades Anthropic, OpenAI, Google, xAI, and Meta on six public practices: logging, monitor efficacy, gated actions, circuit breaking, third-party review, and a containment plan. No company scored above 3/5 (“substantial partial”) on any practice. Overall letter grades are the average of the six 0–5 scores. TechCrunch amplified the containment finding on August 22; that is coverage, not a second ship.
The grades (public evidence through Aug 18)
| Lab | Overall | Containment plan | Notes from Guidelight |
|---|---|---|---|
| Anthropic | C+ (2.50) | 0 | Strongest on logging / monitor / gated actions / circuit-breaking (3s); no public containment protocol even in the Aug 14 Risk Report |
| OpenAI | C+ (2.50) | 3 | Tied overall; led containment (has paused internal workloads after incidents and described resume steps) |
| D+ (1.50) | 2 | AI Control Roadmap (arXiv) is the most specific forward plan; public implementation still limited | |
| xAI | D− (0.83) | 1 | Only lab that did not participate in METR’s Frontier Risk Report; limited public controls |
| Meta | F (0.67) | 0 | Weakest overall; METR disclosures are the main public window |
Scoring used only public materials (system cards, frameworks, risk reports, blogs, third-party write-ups). Two staff scored independently; companies could flag errors. A weak score can mean missing disclosure, not proven absence of internal controls. Not a dual-file of Anthropic’s Aug 14 Risk Report, OpenAI’s Aug 7 Astra / Aug 18 pacing posts, or METR’s May report — this is the first cross-lab control report card.
Limits
- Public-evidence only; weak scores can be disclosure gaps.
- TechCrunch Aug 22 is amplification, not a new assessment.
- Whether labs answer with public kill-switch protocols is open.
Sources
- Guidelight: Control assessment of frontier AI companies (August 2026; info through Aug 18)
- Steven Adler: Guidelight’s first scorecard launching today (August 18, 2026)
- The Decoder: No AI company fully applies basic control measures (August 19, 2026)
- TechCrunch: Frontier AI labs still won’t say how they’d contain a rogue model (August 22, 2026)
- Times of AI:
anthropic-august-2026-risk-report(August 14);openai-pacing-cyber-critical-rl-pause(August 18)