Sunday, Aug 23 | --:--
Back to home

Guidelight Grades Frontier Labs on Control: Nobody Above a C+

On August 18, 2026, Guidelight AI Standards published its first public Control assessment of Anthropic, OpenAI, Google, xAI, and Meta. Information current through August 18; no lab scored above 3/5 on any of six practices. Anthropic and OpenAI tied at C+ (2.50); Google D+ (1.50); xAI D− (0.83); Meta F (0.67).

Tech Insights Reporter 5 min read San Francisco, CA
Cover illustration for Guidelight Grades Frontier Labs on Control: Nobody Above a C+

TLDR

Guidelight AI Standards, a nonprofit founded by former OpenAI safety leads Page Hedley and Steven Adler, published its first Control scorecard on Tuesday, August 18, 2026 (Adler: “launching today”). It grades Anthropic, OpenAI, Google, xAI, and Meta on six public practices: logging, monitor efficacy, gated actions, circuit breaking, third-party review, and a containment plan. No company scored above 3/5 (“substantial partial”) on any practice. Overall letter grades are the average of the six 0–5 scores. TechCrunch amplified the containment finding on August 22; that is coverage, not a second ship.

The grades (public evidence through Aug 18)

Lab Overall Containment plan Notes from Guidelight
Anthropic C+ (2.50) 0 Strongest on logging / monitor / gated actions / circuit-breaking (3s); no public containment protocol even in the Aug 14 Risk Report
OpenAI C+ (2.50) 3 Tied overall; led containment (has paused internal workloads after incidents and described resume steps)
Google D+ (1.50) 2 AI Control Roadmap (arXiv) is the most specific forward plan; public implementation still limited
xAI D− (0.83) 1 Only lab that did not participate in METR’s Frontier Risk Report; limited public controls
Meta F (0.67) 0 Weakest overall; METR disclosures are the main public window

Scoring used only public materials (system cards, frameworks, risk reports, blogs, third-party write-ups). Two staff scored independently; companies could flag errors. A weak score can mean missing disclosure, not proven absence of internal controls.

Product-line de-dupe: not a dual-file of Anthropic’s Aug 14 Risk Report, OpenAI’s Aug 7 Astra / Aug 18 pacing posts, or METR’s May report. This is the first cross-lab control report card.

Why this story matters

The industry is better at spotting misbehavior than at stopping it. After a month of eval-breakout disclosures (OpenAI Hugging Face, Anthropic three orgs, Meta/Irregular), an independent grader put containment plans at the bottom of the sheet—and gave Anthropic a zero on that cell despite the safety brand. Watch: whether labs answer with public kill-switch protocols, or treat the C+ as a disclosure problem.

Sources

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading