Anthropic Red Team: Open-Weight GLM-5.3 Nearly Matches Mythos Preview on Exploit Building; Safeguards Fall to Simple Bypasses
In a Sept. 29 Frontier Red Team post, Anthropic reports that Z.ai’s open-weight GLM-5.3 built end-to-end V8 exploits in 50 of 410 ExploitBench attempts versus 56 for Claude Mythos Preview, and reached full control-flow hijack on 4% of an internal 100-task binary set versus Mythos Preview’s 6%. In simulated harmful-request tests, Anthropic says a cover story, prefilled reasoning, and an abliterated copy got GLM-5.3 to engage 64%, 92%, and 100% of the time; safeguarded Claude models stayed at 0 under API safeguards. Separately, NIST’s CAISI (Sept. 17) calls GLM-5.3 the most cyber-capable open-weight model released to date but about four months behind the U.S. frontier — under CAISI’s own harness, not Anthropic’s.

A competitor lab and a U.S. government evaluator now broadly agree on the same uncomfortable fact: near-frontier exploit capability is downloadable. They do not agree on a single scoreboard — and the lab making the loudest case is selling wider defender access to its own models.
Anthropic’s Frontier Red Team, in a September 29, 2026 research post, says Z.ai’s open-weight GLM-5.3 can autonomously build end-to-end cyber exploits at rates close to Claude Mythos Preview, and that GLM-5.3’s built-in safeguards fall to simple bypasses in Anthropic’s simulated tests. Two weeks earlier, NIST’s Center for AI Standards and Innovation (CAISI) separately called GLM-5.3 “the most cyber-capable open-weight model released to date,” while placing it about four months behind the U.S. frontier on CAISI’s own composite. Keep Anthropic’s numbers attributed to Anthropic — they are not the same measurements as CAISI’s.
What Anthropic measured
On ExploitBench (known V8 vulnerabilities → working end-to-end exploits), Anthropic reports GLM-5.3 succeeded in 50 of 410 attempts versus 56 of 410 for Claude Mythos Preview. On an internal binary-exploitation set of 100 randomly selected OSS-Fuzz-style tasks, full control-flow hijack rates were 4% for GLM-5.3 and 6% for Mythos Preview; Claude Opus 4.6 and GLM-5.2 scored 0. Anthropic also describes researcher-in-the-loop sessions in which GLM-5.3 found and chained previously unknown flaws in a sandboxed Linux browser build, and in which GLM-5.3-Flash turned a known Chrome CVE into an ARM64 exploit chain in about eight hours of model work (~$20.40 at Zhipu API prices, per Anthropic).
On safeguards, Anthropic says GLM-5.3 often refuses clearly harmful requests out of the box, but:
| Bypass | GLM-5.3 engagement (Anthropic simulation) | Safeguarded Claude (API) |
|---|---|---|
| Deceptive red-team cover story | 64% | 0% |
| Prefill model reasoning to “proceed” | 92% | Not feasible via Anthropic API |
| Abliterated open-weight copy | 100% | Weights not public |
Anthropic reports producing an abliterated GLM-5.3 copy in about 2,200 GPU hours (~$4,400), with little change on GPQA-Diamond and only a few percent drop on a CyberGym subset. Public abliterated copies appeared within days of release, the post says. Claude weights cannot be abliterated the same way because they are not publicly released.
What CAISI measured — separately
CAISI’s September 17 assessment evaluates GLM-5.3 on four cyber benchmarks under a ReAct harness with U.S. models tested with cyber safeguards disabled when applicable. Headline findings: GLM-5.3 is the most cyber-capable open-weight model CAISI has evaluated to date, and it lags the U.S. frontier by about four months on CAISI’s aggregate cyber-capability index.
Under CAISI’s harness (not Anthropic’s), example scores include ExploitBench 61.1% for GLM-5.3 versus 100% for the U.S. frontier best, SEC-Bench Pro 40.4% vs 90.2%, ExploitGym 9.4% vs 44.4%, and CAISI OSS-Fuzz 7.7% vs 23.2%. Those percentages are not comparable to Anthropic’s 50/410 and 4%/6% figures; different task counts, grading, and harness settings.
How to read the two posts together
Anthropic is a direct competitor arguing that freely downloadable near-Mythos exploit skill, plus weak open-weight safeguards, raises attacker capability — and that defenders should get broader access to frontier Claude cyber tools (Project Glasswing / trusted programs). CAISI supplies an independent government line that open-weight cyber capability has jumped, while still trailing released U.S. frontier models by roughly a fiscal quarter on CAISI’s scale. Neither post establishes GLM-5.3 as today’s U.S. frontier without that lag qualification.
The timing also matters next to Korea’s ARTEX bank-breach wave: open tools are already showing up in sector incidents even when the model in those cases was a pen-testing agent, not GLM-5.3.
Limits
- ExploitBench counts, binary-hijack rates, bypass percentages, and abliteration cost are Anthropic’s evaluations; Anthropic is a competitor to Z.ai.
- CAISI’s percentages and “four months behind” claim use CAISI’s harness and composite; U.S. models were often tested with safeguards off. Those figures are not interchangeable with Anthropic’s.
- Researcher-driven 0-day anecdotes are Anthropic’s sandboxed sessions; maintainers were notified per the post.
- Artificial Analysis Intelligence Index lists GLM-5.3 (max) at 45 on a general index — not a cyber board — and is not used here as a cyber ranking.
- Digitimes’ Oct. 5 follow was paywalled beyond headline/lede and supports no claim here.