Anthropic Frontier Red Team: open-weights GLM-5.3 nearly matches Claude Mythos Preview at exploit development, with weak safeguards
On Sept 29, 2026 Anthropic's Frontier Red Team published an evaluation of Zhipu's open-weights GLM-5.3 and found exploit-development skills close to Claude Mythos Preview (ExploitBench V8 12% vs 14%; binary exploitation 4% vs 6%, all other tested models 0%). Its safeguards were bypassed 64–100% of the time with simple techniques, and the community had removed them by "abliteration" for $1,200–$4,400 of compute. Anthropic called it "a meaningful step change in the cyber capabilities available to attackers".
Key facts
- ExploitBench (Chrome V8): GLM-5.3 12%, Claude Mythos Preview 14%, Claude Opus 4.6 and GLM-5.2 near 0%
- Internal binary-exploitation benchmark, 100 random tasks: GLM-5.3 achieved full control-flow hijacks in 4% of trials, Mythos Preview in 6%, every other tested model 0%
- Engagement with malicious requests: 0% unmodified, 64% with a deceptive prompt, 92% with prefilled reasoning, 100% for an abliterated version
- Abliteration (removing refusals from open weights) cost roughly $1,200–$4,400 of compute and took community developers days
- Recommendations: government safety testing of advanced models before release, safeguards on such capabilities, and wider vetted defender access to frontier models
What happened
Anthropic ran GLM-5.3, released with open weights in August, through the same offensive-cyber evaluations it uses for its own models. On the hardest tasks, building working exploits for browser and binary targets, it came within a few points of Claude Mythos Preview, the model Anthropic had kept restricted to vetted defenders under Project Glasswing. Because the weights are public, its refusals can be bypassed or removed.
Why it matters
It is the first time a frontier lab has published evidence that an open-weights model reached the level of exploit capability it had judged too risky to release widely. That undercuts restricted-release strategies and strengthens the case for pre-release government testing. The source is a competitor's evaluation of a Chinese model, so independent replication would help.
Changelog
- 2026-09-30: created (sweep 2026-09-29, via Simon Willison's Sept 29 quote)
Related events
- Zhipu (Z.ai) releases GLM-5.3, top open-weights coding/agent model ★★★
- Anthropic reveals Claude Mythos Preview, withholds it over cyber risk and launches Project Glasswing ★★★★★
- Z.ai says GLM-5.3 largely built the inference stack that serves GLM-5.3-Flash, calling it an early step toward recursive self-improvement ★★★
- Artificial Analysis launches the Cyber Index and an industry alliance (IBM, NVIDIA, Vercel, Collinear) for AI vulnerability-fixing evals ★★
Sources (2)
- officialAnthropic: GLM-5.3 and the spread of advanced cyber capabilities
- discussionSimon Willison: Quoting Anthropic Frontier Red Team
id: 2026-09-29-anthropic-glm-5-3-spread-of-cyber-capabilities · updated 2026-09-30 · open in the interactive timeline