Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. Anthropic Frontier Red Team: open-weights GLM-5.3 nearly…

Anthropic Frontier Red Team: open-weights GLM-5.3 nearly matches Claude Mythos Preview at exploit development, with weak safeguards

★★★★after cutoffpolicy-safetyAnthropicZhipu AIconfidence: high

On Sept 29, 2026 Anthropic's Frontier Red Team published an evaluation of Zhipu's open-weights GLM-5.3 and found exploit-development skills close to Claude Mythos Preview (ExploitBench V8 12% vs 14%; binary exploitation 4% vs 6%, all other tested models 0%). Its safeguards were bypassed 64–100% of the time with simple techniques, and the community had removed them by "abliteration" for $1,200–$4,400 of compute. Anthropic called it "a meaningful step change in the cyber capabilities available to attackers".

Key facts

What happened

Anthropic ran GLM-5.3, released with open weights in August, through the same offensive-cyber evaluations it uses for its own models. On the hardest tasks, building working exploits for browser and binary targets, it came within a few points of Claude Mythos Preview, the model Anthropic had kept restricted to vetted defenders under Project Glasswing. Because the weights are public, its refusals can be bypassed or removed.

Why it matters

It is the first time a frontier lab has published evidence that an open-weights model reached the level of exploit capability it had judged too risky to release widely. That undercuts restricted-release strategies and strengthens the case for pre-release government testing. The source is a competitor's evaluation of a Chinese model, so independent replication would help.

Changelog

  • 2026-09-30: created (sweep 2026-09-29, via Simon Willison's Sept 29 quote)

Related events

  1. Zhipu (Z.ai) releases GLM-5.3, top open-weights coding/agent model ★★★
  2. Anthropic reveals Claude Mythos Preview, withholds it over cyber risk and launches Project Glasswing ★★★★★
  3. Z.ai says GLM-5.3 largely built the inference stack that serves GLM-5.3-Flash, calling it an early step toward recursive self-improvement ★★★
  4. Artificial Analysis launches the Cyber Index and an industry alliance (IBM, NVIDIA, Vercel, Collinear) for AI vulnerability-fixing evals ★★

Sources (2)

id: 2026-09-29-anthropic-glm-5-3-spread-of-cyber-capabilities · updated 2026-09-30 · open in the interactive timeline