NIST CAISI: GLM-5.3 is the most cyber-capable open-weight model yet, but trails the US frontier by about four months
On Sept 17, 2026 NIST's Center for AI Standards and Innovation (CAISI) published an assessment of Z.ai's GLM-5.3. It called GLM-5.3 "the most cyber-capable open-weight model released to date", ahead of Kimi K3, the previous best Chinese model. It also found the model's cyber capabilities "significantly lower than those of current U.S. frontier models", about four months behind them (e.g. SEC-Bench Pro 40.4% vs 90.2%).
Key facts
- Published by NIST CAISI on 2026-09-17; GLM-5.3 was announced 2026-08-14 and its weights were released about two weeks later
- SEC-Bench Pro: GLM-5.3 40.4% vs 90.2% for the best US frontier model
- ExploitBench: 61.1% (9.8/16) vs 100.0% (16.0/16)
- ExploitGym: 9.4% (47/498) vs 44.4% (223/502)
- CAISI OSS-Fuzz: 7.7% (23/297) vs 23.2% (69/297)
- Method: models run as agents in a ReAct harness with bash and python; turn limits 200 (SEC-Bench Pro), 300 (ExploitBench), 200 (ExploitGym), 300 (OSS-Fuzz); US models tested with cyber safeguards disabled
- CAISI estimate: GLM-5.3 lags the US frontier's cyber capability by about four months; it beats Kimi K3, the previous best PRC model
What happened
The US government's AI evaluation center (CAISI, part of NIST) ran GLM-5.3 and US frontier models through four agentic offensive-security benchmarks: SEC-Bench Pro, ExploitBench, ExploitGym and its own OSS-Fuzz task set. All models ran in the same ReAct agent harness with bash and python tools, and the US models had their cyber safeguards switched off. GLM-5.3 scored higher than any earlier open-weight model, including Moonshot's Kimi K3. On every benchmark it scored well below the best US model, and CAISI put the gap at about four months of progress.
Why it matters
This is the US government's own measurement of the open-weight cyber gap, published twelve days before Anthropic's Frontier Red Team report on the same model. The two reports frame the result differently. CAISI stresses that the US frontier is still clearly ahead. Anthropic stresses that GLM-5.3 nearly matched Claude Mythos Preview on its hardest exploit tasks and that its safeguards can be removed. Read together, they suggest that capabilities the US frontier had a few months earlier are now available to anyone who downloads the weights.
Changelog
- 2026-10-03: created
Related events
- Zhipu (Z.ai) releases GLM-5.3, top open-weights coding/agent model ★★★
- Anthropic Frontier Red Team: open-weights GLM-5.3 nearly matches Claude Mythos Preview at exploit development, with weak safeguards ★★★★
- Moonshot AI releases Kimi K3, a 2.8T-parameter open-weights multimodal model ★★★★★
- Hugging Face disables an abliterated GLM-5.3 repo branded "for offensive cyber"; it is re-uploaded under a new name and mirrored on Pirate Face ★★★
Sources (1)
id: 2026-09-17-caisi-glm-5-3-cyber-assessment · updated 2026-10-03 · open in the interactive timeline