Open Source · AUGUST 19, 2026
Z.ai holds GLM-5.3 weights after 84.5% CyberGym result and faster-than-expected exploit-chain lift
GLM-5.3 shares its base model with GLM-5.2; the entire jump — CyberGym 84.5%, ExploitBench 54.4%, 105 ExploitGym tasks in two hours — comes from scaled post-training. Z.ai is delaying open weights by roughly two weeks.
Z.ai shipped GLM-5.3 on August 14, and for the first time in the GLM line, the open weights aren't landing with the launch. Hugging Face gets them roughly two weeks late. The rest of the release notes explain the pause without ever saying so directly.
GLM-5.3 sits on the same 753-billion-parameter mixture-of-experts base and the same 1-million-token context window as GLM-5.2, which shipped in mid-July. Nothing about the substrate changed. What changed is everything above it. "Scaling post-training is all we did for GLM-5.3," the company said in its technical announcement, and the benchmark deltas make that a plausible claim rather than a modest one.
On CyberGym, the 1,507-task white-box vulnerability-discovery benchmark, GLM-5.3 posts 84.5% Pass@1, edging past Anthropic's Mythos 5 at 83.8% and OpenAI's GPT-5.6 Sol at 83.6%. GLM-5.2 scored 77.2% on the same eval a month earlier. The run used temperature 1.0, top_p 1.0, a 128,000-token output cap, and the Claude Code 2.1.207 harness.
ExploitBench is where the shape of the jump gets harder to wave off. GLM-5.2 scored 24.4%. GLM-5.3 scored 54.4%. That's more than a doubling from post-training alone, on a benchmark measuring end-to-end exploit chains rather than bug spotting. Mythos 5 still leads at 78% and GPT-5.6 Sol at 76.5%, but the frontier labs no longer have a moat measured in orders of magnitude.
ExploitGym tells the same story with a clock attached. In a two-hour budget, GLM-5.3 completes 105 tasks against GLM-5.2's 29; at six hours, 130 versus 39. Fable 5 hits 181 and 247 in those windows, GPT-5.6 Sol 216 and 293. GLM-5.3 isn't the leader here, but it closed most of the distance to the leaders in a single post-training cycle.
Alongside the benchmarks, Z.ai disclosed 2,436 vulnerability findings developed with Chinese security teams across 269 projects after expert review and deduplication. Of those, 1,097 are rated medium-to-high or critical. Only 53 are public at launch; 2,383 remain under embargo on a new Security Disclosure Ledger. The oldest disclosed flaw traces back to software introduced in 1981.
The GLM Coding Plan and ZCode environment are live now for paying users. That's the interesting sequencing. Weights that can drive an 84.5% CyberGym run against 269 projects' worth of real software are the kind of artifact governments have historically wanted registered before, not after, release. The 2016 Wassenaar debate over intrusion-software export controls asked exactly this question and never resolved it. Z.ai's two-week hold is the first time an open-weights lab has answered it unilaterally.
Sources
- https://siliconangle.com/2026/08/14/z-ai-debuts-glm-5-3-long-horizon-coding-cybersecurity-upgrades/
- https://venturebeat.com/technology/glm-5-3-is-here-with-advanced-cyber-capabilities-and-reportedly-already-found-a-serious-vulnerability-in-cursor
- https://www.axios.com/2026/08/14/china-open-source-ai-glm-53
- https://cybersecuritynews.com/glm-5-3-major-enhancements/
- https://www.developer-tech.com/news/z-ai-glm-5-3-cybergym-cybersecurity-ai-model-benchmark/