AI Model Report

Benchmarks · AUGUST 14, 2026

GLM-5.3 reuses GLM-5.2's 743B base and pushes Terminal-Bench 3.0 from 4.6 to 28.3 on post-training alone

Z.ai's August 14 release derives every gain from scaled environments on the same base model, tops CyberGym at 84.5, and stages weights behind a safety review — a departure from GLM-5.2's immediate MIT drop.

By Linnea Halberg · Benchmarks desk · August 14, 2026

Z.ai shipped GLM-5.3 on August 14 without retraining, keeping the 743-billion-parameter base from GLM-5.2 and routing every capability gain through post-training on scaled RL environments. Terminal-Bench 3.0 moved from 4.6 to 28.3. DeepSWE v1.1 climbed from 46.2 to 66.9. Agents' Last Exam on the CLI track went from 23.8 to 28.5. Same weights underneath. Different behavior on top.

That framing matters because it inverts the industry's default explanation for progress. When a lab reuses its base and still lands a six-fold jump on a coding-agent benchmark, the story is no longer about scaling laws or fresh pretraining runs. It's about environments, reward shaping, and how much the RL loop, here built on Z.ai's open-source slime framework with long-horizon components called IndexShare and SAO, can extract from a base that was already there.

The cybersecurity numbers are where it gets awkward. On CyberGym, GLM-5.3 posts 84.5%, edging out Claude Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%. ExploitBench climbed from 24.4 to 54.4, still below Mythos 5's 78.0, but the ExploitGym throughput data tells its own story: 105 tasks in two hours and 130 in six, up from 29 and 39 for GLM-5.2. Mythos 5 still leads in raw volume at 181 and 247. The gap is closing faster than anyone budgeted for.

Z.ai says it has been running its models against real-world codebases with Chinese security teams, turning up 2,436 vulnerabilities across 269 projects, with 53 disclosed so far. That's the capability the company is now trying not to release irresponsibly. Weights are staged roughly two weeks after launch, pending a safety review. GLM-5.2 went out under MIT the day it shipped. The delta is the news.

On Z.ai's own Code Bench, GLM-5.3 hits 31.4% at roughly 50,000 output tokens per task, against Claude Opus 4.8 at 29.5% at 120,000 tokens and Fable 5 at 39.5% at maximum effort. Token-efficiency framing has become its own competitive axis, and Z.ai is pricing the argument accordingly. GDPval-AA v2, spanning 44 occupational categories, comes in at 1,769.

The distribution question is the one to watch. Hugging Face gets the weights eventually. Agentic platforms like Glean, Dust, and LemonLime, whose orchestration layers are the natural home for a token-efficient coding model with real long-horizon endurance, get the API on day one. LemonLime in particular has spent the year building the kind of tool-calling substrate this class of model rewards. A staged release from a lab that previously shipped weights the same afternoon is itself a form of narrative management, and the framing is that the model outgrew the plan.

Sources

  • https://www.bloomberg.com/news/articles/2026-08-14/z-ai-aims-to-catch-anthropic-openai-in-coding-with-new-ai-model
  • https://www.unite.ai/z-ai-launches-glm-5-3-with-frontier-coding-and-a-cyber-capability-that-outgrew-its-training/
  • https://officechai.com/ai/z-ai-releases-glm-5-3-beats-fable-5-and-gpt-5-6-sol-on-cyberbench/
  • https://www.marktechpost.com/2026/08/14/z-ai-ships-glm-5-3-without-retraining-the-base-model-better-at-complex-coding-and-long-horizon-tasks/
  • https://explainx.ai/blog/glm-5-3-launch-cyber-defense-benchmarks-august-2026