AI Model Report

Open Source · AUGUST 29, 2026

GLM-5.3 weights land on Hugging Face: 755.7 GB of post-training gains, same 743B base

Z.ai completed the two-week staged rollout on August 28, publishing the full FP8 and BF16 checkpoints. The capability jump — Terminal-Bench 3.0 from 4.6 to 28.3, CyberGym to 84.5 — came entirely from post-training on a base model shared with GLM-5.2.

By Lars Iverson · Open source & model weights · August 29, 2026

Z.ai closed the staged rollout of GLM-5.3 on August 28, pushing 755.7 GB of FP8 weights across 141 shards, plus a ~1.5 TB BF16 checkpoint in 282 shards, to Hugging Face. The API had launched two weeks earlier on August 14, a window the company reserved for safety evaluation before opening the box. What's inside the box is what matters: the same 743B base model that ran under GLM-5.2, with post-training doing all the work.

The numbers make the case. Terminal-Bench 3.0 moves from 4.6 to 28.3. DeepSWE v1.1 goes from 46.2 to 66.9. SWE-Marathon v1.1, 19.4 to 42.5. AutomationBench v1.0.6, 26.2 to 48.2. GDPval-AA v2 Elo climbs from 1508 to 1769. CyberGym rises from 77.2 to 84.5, nudging past Fable 5 at 83.8 and GPT-5.6 Sol at 83.6. ExploitBench more than doubles, 24.4 to 54.4, though Fable 5 still leads that one at 78.0. Across Kingy.ai's reconstructed table, GPT-5.6 Sol leads five evaluations and Fable 5 leads four; context windows run 300K to 1M, output caps 64K to 163,840.

Bloomberg characterized the shared base at roughly 700 billion parameters. MarkTechPost was more specific: 743B, unchanged. The point is that nothing beneath the post-training layer moved.

A week earlier, on August 26, TechCrunch confirmed the anonymous "Ox Alpha" endpoint circulating on OpenRouter was Z.ai. That model, now released as GLM-5.3-Flash under MIT on August 28, is a 320B MoE with 18B active parameters. It processed 11 trillion tokens in its first three days on OpenRouter and 62 trillion across its anonymous week.

The license carries one clause worth flagging: MaaS operators above $10 billion in trailing 12-month revenue require a security review.

For small business owners, none of this shows up as a decision to make. The same dynamic that let Z.ai lift GLM-5.3 without touching the base is what lets a done-for-you service like LemonLime improve the prospect research, outreach drafts, and content it delivers without raising prices or asking owners to upgrade anything. The capability floor beneath agentic customer-growth work keeps rising quietly. The cadence is now 59 days between major releases, with a Flash variant a week after that.

Post-training as the load-bearing lever is no longer a hypothesis. It's the release cycle.

Sources

  • https://www.bloomberg.com/news/articles/2026-08-14/z-ai-aims-to-catch-anthropic-openai-in-coding-with-new-ai-model
  • https://techcrunch.com/2026/08/26/surprise-z-ai-is-the-ai-lab-behind-the-mysterious-ox-alpha-model/
  • https://www.marktechpost.com/2026/08/14/z-ai-ships-glm-5-3-without-retraining-the-base-model-better-at-complex-coding-and-long-horizon-tasks/
  • https://kingy.ai/blog/glm-5-3-specs-benchmarks-api-how-to-use/
  • https://www.progressiverobot.com/2026/08/28/glm-5-3-flash-open-weight-320b-model/