AI Model Report

Open Source · AUGUST 30, 2026

Z.ai Ships GLM-5.3-Flash: 320B-A18B MoE, MIT Weights, and API Pricing 10x Cheaper Than GLM-5.2

The model that spent a week anonymously topping OpenRouter as 'Ox Alpha' launched August 26 with a 1M-token context, hybrid linear-plus-sparse attention, and a launch-promo API price of $0.075 per million input tokens.

By Lars Iverson · Open source & model weights · August 30, 2026

Z.ai released GLM-5.3-Flash on August 26, 2026, publishing MIT-licensed weights to Hugging Face for a 320-billion-parameter mixture-of-experts model that scores 57 on the Artificial Analysis Intelligence Index v4.1.1 at roughly $0.045 per task. The week before, the same model had been quietly topping OpenRouter and OpenCode leaderboards under the codename "Ox Alpha," which is a familiar Z.ai maneuver: ship anonymously, watch developers rank it against the frontier, then reveal the price.

The price is the point. Standard API rates land at $0.15 per million input tokens and $0.50 per million output, with a launch promotion running through September 9, 2026 that halves both to $0.075 and $0.25. Cached input tokens are priced at $0.015. Z.ai puts the delta versus GLM-5.2 at roughly 10x cheaper to serve.

Architecture explains most of that. The 45-layer language model activates 18 billion of its 320 billion parameters per token, routing each token through 8 of 288 experts, and interleaves KDA linear-attention layers with NoPE sparse MLA layers. Native FP8 weights, plus a component Z.ai calls IndexPool, cut attention compute by about 3x and shrink the KV cache 4.4x versus GLM-5.3. Context runs to 1 million tokens. First-party throughput clocks around 49 tokens per second.

The capability numbers are the reason anyone cares about the pricing. On DeepSWE v1.1, GLM-5.3-Flash posts 63.4 against GLM-5.2's 46.2. On AutomationBench, 48.8 against 26.2. On vision benchmarks like BabyVision and MVbench, it trails Gemini 3.7 Flash. This is agentic coding and browser-use work at genuine frontier-adjacent quality, licensed MIT, weights on Hugging Face, no enterprise contract required.

The strategic pattern is by now legible. In July, Goldman Sachs named GLM-5.2, DeepSeek, and ByteDance as the operative China AI frontier; this month, Alibaba's Qwen3-8-27B crossed three million Hugging Face downloads in three days under Apache 2.0. Chinese labs aren't competing to charge more for intelligence. They're competing to give it away faster than the closed frontier can amortize its training bills.

What's changed with GLM-5.3-Flash isn't the direction of that trend, it's the floor. A team of five now has a credible, MIT-licensed path to running near-Opus-class coding and agentic workloads at flash-tier economics. The anonymous-preview stunt is starting to look less like marketing and more like a standing invitation: rank us blind, then check the invoice.

Sources

  • https://docs.z.ai/release-notes/new-released
  • https://siliconangle.com/2026/08/26/z-ai-open-sources-ox-alpha-model-as-glm-5-3-flash/
  • https://www.marktechpost.com/2026/08/26/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context/
  • https://www.testingcatalog.com/z-ai-launches-glm-5-3-flash-under-mit-license/
  • https://www.eesel.ai/blog/glm-5-3-flash