Open Source · AUGUST 29, 2026
GLM-5.3-Flash lands MIT-licensed at ~$0.10/M blended after a week atop OpenRouter as Ox Alpha
Z.ai's 320B/18B-active multimodal MoE ships day-one MIT weights and a 1M-token context, arriving days after the full GLM-5.3 checkpoints cleared a two-week cyber-capability safety hold.
Z.ai unmasked "Ox Alpha," the anonymous model that spent the last week topping OpenRouter and OpenCode leaderboards, as GLM-5.3-Flash: a 320B-parameter / 18B-active multimodal MoE shipping under MIT license with a 1M-token context and blended pricing that lands near $0.10 per million tokens. The reveal came the same week Z.ai finally published FP8 and BF16 checkpoints for the full 744B-class GLM-5.3 on August 28, closing out a staged rollout that had been paused for roughly two weeks after hosted launch.
Kingy.ai's release tracking attributes that delay to unexpectedly fast cyber-capability gains during post-training. In the current climate around Chinese frontier releases, that framing does real work: it lets Z.ai signal safety discipline without conceding the underlying capability jump. The pause becomes evidence of seriousness rather than an embarrassment.
The pricing is the story small teams will actually feel. Standard API rates run $0.15 per million input tokens, $0.03 cached, and $0.50 output, per MarkTechPost, with a 50% launch discount active through September 9, 2026, per Progressive Robot. That's roughly a 9x reduction against the ~$0.90/M blended rate for the full GLM-5.3, and considerably more against proprietary flagships. For the first time, long-horizon agentic coding, the kind that quietly burns tokens iterating on a build for hours, sits inside the budget of a two-person shop.
The benchmarks Z.ai chose to publish tell a coherent story. On its own Z.ai Code Bench v1.0, GLM-5.3-Flash scores 29.0 at max effort against Claude Opus 4.8's 29.5, per OfficeChai. On Terminal-Bench 2.1 it posts 84.3, with Opus 4.8 at 85.0 and GPT-5.6 Terra ahead at 87.4. On DeepSWE v1.1 it hits 63.4 against predecessor GLM-5.2's 46.2. Artificial Analysis clocks its Intelligence Index at 57, matching Opus 4.8, though measured throughput of 49.4 tokens/sec sits under the 64.8 median. These are self-reported figures; independent reproductions over the next two or three weeks will decide how much of the parity survives contact with other people's evals.
The MIT license is the multiplier. The FP8 checkpoint weighs in around 331 GB and needs roughly 386 GB of VRAM at TP4 to serve, which keeps casual self-hosting out of reach, but B2B software builders can embed the weights in customer-facing products without licensing friction. Combined with Goldman Sachs' recent framing of China's AI frontier and the three-million-download surge behind Alibaba's Apache-2.0 Qwen3, the pattern is legible: Chinese labs are competing on permissive licensing as aggressively as on scores. Ox Alpha's week of anonymous benchmarking wasn't a marketing stunt. It was proof of work.