Open Source · AUGUST 31, 2026
GLM-5.3-Flash Debuts at $0.075/$0.25 as DeepSeek Warns of 'Significant' Hike
Z.ai's MIT-licensed 320B-A18B MoE landed August 26 at half its list price through September 9, while DeepSeek's price-floor dominance shows visible cracks.
Three weeks after DeepSeek told developers to brace for a "significant" API price hike, Z.ai shipped GLM-5.3-Flash as a 320B-total, 18B-active MoE under MIT at a promotional $0.075 input and $0.25 output per million tokens, undercutting DeepSeek's incumbent V4-Flash input rate by roughly half. The API price floor, long anchored by DeepSeek's willingness to price near cost, is being pried loose from two directions at once.
DeepSeek's August 6 developer-platform notice, reported by Bloomberg per Dataconomy, warned of a "significant" price hike across its API services "in the near future," and told customers a "significant increase" is expected and that they "should plan their usage accordingly." No replacement rate card was published, no effective date given. Dataconomy noted it was the second pricing change in under a month, following mid-July peak and off-peak tiers. AI Pricing Guru's tracker, updated August 14, then recorded a subsequent rate card with peak and off-peak billing beginning 16:00 UTC on August 16. Neither AI Pricing Guru nor BenchLM's August 12 sync captured the "significant" replacement multiplier itself.
V4-Flash-0731 today still carries $0.14 input, $0.28 output, and $0.0028 cache-hit input per million tokens, per BenchLM. BenchLM also notes DeepSeek hadn't published the 0731 weights when it checked on July 31, even though V4 Preview shipped as open weights back in April.
Enter GLM-5.3-Flash. FelloAI's model summary puts its Intelligence Index at 57 against 52 for the V4-Flash July 31 checkpoint, and confirms the list price of $0.15/$0.50 is halved through September 9. The GLM-5.2 lineage was already the open leaderboard's benchmark, and in July Goldman Sachs flagged it alongside DeepSeek as the Chinese frontier tightened. What's new is the delivery: MIT license, better benchmark, lower headline price than the model that defined the floor.
Compare that with the American end of the market. OpenAI put GPT-5.6 Luna at $0.20/$1.20 after an 80% cut three weeks after launch, then followed with a broader 20% API cut across GPT-5.6 Sol. Luna's output rate still runs roughly 4.3x V4-Flash's. The Chinese-open and American-closed markets are converging on token price from opposite sides, and the compression is happening inside a single month.
For a small team running outreach or content workflows on V4-Flash economics, the math has an expiration date. Promotional pricing ends September 9. DeepSeek's replacement multiplier remains unpublished. The cheapest usable token in production today may not be the cheapest usable token in ten.
Sources
- Best AI Models in August 2026: Updated Rankings and Comparisons
- DeepSeek Warns Developers Of Significant API Price Increases
- DeepSeek signals 'significant' price hike amid surge in demand for low-cost AI models
- DeepSeek API Price Increase: What It Means (August 2026)
- DeepSeek API Pricing (August 2026): V4 Pro & Flash Rates