AI Model Report

Benchmarks · AUGUST 1, 2026

DeepSeek V4-Flash-0731: a re-post-trained 284B beats its own 1.6T flagship

DeepSeek pushed a re-post-trained V4-Flash into public beta on July 31 — same architecture as the April preview, but scoring 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE, beating V4-Pro-Preview on all nine published agent benchmarks at a third of the price.

By Linnea Halberg · Benchmarks desk · August 1, 2026

DeepSeek moved its V4-Flash API into public beta on July 31, 2026, under the build designation V4-Flash-0731, and the interesting number isn't the release itself. It's that the 284B-total, 13B-active MoE now outscores DeepSeek's own V4-Pro-Preview on all nine of the agent benchmarks the company publishes, at roughly a third of the price. Same architecture as the April preview. New post-training run. That's the entire delta.

The headline numbers do the arguing. Terminal Bench 2.1 jumps from 61.8 on the Flash Preview to 82.7 on the 0731 build, putting it 2.3 points behind Anthropic's Opus-4.8 (85.0) and 10.6 points ahead of V4-Pro-Preview (72.1). DeepSWE goes from 7.3 to 54.4, a 645% increase that reads less like an improvement and more like a phase transition. Agents' Last Exam lands at 25.2, effectively tied with Opus-4.8's 25.7. Toolathlon hits 70.3, Cybergym 76.7, NL2Repo 54.2, and the internal DSBench-FullStack climbs from 37.0 to 68.7.

DeepSeek attributes the gains to a fresh round of post-training targeting agent capability. The TechTimes write-up is more specific: reinforcement learning against verifiable coding rewards, where code executes, tests pass or fail, and the update favors sequences leading to passing tests. This is the same recipe that made o1 legible to institutional buyers a year and a half ago, now applied to a mid-sized open-weight MoE under an MIT license with a 1M-token context window.

Artificial Analysis scores the model at 50 on its Intelligence Index v4.1 (a nine-eval composite spanning Terminal-Bench v2.1, τ³-Banking, SciCode, GPQA Diamond, Humanity's Last Exam, CritPt, AA-Omniscience, AA-LCR, and GDPval-AA v2), against a median of 25 for open-weight models of comparable size. Pricing sits at $0.14 per million input tokens and $0.28 per million output, versus a comparable-model median of $0.58 and $2.20. The full Intelligence Index run cost $72.02.

The catch sits in two places. First, verbosity: V4-Flash-0731 generated 210 million output tokens across the Intelligence Index, against a reference-field median of 100 million, a 2.1× ratio that widens to roughly 3.4× on reasoning-heavy subsets. Cheap per token isn't cheap if the model writes twice as much. Second, and structurally more important, every score in the DeepSeek changelog carries a footnote: the numbers were produced by "DeepSeek Harness minimal mode (to be released soon)" at max effort, topp=0.95, temperature=1.0. The harness is unreleased. TechTimes also notes that third-party aggregators such as OpenRouter may still be serving the April preview until routing updates propagate, meaning most developers currently benchmarking "V4-Flash" aren't benchmarking the 0731 build at all.

The lab is publishing frontier-adjacent agent numbers on an eval harness only it can run. That's the story worth watching, not the leaderboard.

Sources

  • https://api-docs.deepseek.com/updates/
  • https://www.bloomberg.com/news/articles/2026-07-31/deepseek-unveils-public-beta-api-for-flagship-ai-model
  • https://artificialanalysis.ai/models/deepseek-v4-flash
  • https://technode.com/2026/07/31/deepseek-puts-v4-flash-api-into-public-beta/
  • https://www.techtimes.com/articles/322513/20260731/deepseek-retrained-v4-flash-beats-its-flagship-pro-nine-agent-benchmarks.htm