AI Model Report

Benchmarks · AUGUST 2, 2026

DeepSeek V4 Flash 0731: same 284B MoE, dramatically better agent scores from post-training alone

DeepSeek's July 31 public-beta refresh leaves architecture, parameter count, and price untouched but reports a 20.9-point Terminal-Bench 2.1 jump and a 47.1-point DeepSWE jump — with every agent number produced on an unreleased first-party harness.

By Linnea Halberg · Benchmarks desk · August 2, 2026

DeepSeek pushed DeepSeek-V4-Flash-0731 to Hugging Face under an MIT license on July 31, 2026, and opened a public-beta API the same day. The architecture didn't change: 284B total parameters, 13B active per token, a 1M context window, a 384K output ceiling on high and max reasoning. Pricing didn't change either: $0.14 per million cache-miss input tokens, $0.0028 on cache hits, $0.28 on output. What changed is every agent number on the card.

Against the April Preview, Terminal-Bench 2.1 climbs from 61.8 to 82.7, a 20.9-point gain. DeepSWE moves from 7.3 to 54.4, a 47.1-point swing that also leaves the V4-Pro Preview's 12.8 far behind. Nine agent benchmarks in total sit on the model card, all run at temperature 1.0 and top_p 0.95, all executed inside an unreleased first-party rig DeepSeek calls DeepSeek Harness in "minimal mode." No third party has reproduced any of them yet.

That caveat matters, because same-weights, better-scores is the story DeepSeek wants told. If the underlying model is unchanged and only post-training and scaffolding shifted, the delta is a claim about harness engineering as much as about capability. The comparison rows against Claude Opus 4.8 tell the honest version: 0731 trails 85.0 to 82.7 on Terminal-Bench 2.1, 25.7 to 25.2 on Agents' Last Exam, 69.7 to 54.2 on NL2Repo, and 71.7 to 59.6 on DSBench-Hard. Repository-scale coding still runs 12–16 points behind Opus.

Artificial Analysis, running its own eval stack rather than DeepSeek's, puts the model at 50 on Intelligence Index v4.1, up from 40 for the April Flash and one point below GPT-5.6 Luna's 51. AA-Omniscience moves from −23 to −16, driven by an 11-point drop in hallucination rate while raw accuracy stays flat at 37%, exactly what you'd expect from a same-size checkpoint learning to say "I don't know."

The pricing context is the part that gets underplayed. OpenAI cut GPT-5.6 Luna to $0.20 input and $1.20 output the day before DeepSeek's beta opened. Artificial Analysis estimates 0731 still runs about 60% cheaper per task once caching and reasoning-token overhead are factored in, helped by DeepSeek's roughly 98% first-party cache-hit discount. The competitive posture, in other words, is unchanged: match the frontier within a few points on published evals, undercut on price, ship weights under MIT.

The interesting artifact here isn't the score jump. It's that a Chinese lab now feels comfortable publishing nine agent benchmarks entirely on infrastructure nobody else can run, and letting the market take the number on faith for a beta window. Bloomberg covered the launch. benchlm.ai wrote it up. The harness stays inside the building.

Sources

  • https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
  • https://artificialanalysis.ai/articles/deepseek-v4-flash-0731-scores-50-on-the-artificial-analysis-intelligence-index-10-points-above-previous-deepseek-v4-flash
  • https://www.bloomberg.com/news/articles/2026-07-31/deepseek-unveils-public-beta-api-for-flagship-ai-model
  • https://xenospectrum.com/en/deepseek-v4-flash-0731-pricing/
  • https://benchlm.ai/blog/posts/deepseek-v4-flash-0731