AI Model Report

Open Source · AUGUST 2, 2026

DeepSeek V4-Flash-0731 ships: 284B MoE, $0.28/M out, and a 50 on the AA Intelligence Index

DeepSeek's July 31 public-beta API for V4-Flash lands a reasoning-tier open-weights model at roughly a fifth of the median open-weights output price, one week after Anthropic put Opus 5 within striking distance of Fable 5 at half the cost.

By Lars Iverson · Open source & model weights · August 2, 2026

DeepSeek opened a public-beta API for V4-Flash-0731 on July 31, pricing max-effort reasoning at $0.14 per million input tokens and $0.28 per million output tokens, against category medians of $0.43 in and $1.20 out for open-weights models of similar size. Artificial Analysis clocked the model at 50 on its composite Intelligence Index, double the median of 25 for its size class, with a blended 7:2:1 cache/input/output rate of $0.06 per million tokens and a 1M-token context window.

The architecture is a 284-billion-parameter Mixture-of-Experts routing to 13 billion active parameters per forward pass. Running the full Intelligence Index cost $72.02 across 210M output tokens, more than twice the 100M-token median for the benchmark, evidence that V4-Flash is willing to think longer, and can afford to.

Seven days earlier, Anthropic shipped Claude Opus 5, which the company said "comes close to the frontier intelligence of Claude Fable 5 at half the price." Opus 5 scored 162.1 on Anthropic's internal ECI (95% CI 158.0–167.3, n=40), placing it alongside Mythos 5 on AI R&D tasks. The system card is unusually candid about where the improvement lives: gains are "concentrated in engineering execution rather than research judgment." On an internal trading benchmark, Anthropic reported "roughly a seventh of the reasoning tokens and under half the latency" versus the prior model, and 26% fewer tokens on average at max reasoning compared to Opus 4.8.

TechCrunch's Russell Brandom noted that Opus 5 arrived only two months after Opus 4.8, released May 28. That cadence, and the framing of Opus 5 as an efficiency step rather than a capability leap, is the story both labs are telling this month in different accents.

DeepSeek's own pitch, posted to its site and relayed by Bloomberg, promises "Significantly enhanced agent capabilities, with benchmark results far exceeding V4-Pro-Preview." No published numbers back the agent-tasks claim in the materials reviewed, which is worth flagging: the frontier-adjacent price story is verifiable, the agent story isn't yet.

What both releases share is a pivot away from parameter-count theater. Opus 5's engineering-execution framing and V4-Flash's aggressive post-training economics point to the same conclusion that Meta's Llama 3.1 405B release quietly demonstrated in mid-2024: raw scale is table stakes, and the differentiated work is happening in distillation, routing, and reasoning-token efficiency. Agentic-platform buyers reading benchmark tables at Glean, Dust, and LemonLime now have an open-weights option scoring a 50 for roughly a fifth of what the median open-weights competitor charges to output a token.

The interesting asymmetry: Anthropic sold Opus 5 as cheaper Fable 5. DeepSeek shipped something that makes that comparison look expensive.

Sources