Open Source · SEPTEMBER 11, 2026
DeepSeek V4.1 Flash: 552B MoE, MIT weights, and $0.15/M off-peak input
DeepSeek's new Causal Encoder-Decoder Flash tier posts 74.2 on DeepSWE v1.1 and 54.8 on AutomationBench while pricing input at a fraction of GPT-5.6 Sol and Claude Opus 5.
DeepSeek released V4.1 Flash on September 10, 2026, publishing the weights under an MIT license and pricing input tokens at $0.15 per million off-peak, roughly one-fortieth what Anthropic charges for Opus 5, and the first open-weight release to edge past GPT-5.6 Sol on an agentic-coding benchmark.
The scores explain the noise. On DeepSWE v1.1, V4.1 Flash posts 74.2, sitting 0.2 points above Claude Opus 5 (74.0) and 1.2 points above GPT-5.6 Sol (73.0). On Automation-Bench it lands at 54.8, 4.5 points clear of Opus 5. Terminal-Bench 2.1 comes in at 90.6, CyberGym at 88.1, GPQA Diamond at 90.9. The frontier labs still win on harder reasoning suites, Opus 5 takes HLE 56.3 to 36.8 and Terminal-Bench 3.0 43.3 to 30.0, but the ceiling gap has narrowed to a small handful of benchmarks.
The pricing is where the release becomes structurally interesting. Peak input runs $0.30 per million and peak output $1.20; off-peak (everything outside weekdays 01:00–04:00 and 06:00–10:00 UTC) drops those to $0.15 and $0.60. Cache-hit input off-peak is $0.003 per million. Opus 5 lists at $5 input, $25 output. GPT-5.6 Sol in Standard mode: $5 and $30. DataCamp pegs V4.1 Flash at roughly 30× cheaper on input and 40× cheaper on output than GPT-6 Astra or Claude Fable 5.1 at peak rates.
Fifty million cached input tokens costs about $0.15 on V4.1 Flash off-peak, versus about $25 on Opus 5. For a small team running lead-research agents or outreach pipelines against a stable corpus, that's the difference between a hobby project and production.
"Flash" is now a misnomer at the infrastructure layer. The architecture is a 552B-parameter Causal Encoder-Decoder MoE, activating 8B during prefill and 16B during decode, with a one-million-token context. The predecessor V4-Flash carried a 284B backbone. DeepSeek invites organizations at roughly 2,000 GPUs of capacity to contact them directly about self-hosting, which locates practical on-prem deployment somewhere north of most Series B budgets. The MIT license matters for fine-tuning economics; it doesn't shrink the weights.
VentureBeat's coverage cites Pulse data indicating 53% of enterprises above 100 employees don't rigorously track compute cost and ROI, which is one reason API price wars keep working as marketing. The reference-class here's Z.AI's GLM-5.3 Flash release, another MIT-licensed MoE playing the same disruption ladder, and Anthropic's Fable 5.1 automation gains at the opposite end of the price curve.
On September 14, starting 12:00 Beijing Time, deepseek-v4-pro requests route to V4.1 Flash at V4.1 Flash rates until V4.1 Pro ships. The retirement is the tell: DeepSeek is confident the cheaper model absorbs the older Pro workload cleanly, and it's pricing accordingly.