AI Model Report

Benchmarks · SEPTEMBER 12, 2026

DeepSeek V4.1 Flash undercuts the frontier on cache-hit pricing

A 552B MoE with an asymmetric causal encoder-decoder, MIT-licensed weights, and a $0.003/M off-peak cached-input rate — DeepSeek retires V4 Pro to route into it.

By Linnea Halberg · Benchmarks desk · September 12, 2026

DeepSeek shipped V4.1-Flash on September 10, priced cached input at $0.003 per million tokens off-peak, and announced V4 Pro would be retired starting September 14 by routing its traffic into the new model. MiniMax and Z.ai closed off more than 8% in Hong Kong; Alibaba fell more than 2%. The market read the release as another repricing event, and it isn't wrong.

The headline number to internalize isn't the $0.15 uncached input rate or the $0.60 output rate. It's 890 bytes per token. That's where V4.1-Flash's global KV cache lands, roughly a quarter of V4-Flash's footprint, and it's the mechanical reason a 500,000-token reusable prefix reused across 100 requests, 50 million cached tokens in total, costs $0.15 on V4.1-Flash at off-peak rates versus $15 on Kimi K3, $20 on GPT-5.6 Sol, and $25 on Claude Opus 5.

Under the hood is an asymmetric Causal Encoder-Decoder: 20 encoder layers, 20 decoder layers, decoder KV projected from the encoder's final hidden states rather than recomputed per layer. Compressed Sparse Attention 2 sits alongside it. The 552-billion-parameter MoE backbone activates 8B parameters per token on prefill and 16B on decode. Prefill compute halves for long sequences. Single-token decode FLOPs rise only about 25% as context grows 256-fold from 4K to 1M.

Benchmarks land where they need to. V4.1-Flash scores 74.2 on DeepSWE v1.1 against Claude Opus 5 at 74.0 and GPT-5.6 Sol at 73.0, and 88.1 on CyberGym. A 1-million-token context window, weights on Hugging Face under MIT license, peak rates at double the off-peak schedule.

The relevant comparison isn't Claude Opus 5 on coding evals. It's the shape of the agentic loop that reads the same playbook, the same CRM extract, the same prospect history on every step. That loop was economically ruinous at frontier prices and merely expensive at flash prices. At $0.003 per million cached tokens off-peak, the loop is a rounding error. VentureBeat's number that 47% of enterprises rigorously track AI compute cost suggests the other 53% are about to discover why they should've.

The parallel worth naming is the 2023 collapse in vector-database pricing after pgvector went mainstream: a capability that had justified a standalone budget line stopped justifying one, and the workloads reorganized around the new floor. V4.1 Pro is expected next, and it'll reset the top of the comparison. The floor it's being built on top of is the story.

Sources

  • https://www.bloomberg.com/news/articles/2026-09-10/deepseek-s-new-low-cost-model-deals-a-fresh-blow-to-openai-z-ai
  • https://venturebeat.com/technology/deepseek-v4-1-flash-debuts-with-0-003-1m-off-peak-cached-input-rate-and-benchmarks-eclipsing-gpt-5-6-sol-claude-opus-5
  • https://dataconomy.com/2026/09/11/deepseek-v4-1-flash-ultralow-token-pricing/
  • https://www.datacamp.com/blog/deepseek-v4-1-flash
  • https://www.intelligentliving.co/deepseek-v4-1-flash-pricing-release/