Benchmarks · SEPTEMBER 2, 2026
Gemini 3.8 Flash Hits 73.7% on DeepSWE v1.1 at $0.75/$3.75 — Third Google Flash Release in 43 Days
Google's new Flash model ties Claude Opus 5 on long-horizon coding at roughly one-seventh the price, with the introductory rate expiring December 31.
Google DeepMind shipped Gemini 3.8 Flash on September 2, its third Flash release in 43 days after 3.6 Flash in July and 3.7 Flash on August 13. The cadence is the story as much as the model. Google is now iterating on the mid-tier at roughly the frequency competitors iterate on quarterly earnings framing.
The numbers matter because they land in agentic territory. Gemini 3.8 Flash posts 73.7% on DeepSWE v1.1 per Google's evaluation PDF, within a rounding error of Claude Opus 5's 74.0%. Terminal-Bench 2.1 climbs to 90.8% from 3.7 Flash's 81.6%, per DataCamp. τ³-Banking, a tool-use benchmark, jumps 12 points to 45%, per Artificial Analysis. HLE-Verified sits at 54.9%, and 9to5Google flags additional gains on Vals Finance Agent V2 and Harvey's Legal Agent Benchmark. The Artificial Analysis Intelligence Index rises three points at high reasoning, to 59, with Artificial Analysis crediting the gain primarily to agentic evaluations rather than raw reasoning.
Pricing is where the strategic play sharpens. Introductory rates are $0.75 per million input tokens and $3.75 per million output, roughly one-seventh Claude Opus 5's tier and well below the $5/$30 GPT-5.6 Sol schedule DataCamp lists, or Claude Fable 5.1 at $10/$50. Two asterisks apply. Both rates double on January 1. And cost per Intelligence Index task at high reasoning is $0.58, about 40% above 3.7 Flash's $0.40, because average output per task rose 30% to 48,000 tokens and time per task drifted from 2.2 to 2.5 minutes.
Google addresses this obliquely in its launch post: "At times, the model might use more tokens to maximize performance, especially at higher effort levels." Translation, verbosity is now a pricing variable, not a bug.
There's also a benchmark discrepancy worth flagging. Fello AI and AlphaCorp AI note the model card reports 71.0% on DeepSWE v1.1, not the 73.7% headline figure. Two numbers, one model, one benchmark. Buyers should treat the delta as a reminder that headline scores travel faster than methodology.
The structural read: Google is compressing the release cycle to keep Flash inside every serious agent buildout's evaluation shortlist, while pricing the introductory window to make migration decisions feel time-boxed against December 31. That mirrors, in structure if not scale, how open-weights alternatives are shipping preview-first to force the same evaluation.
Anyone running long-horizon agent pipelines should benchmark 30 to 50 real tasks before migrating. The per-token price is generous. The per-task price is already climbing.
Sources
- https://artificialanalysis.ai/articles/gemini-3-8-flash
- https://9to5google.com/2026/09/02/gemini-3-8-flash-launch/
- https://www.datacamp.com/blog/gemini-3-8-flash-cyber
- https://felloai.com/gemini-3-8-flash/
- https://alphacorp.ai/blog/gemini-3-8-flash-launch-pricing-benchmarks-and-whats-new