Benchmarks · SEPTEMBER 3, 2026
Gemini 3.8 Flash hits DeepSWE 73.7% at $0.75/M input, level with Opus 5
Google's third Flash release in six weeks matches Claude Opus 5 on long-horizon coding at roughly one-sixth the input-token price — until the introductory rate doubles on January 1.
Google released Gemini 3.8 Flash on September 2, 2026, its third Flash model in six weeks, and the benchmark line that matters is this: 73.7% on DeepSWE v1.1, the long-horizon software engineering suite, a few fractions of a percent behind Claude Opus 5 and a point ahead of GPT-5.6 Sol. The list price is $0.75 per million input tokens and $3.75 per million output.
That's the story, but the story has an expiration date.
The introductory tokens are held through the end of 2026, matching the pattern Google used on the Gemini 3.6 Flash, 3.5 Flash Lite, and 3.5 Flash Cyber release in July. On January 1, 2027, per-token pricing doubles to $1.50 and $7.50. Builders locking in agentic pipelines now are effectively pricing four months of frontier-adjacent inference against a January step-up they can already see.
The comparative economics are stark. Artificial Analysis clocks Gemini 3.8 Flash at $0.58 per Intelligence Index task and calls it "the cheapest model at its level of intelligence." Anthropic's Claude Fable 5.1 comes in at $3.76 on the same measure, roughly six times more. On the Intelligence Index itself, 3.8 Flash scores 59 in high-reasoning mode, level with GPT-5.6 Sol (xhigh) and Grok 4.6 (medium), three points up from 3.7 Flash. Opus 5 sits at 63, Fable 5.1 at 66. The frontier isn't being caught. The gap to it's being narrowed at a fraction of the price.
The lift comes from working harder, not thinking differently. Google's Tulsee Doshi and Raluca Ada Popa wrote in the launch post that "3.8 Flash works harder... executing extra reasoning steps, and calling tools iteratively." Artificial Analysis backs it out mechanically: about 30% more output tokens per task (~48K), time per task up from 2.2 to 2.5 minutes on high reasoning, throughput around 300 output tokens per second, and cost per task roughly 40% higher than 3.7 Flash despite unchanged per-token rates. The gains show up where iteration compounds: a 12-point jump to 45% on 𝜏³-Banking tool use, and The Register notes further agentic-workflow gains on Vals Finance Agent V2, Harvey's Legal Agent Benchmark, and HLE-Verified.
Built directly on 3.7 Flash, the model retains a 1M-token context window, 64K output tokens, and text, image, audio, and video inputs. A low-reasoning mode scores 52, matching 3.6 Flash on high reasoning at 30% lower cost and roughly a third the time.
This is the same pattern that reshaped API economics when OpenAI cut GPT-5.6 Sol pricing 20% in August: the second-tier model catches the top tier on the benchmarks buyers actually run, and the price collapses first. What used to require Opus-class budgets now runs on Flash-class ones. Until January 1.
Sources
- Google releases Gemini 3.8 Flash, its third Flash model in six weeks
- With Gemini 3.8 Flash, Google reminds everyone it's still in the race
- Google launches two Gemini 3.8 models with cutting-edge reasoning capabilities
- Gemini 3.8 Flash reaches the Intelligence vs. Cost per Task Pareto frontier
- Gemini 3.8 Flash, Model Card