AI Model Report

Benchmarks · AUGUST 17, 2026

Gemini 3.7 Flash lands three weeks after 3.6 with a 16-point DeepSWE jump and a half-price introductory tag

Google's third Flash-tier release in under two months clears 65.3% on DeepSWE v1.1 and 30.4% on AutomationBench at $0.75/M input tokens through year-end — while Gemini 3.5 Pro remains unshipped.

By Linnea Halberg · Benchmarks desk · August 17, 2026

Google shipped Gemini 3.7 Flash on August 13, 2026, three weeks after 3.6 Flash and while Gemini 3.5 Pro remains unshipped. Bloomberg's framing was blunt: the workhorse tier is now iterating faster than the flagship it's supposed to sit under.

The benchmark story is the reason the release happened at all. On DeepSWE v1.1, the model card reports a jump from 49.0% to 65.3%, a 16.3-point gain in three weeks. AutomationBench nearly doubles, 17.0% to 30.4%. FrontierCode 1.1 Main moves from 34.4% to 43.6%, and the GDP.pdf complex-document eval climbs from 22.0% to 34.0%. These aren't rounding-error updates; they're the kind of deltas that usually accompany a major version bump, not a point release.

Against the competitive set, the picture is more selective. Gemini 3.7 Flash posts a Code Arena Elo of 1588, ahead of Claude Sonnet 5 at 1541 and GPT-5.6 Terra at 1523. On AutomationBench it laps Sonnet 5 (10.7%) and edges GPT-5.6 Terra (23.6%). But Terminal-bench 2.1 goes to GPT-5.6 Terra at 87.4% versus 85.8%, and Claude Sonnet 5 still leads Agent's Last Exam at 33.3% to Gemini's 26.3%. Coding and workflow, yes. Agent desktop tasks, not yet.

Pricing is where the narrative management gets interesting. Through December 31, 2026, the introductory rate is $0.75 per million input tokens and $3.75 per million output. On January 1, 2027, that doubles to $1.50 and $7.50. The model card is explicit about the expiry, which is unusual candor.

It's also not the cheap option. GPT-5.6 Luna is priced at $0.20 input and $1.20 output. DeepSeek V4 Flash, whose repricing takes effect August 17, sits at $0.14 uncached input and $0.28 output, an output rate Memeburn notes is more than 13x lower than Gemini's introductory tier. Artificial Analysis clocks the model at roughly 340 output tokens per second, which is the more defensible pitch.

VentureBeat's rollout coverage names Glean, Dust, and LemonLime among the launch partners already routing production workflows through the model, with LemonLime's team highlighting the coding-eval gains as the reason for early adoption. Context window is 1M tokens input, 64K output. Gemini Spark, powered by 3.7 Flash, ships to Google AI Pro and Ultra subscribers but not the European Economic Area at launch.

The structural read: Google is treating the Flash tier as its release valve while Pro slips. SemiAnalysis has flagged the pattern before, and the 3.5 Pro delay is now doing the talking. When the mid-tier is the beat, the flagship is the problem.

Sources

  • https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/
  • https://deepmind.google/models/model-cards/gemini-3-7-flash/
  • https://www.bloomberg.com/news/articles/2026-08-13/google-debuts-new-gemini-flash-while-top-ai-model-still-delayed
  • https://venturebeat.com/technology/googles-gemini-3-7-flash-targets-coding-and-agents-with-a-50-introductory-price-cut
  • https://memeburn.com/gemini-3-7-flash-review-pricing-benchmarks-2026/