AI Model Report

Reviews · AUGUST 17, 2026

Gemini 3.7 Flash ships 23 days after 3.6, jumps 16 points on DeepSWE

Google's fastest Flash cycle yet pairs a same-architecture refresh with a 50% introductory price cut — and no sign of Gemini 3.5 Pro.

By Karl Strauchman · Senior model reviewer · August 17, 2026

Google released Gemini 3.7 Flash on August 13, 2026, twenty-three days after 3.6 Flash, the tightest Flash-tier cycle Google has ever run. The model card describes it as "the next iteration in the Gemini 3 model family, featuring algorithmic improvements to its core reasoning foundation," which is corporate for: same architecture, retrained head, aimed squarely at coding and agents.

The benchmark deltas are the story. DeepSWE v1.1 goes from 49.0% to 65.3%, a 16.3-point jump. FrontierCode 1.1 Main moves 9.2 points to 43.6%. AutomationBench climbs 13.4 points to 30.4%, which is worth reading twice: roughly seven of ten multi-step automation tasks still fail. The frontier is closer than it was three weeks ago, and it's still not close.

Against the field, Google is winning where it wants to win and quietly losing where it doesn't. On WebDev Arena, 3.7 Flash posts an Elo of 1588, ahead of 3.6 Flash at 1538, Claude Sonnet 5 at 1541, and GPT-5.6 Terra at 1523. On GDP.pdf, it hits 34.0%, six points ahead of Sonnet 5 and 9.3 ahead of GPT-5.6 Terra. On Terminal-bench 2.1, it lands at 85.8%, trailing GPT-5.6 Terra's 87.4%. On Agent's Last Exam, Sonnet 5's 33.3% clears 3.7 Flash's 26.3%. The pattern: Flash is now the web-dev and document-reasoning leader at its price tier, and a credible second on longer-horizon agent tasks.

The pricing is the lever. Introductory rates are $0.75 per million input tokens and $3.75 per million output, roughly a 50% cut from 3.6 Flash. Artificial Analysis clocks a blended price of $0.58 per million tokens against 3.6's $1.16, and ranks 3.7 Flash first of 186 models on output speed at 340.1 tokens per second. On January 1, 2027, pricing snaps to $1.50 input and $7.50 output. Developers have a five-month window to lock in dependencies before the meter changes.

Tulsee Doshi, senior director of product management at Google, framed the release around the 950 million monthly users of the Gemini app. The context window stays at 1,048,576 input tokens with a 65,536 output cap, a March 2026 knowledge cutoff, and three exposed thinking levels defaulting to medium. It ships across the Gemini API, Android Studio, Google Antigravity, the Gemini Enterprise Agent Platform, and Spark.

Notable by absence: Gemini 3.5 Pro. Google is running the Flash tier as its narrative engine while Pro stays quiet, an inversion of the usual model-launch choreography where the flagship carries the story. The economics of inference now favor whoever ships the cheapest fast model that clears the bar. Google just moved the bar and cut the price on the same day.

Sources

  • https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/
  • https://deepmind.google/models/model-cards/gemini-3-7-flash/
  • https://siliconangle.com/2026/08/13/google-launches-gemini-3-7-flash-coding-ai-agent-projects/
  • https://venturebeat.com/technology/googles-gemini-3-7-flash-targets-coding-and-agents-with-a-50-introductory-price-cut
  • https://datanorth.ai/news/google-releases-gemini-3-7-flash