Reviews · JULY 22, 2026
Google ships three Flash-tier Geminis and confirms 3.5 Pro is still slipping
Gemini 3.6 Flash lands at $1.50/$7.50 per million tokens with 17% fewer output tokens than 3.5 Flash, Flash-Lite arrives at 350 tok/s, a government-only Cyber variant enters CodeMender — and 3.5 Pro is still in partner testing after missing June.
Google DeepMind pushed three Flash-tier Gemini models on July 21, 2026: Gemini 3.6 Flash, 3.5 Flash-Lite, and a government-only variant called 3.5 Flash Cyber. The frontier model everyone was actually waiting on, Gemini 3.5 Pro, wasn't among them. Bloomberg had already reported five days earlier that Pro was "months behind schedule" after a late-June training-data refresh produced results insiders called "disappointing."
The Flash release is real work, priced to fight. 3.6 Flash lists at $1.50 per million input tokens and $7.50 per million output, down from the $9 output rate on 3.5 Flash, and Google claims a 17% reduction in output tokens consumed on the Artificial Analysis Index. The benchmark deltas are consistent: 49% on DeepSWE against 37% for the prior generation, with up to 65% fewer output tokens; 63.9% on MLE Bench versus 49.7%; 83% on OSWorld-Verified versus 78.4%; 1421 on GDPval-AA versus 1349. Knowledge cutoff moves from January 2025 to March 2026.
Flash-Lite is the throughput play. At $0.30 in and $2.50 out per million tokens, Google clocks it at 350 output tokens per second on the Artificial Analysis Index. It posts 54% on Terminal-Bench 2.1 (from 31% for 3.1 Flash-Lite), 72.2% on GDM-MRCR v2 (from 60.1%), 1140 on GDPval-AA v2 (from 642), and 54.2% on SWE-Bench Pro against 3 Flash's 49.6%.
The Cyber variant is the strategic curiosity. Google says it's "exclusively available to governments and trusted partners" and pairs with the company's CodeMender agent. On CyberGym, Google claims "competitive performance at the frontier"; on the internal Big Sleep evaluation, which probes vulnerabilities in Chrome and Safari codebases, the company says Cyber "significantly surpassed" both mainline 3.5 and 3.6 Flash. This is the first time Google has shipped a Gemini SKU gated by customer identity rather than usage tier, which locates it inside the same procurement channel Palantir and Anduril have spent a decade building out.
The Pro absence is doing more talking than the release. Tulsee Doshi, senior director of product management at Google DeepMind, framed the day around a "most ambitious pre-training run yet, for Gemini 4," and said Pro would arrive "as soon as it's ready." Product lead Logan Kilpatrick told TechCrunch it would "land soon." Neither committed to a window.
Skipping the tentpole to tease the generation after it's a familiar recovery pattern. Intel ran a version of it during the 10nm delays, promising 7nm process leadership while shipping refreshed 14nm parts year after year. Alphabet shares fell on the CNBC report about the Pro delay on July 16; the Flash triple-header five days later is the response, and it's a strong one on price and benchmarks. It just isn't the response investors were pricing in.
Sources
- https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/
- https://techcrunch.com/2026/07/21/google-releases-three-new-gemini-models-but-no-3-5-pro/
- https://www.bloomberg.com/news/articles/2026-07-16/google-gemini-launch-delayed-as-tech-falls-short-of-internal-goals
- https://www.cnbc.com/2026/07/16/alphabet-stock-gemini-3-5-pro-ai.html
- https://9to5google.com/2026/07/21/gemini-3-6-flash-launch/