Reviews · SEPTEMBER 27, 2026
Gemini 3.8 Flash TTS Ships at 1.35 Cents a Minute — Until January 1 Doubles the Bill
Google's new text-to-speech pair replaces a 30-voice preset library with a prompt-driven design surface and 2,000+ production voices, and prices the flagship 70% below the preview it replaces — for 99 days.
Google shipped Gemini 3.8 Flash TTS and Flash-Lite TTS into the Gemini API and AI Studio on September 23, 2026, priced at $9 per million audio output tokens for the flagship, roughly 1.35 cents per minute of generated speech at Google's stated 25 audio tokens per second. On January 1, 2027, every rate doubles. The introductory window runs 99 days, and the pricing behavior echoes the Gemini 3.8 Flash text launch three weeks earlier, which also held the prior generation's $0.75/M input rate before a scheduled doubling.
The framing choice matters. Google isn't lowering a price; it's staging one, then re-anchoring upward against usage already booked into pipelines. Flash-Lite TTS lands at 0.9 cents per minute, which CellCog clocks at 70% below the $20/M gemini-3.1-flash-tts-preview it replaces.
The catalog got the same treatment as the price sheet. The old 30 curated voices are gone, replaced by what Google's launch post calls "2,000+ production-ready voices" (the changelog says 150+, an internal inconsistency CellCog flagged on release day) and a new Voice Design surface that generates a persona from a natural-language prompt. Leland Rechis and Alan Cowen, Google's Group PM and Director of Research Science on the launch, describe the shift as moving "from static presets into a dynamic creative studio." A 30-second sample is enough to replicate a voice; stored customs cap at 200 per project and persist one year, stateless replication keys expire in seven days, and every clip carries a SynthID watermark with C2PA credentials layered onto replicated voices.
On the arena, the pricing looks even more aggressive. CellCog's read of Artificial Analysis' blind voice leaderboard puts Flash TTS at 1260 Elo, second behind Cartesia Sonic 3.6 at 1273, with ±17 intervals that render the gap a statistical tie. Flash-Lite lands sixth at 1235. ElevenLabs v3 Conversational, the house's strongest entry, sits twelfth at 1196. Google also claims first on Hume AI's Voice Design Benchmark at 71.4 overall and 60.8 on accent modeling.
The economics collapse a category. An hour of narration costs 81 cents on Flash TTS at the introductory rate; a 30-second ad read runs under a penny. Small operators who had priced voice-over as a studio line item now price it as an API call, and the January cliff turns Q4 into a procurement decision rather than an experimentation window. That's the point of a 99-day rate.
The models accept 8K tokens of text and return up to 64K tokens of audio, per Unite.AI, and Google's model card lists Gemini 3 Pro as the base. After January 1, the flagship rises to 2.7 cents per minute and Flash-Lite to 1.8. Every buyer who benchmarked at 1.35 will renew at double, which is the pricing story the launch was designed to tell.
Sources
- https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/
- https://aistudio.google.com/learn/gemini-3-8-flash-tts-developer-guide
- https://www.unite.ai/google-rolls-out-gemini-3-8-speech-models-in-api-and-ai-studio/
- https://www.artificialintelligence-news.com/news/google-gemini-3-8-flash-tts-voice-models/
- https://cellcog.ai/blog/gemini-3-8-flash-tts/