Reviews · SEPTEMBER 16, 2026
Gemini 3.8 Live lands at $1.38 an hour, collapsing the voice-agent stack into one endpoint
Google's native speech-to-speech pair scored 82.6 on Artificial Analysis' Speech-to-Speech Quality Index and shipped September 15 at $0.005 per input minute and $0.018 per output minute — the first frontier voice model with a published per-minute price small teams can budget against.
Google shipped Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15 at $0.005 per minute of audio input and $0.018 per minute of output, roughly $1.38 an hour for a two-way conversation. That's the news. The interesting part is that a frontier voice model now has a per-minute list price a small business can put in a spreadsheet.
The benchmark card is easy to describe: Extended Thinking took the #1 slot on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6, posted 68.6% on the τ-Voice benchmark, 35.1% on Sierra's τ-Voice-banking benchmark, and 97.7% on Big Bench Audio. Extended Thinking's Quality Index score surpassed GPT-Live-1-Astra and Grok Voice Think Fast 2.0, per SiliconAngle, while the base 3.8 Live model came in second on Artificial Analysis' Speech Agent Arena. Google's post credits ServiceNow's EVA-Bench, run through the Live API on the Gemini Enterprise Agent Platform.
Benchmarks travel poorly. Pricing doesn't.
For most of 2024 and 2025, standing up a production voice agent meant paying separately for speech-to-text, an LLM turn, and a text-to-speech vendor, and then paying an integrator to make the round trip feel human. There was no single number to underwrite the business case against. Gemini 3.8 Live is billed as one endpoint: 16-bit PCM at 16kHz in, 24kHz PCM out, JPEG frames at up to 1 fps if you want vision, 97 languages with mid-conversation switching, and SynthID watermarking on every output. The Live API is live in the Gemini API and Google AI Studio today, with Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents handling the WebSocket layer. Salesforce, Genspark, and Lumeris are the named enterprise launch partners.
Run the math on a five-person services business fielding 40 hours of inbound voice a month, split evenly between listening and speaking. That's about $55 in API spend. Not a headcount decision. A line item.
It follows an earlier Gemini 3.8 Flash launch that held pricing flat at $0.75 per million tokens. The pattern rhymes with the way AWS priced S3 in 2006 and EC2 shortly after: publish a per-unit rate low enough that the buyer stops shopping and starts building. Cloud didn't win on features. It won when the CFO could model it.
Extended Thinking is already powering Gemini Live, Gmail, and Google Keep on the consumer side, which is where the model gets its conversational reps. The commercial story sits inside the pricing table.
Sources
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
- Google's new speech model Gemini 3.8 Live supports real-time reasoning
- Google Launches Gemini 3.8 Live and Extended Thinking Voice Models
- Gemini 3.8 Live Extended Thinking powers Gemini Live, Gmail, & Keep
- Google Launches Gemini 3.8 Live for Smarter, Free-Flowing AI Voice Chats