AI Model Report

Model Releases · OCTOBER 4, 2026

Microsoft's MAI-Transcribe-2-Streaming lands at 2.5% WER and 0.13s to final, closing the voice-agent loop on Azure

Microsoft shipped its first streaming STT model on October 1 alongside two refreshed TTS variants, giving developers a complete listen-reason-speak stack on one subscription — at introductory pricing that reshapes who can afford a 24/7 voice agent.

By Karl Strauchman · Senior model reviewer · October 4, 2026

Microsoft AI shipped MAI-Transcribe-2-Streaming on October 1, alongside refreshed MAI-Voice-2.1 and MAI-Voice-2.1-Flash text-to-speech models, and the real news isn't the benchmark numbers. It's that the Azure voice-agent stack no longer requires a third-party STT vendor to close the loop.

The headline figures are respectable. The streaming transcriber posts 2.5% WER on Artificial Analysis's AA-WER Streaming index with 0.13 seconds to final transcript after end-of-speech, and 0.12 seconds to first-partial, which Microsoft says is enough for a #1 placement on the index. At publication, TechTimes noted that ranking hadn't yet surfaced on the publicly visible leaderboard, where Grok Voice Transcribe 2.0 (2.7% WER, 0.49s to final) still held the top line. Muse Voice Transcribe sits at 3.1% WER and 0.16s; Cartesia's Ink-2 is faster at 0.07s but trades accuracy at 4.0% WER. The benchmark audio runs about eight hours, weighted 50% AA-AgentTalk, 25% VoxPopuli, 25% Earnings22, with SileroVAD used to mark end-of-speech.

SiliconAngle reported that first-hypothesis times average 320 milliseconds and depend on network conditions, versus the just-over-100ms figure Microsoft cites for initial partials. The gap between benchmark and production latency is the number that actually matters when you're sitting on a call.

On the TTS side, MAI-Voice-2.1 covers 23 languages and 26 locales at $22 per million characters. The Flash variant generates up to 45 seconds of audio at 150ms end-to-end latency, prices in at $15 per million characters, and claims roughly 55% faster inference and 60% lower cost than comparable models, per Microsoft.

TechTimes assembled the full-loop budget: ~130ms committed transcript, 300–500ms for a fast reasoning step, 150ms to first audio from Flash. Total conversational turn, under good conditions, 580–780 milliseconds. That's inside the window where humans don't notice they're talking to software.

The pricing structure is where the strategic read sharpens. Batch MAI-Transcribe-2 launched in September 2026 at $0.10 per audio hour. The streaming variant debuts at $0.54 per audio hour, introductory, through December 31, 2026. That's a 5.4× premium, which TechTimes attributes to the cost of holding a persistent WebSocket per active session rather than queue-processing files. It's still cheap enough that a small team can run an inbound voice agent on variable cost, paying only when calls arrive.

What's being consolidated is a market structure. The specialist streaming-STT vendors, Deepgram, AssemblyAI, ElevenLabs on the TTS side, built businesses on the fact that hyperscalers didn't have competitive real-time models. Microsoft now does, priced into the same Foundry subscription that already holds the reasoning layer. The 2017 playbook, where AWS absorbed middleware categories by shipping first-party equivalents at infrastructure margins, is the obvious parallel. The specialists survived that cycle by moving upmarket into verticals and tooling. The question is whether 2.5% WER at $0.54 an hour leaves enough oxygen to repeat the trick.

Sources

  • https://microsoft.ai/news/our-first-streaming-transcription-model/
  • https://siliconangle.com/2026/10/01/microsoft-targets-ultra-realistic-voice-agents-with-its-first-streaming-transcription-model/
  • https://www.techtimes.com/articles/328501/20261002/microsoft-ships-first-streaming-transcription-model-completing-voice-agent-pipeline-foundry.htm
  • https://www.marktechpost.com/2026/10/02/microsoft-ai-releases-mai-transcribe-2-streaming-1-real-time-speech-to-text-model-on-artificial-analysis/
  • https://www.unite.ai/microsoft-launches-mai-transcribe-2-streaming-and-two-mai-voice-models/