Microsoft AI launched MAI-Transcribe-2-Streaming on October 1, 2026. The release included two new text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Artificial Analysis ranks the new transcription tool first among 38 models for both final and initial partial transcript accuracy.
In this article
The release
This model is the streaming counterpart to the batch version, MAI-Transcribe-2, which arrived in September. It supports 60 languages with automatic, continuous language detection. Audio streams in continuously, and text streams back while the speaker is still talking.
The model emits its first hypotheses, known as partials, just over 100ms after receiving audio. It revises those partials as context arrives, then commits a stable final transcript. An agent can therefore start reasoning or calling tools mid-sentence. Microsoft team states its internal tests show words appearing 2x faster than its closest competitor.
Artificial Analysis results
The AA-WER Streaming index uses about 8 hours of audio. The mix is AA-AgentTalk (50%), VoxPopuli (25%) and Earnings22 (25%). Latency is timed from the end of speech, as detected by SileroVAD.
- Final transcript: 2.5% WER at 0.13s after end of speech, #1 of 38 models.
- First partial transcript: 2.5% WER at 0.12s after end of speech, also #1.
- Runners-up: Grok Voice Transcribe 2.0 at 2.7% and 0.49s. Muse Voice Transcribe at 3.1% and 0.16s.
- Not the fastest: Cartesia Ink-2 (external endpoints) returns finals in 0.07s, but at 4.0% WER.
The first partial is as accurate as the final transcript. That matters for agents that act before the speaker finishes. Microsoft also places the model on the accuracy versus latency Pareto frontier.
Pricing
MAI-Transcribe-2-Streaming costs $0.54 per hour of audio. This is an introductory price through the end of 2026. Artificial Analysis normalizes it to $9.00 per 1,000 minutes. Batch MAI-Transcribe-2 costs $0.10 per hour. On streaming, Microsoft charges more than xAI and Meta, and roughly matches Google’s estimated rate.
Integration
Microsoft documents 2 integration paths. The Realtime API suits apps already using an OpenAI Realtime-compatible WebSocket. The Azure Speech SDK handles connection management, retries and audio streaming. Both return intermediate and final results.
The model is also available in the MAI Playground, through Vercel and Azure Voice Live. LiveKit support is listed as coming soon.
Microsoft pairs it with MAI-Voice-2.1-Flash for full voice loops. Flash generates 45s of audio at 150ms end-to-end latency for $15 per 1M characters. MAI-Voice-2.1 covers 23 languages and 26 locales at $22 per 1M characters.
Comparison: MAI-Transcribe-2-Streaming vs Closest Streaming Competitors
| Feature | MAI-Transcribe-2-Streaming | Grok Voice Transcribe 2.0 | Muse Voice Transcribe | Gemini 3.5 Transcribe Live |
|---|---|---|---|---|
| Developer | Microsoft AI | xAI | Meta Superintelligence Labs | |
| Released | Oct 1, 2026 | Sep 18, 2026 | Sep 1, 2026 | Aug 26, 2026 |
| AA-WER Streaming (final) | 2.5% | 2.7% | 3.1% | 4.0% |
| Time to final | 0.13s | 0.49s | 0.16s | Not reported by AA source cited |
| Streaming price / hour | $0.54 (intro) | $0.20 | $0.18 | ~$0.54 (token-billed estimate) |
| Languages | 60, continuous auto-detect | Dozens, auto-detect, mid-recording switch | 70+ trained, 25 verified | 85+, auto-detect |
| Speaker diarization in stream | Not stated | Included in API (streaming not confirmed) | Yes, 20+ speakers | Not supported in Live mode |
| Interface | Realtime API (WebSocket) + Azure Speech SDK | WebSocket | WebSocket + file endpoint | Live API (WebSocket) |
| Open weights | No | No | No | No |
Sources: Artificial Analysis, 9to5Mac (Muse pricing), DataCamp (Grok pricing), The Batch (Gemini pricing). Verified October 2, 2026.
What it means
Developers building voice agents can now trigger actions almost instantly. The model provides accurate text before the speaker finishes, allowing systems to respond without waiting for a full sentence to complete.




