Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis

Microsoft AI launched MAI-Transcribe-2-Streaming on October 1, 2026. The release included two new text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Artificial Analysis ranks the…

By Vane October 3, 2026 2 min read
Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis

Microsoft AI launched MAI-Transcribe-2-Streaming on October 1, 2026. The release included two new text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Artificial Analysis ranks the new transcription tool first among 38 models for both final and initial partial transcript accuracy.

The release

This model is the streaming counterpart to the batch version, MAI-Transcribe-2, which arrived in September. It supports 60 languages with automatic, continuous language detection. Audio streams in continuously, and text streams back while the speaker is still talking.

The model emits its first hypotheses, known as partials, just over 100ms after receiving audio. It revises those partials as context arrives, then commits a stable final transcript. An agent can therefore start reasoning or calling tools mid-sentence. Microsoft team states its internal tests show words appearing 2x faster than its closest competitor.

Artificial Analysis results

The AA-WER Streaming index uses about 8 hours of audio. The mix is AA-AgentTalk (50%), VoxPopuli (25%) and Earnings22 (25%). Latency is timed from the end of speech, as detected by SileroVAD.

  • Final transcript: 2.5% WER at 0.13s after end of speech, #1 of 38 models.
  • First partial transcript: 2.5% WER at 0.12s after end of speech, also #1.
  • Runners-up: Grok Voice Transcribe 2.0 at 2.7% and 0.49s. Muse Voice Transcribe at 3.1% and 0.16s.
  • Not the fastest: Cartesia Ink-2 (external endpoints) returns finals in 0.07s, but at 4.0% WER.

The first partial is as accurate as the final transcript. That matters for agents that act before the speaker finishes. Microsoft also places the model on the accuracy versus latency Pareto frontier.

Pricing

MAI-Transcribe-2-Streaming costs $0.54 per hour of audio. This is an introductory price through the end of 2026. Artificial Analysis normalizes it to $9.00 per 1,000 minutes. Batch MAI-Transcribe-2 costs $0.10 per hour. On streaming, Microsoft charges more than xAI and Meta, and roughly matches Google’s estimated rate.

Integration

Microsoft documents 2 integration paths. The Realtime API suits apps already using an OpenAI Realtime-compatible WebSocket. The Azure Speech SDK handles connection management, retries and audio streaming. Both return intermediate and final results.

The model is also available in the MAI Playground, through Vercel and Azure Voice Live. LiveKit support is listed as coming soon.

Microsoft pairs it with MAI-Voice-2.1-Flash for full voice loops. Flash generates 45s of audio at 150ms end-to-end latency for $15 per 1M characters. MAI-Voice-2.1 covers 23 languages and 26 locales at $22 per 1M characters.

Comparison: MAI-Transcribe-2-Streaming vs Closest Streaming Competitors

FeatureMAI-Transcribe-2-StreamingGrok Voice Transcribe 2.0Muse Voice TranscribeGemini 3.5 Transcribe Live
DeveloperMicrosoft AIxAIMeta Superintelligence LabsGoogle
ReleasedOct 1, 2026Sep 18, 2026Sep 1, 2026Aug 26, 2026
AA-WER Streaming (final)2.5%2.7%3.1%4.0%
Time to final0.13s0.49s0.16sNot reported by AA source cited
Streaming price / hour$0.54 (intro)$0.20$0.18~$0.54 (token-billed estimate)
Languages60, continuous auto-detectDozens, auto-detect, mid-recording switch70+ trained, 25 verified85+, auto-detect
Speaker diarization in streamNot statedIncluded in API (streaming not confirmed)Yes, 20+ speakersNot supported in Live mode
InterfaceRealtime API (WebSocket) + Azure Speech SDKWebSocketWebSocket + file endpointLive API (WebSocket)
Open weightsNoNoNoNo

Sources: Artificial Analysis, 9to5Mac (Muse pricing), DataCamp (Grok pricing), The Batch (Gemini pricing). Verified October 2, 2026.

What it means

Developers building voice agents can now trigger actions almost instantly. The model provides accurate text before the speaker finishes, allowing systems to respond without waiting for a full sentence to complete.

Scroll to Top