Microsoft AI launches MAI-Transcribe-2-Streaming (real-time ASR) and MAI-Voice-2.1 / 2.1-Flash
On Oct 1, 2026 Microsoft AI released MAI-Transcribe-2-Streaming, its first streaming speech-recognition model (first hypotheses in just over 100 ms, 60 languages, $0.54/hour through year-end), and two text-to-speech models, MAI-Voice-2.1 ($22 per 1M characters) and MAI-Voice-2.1-Flash (150 ms end-to-end, $15 per 1M characters). It extends Microsoft's in-house voice stack beyond OpenAI models.
Key facts
- MAI-Transcribe-2-Streaming: first hypotheses in just over 100 ms; Microsoft claims words appear 2x faster than the closest competitor and #1 accuracy for final and partial transcripts on Artificial Analysis
- MAI-Transcribe-2-Streaming: 60 languages with automatic language detection; $0.54 per hour of audio through end of 2026
- MAI-Voice-2.1: 23 languages / 26 locales, one voice keeps its native accent across languages, voice cloning with consent guardrails; $22 per 1M characters
- MAI-Voice-2.1-Flash: 150 ms end-to-end latency, 55% faster inference; $15 per 1M characters
- Availability: Microsoft Foundry, MAI Playground, OpenRouter, Vercel, Azure Voice Live; LiveKit coming soon
- Artificial Analysis (Oct 1, ~368k views): #1 of 38 models on AA-WER Streaming, final-transcript WER 2.5% at 0.13 s after end of speech (previous #1 Grok Voice Transcribe 2.0: 2.7% at 0.49 s); first partial transcript 2.5% at 0.12 s. AA puts its $0.54/hour ($9 per 1,000 min) at the high end, above ElevenLabs Scribe v2 Realtime and Deepgram Flux ($6.50)
- Mustafa Suleyman (Oct 1, ~311k views): 'most accurate real time transcription model in the world', '55% faster and 60% cheaper than ElevenLabs' (the price claim conflicts with AA's per-hour comparison; the basis of the 60% figure is not stated)
What happened
Microsoft AI (Mustafa Suleyman's group) shipped a low-latency streaming version of MAI-Transcribe-2 plus an updated voice-generation pair. All three were available on launch day. Benchmark claims are Microsoft's own, citing Artificial Analysis.
Why it matters
Real-time transcription and TTS are the building blocks of voice agents. Microsoft now offers its own models for both at low prices, competing with OpenAI, ElevenLabs and Deepgram.
Changelog
- 2026-10-02: created
- 2026-10-02: added Artificial Analysis results and Suleyman's post (sweep 2026-10-02)
Models
- MAI-Transcribe-2-Streaming Microsoft · current
- MAI-Voice-2.1 / MAI-Voice-2.1-Flash Microsoft · current
People
Related posts (3)
- Microsoft AI original ↗ Microsoft AI @microsoftai · x · 2026-10-01
Cited as a source by: 2026-10-01-microsoft-mai-transcribe-2-streaming-voice-2-1 - Artificial Analysis: MAI-Transcribe-2-Streaming takes #1 on AA-WER Streaming original ↗ Artificial Analysis @ArtificialAnlys · x · 2026-10-01
Independent benchmark (~368k views): 2.5% WER at 0.13 s, #1 of 38 streaming STT models, with price comparison. - Suleyman: 'the most accurate real time transcription model in the world' original ↗ Mustafa Suleyman @mustafasuleyman · x · 2026-10-01
Microsoft AI CEO's launch post (~311k views) claiming #1 accuracy and '55% faster and 60% cheaper than ElevenLabs'.
Related events
- Microsoft launches MAI-Transcribe-2, claiming the most accurate and cheapest speech recognition at $0.10/hour ★★★
- Microsoft launches seven in-house MAI models at Build 2026, led by MAI-Thinking-1 ★★★★
- Microsoft makes its own Azure Realtime speech-to-speech model generally available in the Voice Live API ★★
Sources (4)
- discussionArtificial Analysis on X: MAI-Transcribe-2-Streaming #1 on AA-WER Streaming
- officialMustafa Suleyman on X: most accurate real-time transcription model
- officialMicrosoft AI: Our first streaming transcription model
- official@MicrosoftAI on X: all three models available today
id: 2026-10-01-microsoft-mai-transcribe-2-streaming-voice-2-1 · updated 2026-10-02 · open in the interactive timeline