MAI-Voice-2 / MAI-Voice-2-Flash
Launched at Build 2026-06-02 (MAI-Voice-2); Flash followed 2026-07-23 (date per secondary sources). Both public preview in Azure Speech. Languages include en-US/AU, de, fr, es-ES/MX, pt-BR/PT, it, ko, zh-CN, tr, ru, th, nl, ro, hu, hi. Also used in Copilot (Audio Expressions). Predecessor MAI-Voice-1 no longer listed on the MAI-Voice docs page. Also on Fireworks and Baseten (ids not verified).
- Input
- text
- Output
- audio
- License
- proprietary
- Pricing
- per million characters: $22 · per million characters flash: $15 (USD per 1M characters (MAI-Voice-2 / MAI-Voice-2-Flash, 'starting at')) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Azure Speech in Microsoft Foundry (SSML voice name) | en-US-Harper:MAI-Voice-2 | https://{region}.tts.speech.microsoft.com/cognitiveservices/v1 | docs |
| Azure Speech in Microsoft Foundry (SSML voice name) | en-US-Harper:MAI-Voice-2-Flash | — | docs |
| Azure Voice Live (TTS output) | MAI-Voice-2-Flash | — | docs |
| OpenRouter | microsoft/mai-voice-2 | https://openrouter.ai/api/v1/audio/speech | — |
| OpenRouter | microsoft/mai-voice-2-flash | — | — |
| Web app (MAI Playground) | — | playground.microsoft.ai/ | — |
Notable capabilities (3)
- Gated instant voice cloning: Matches a consented reference voice from a 5-60 s clip without training; only approved (Limited Access) licensed voices can be synthesized. source
- SSML emotion/style control: mstts:express-as styles (angry, fearful, joyful, whispering, shouting, etc.) with styledegree, across 15 languages / 18 locales. source
- Low-latency Flash tier (found after launch): MAI-Voice-2-Flash (public preview from 2026-07-23) targets voice agents/IVR; Microsoft quotes ~225 ms latency vs ~1 s for MAI-Voice-2 (for a 45 s clip). source
Microsoft's in-house expressive TTS, used through standard Azure Speech SSML.
curl -X POST "https://$REGION.tts.speech.microsoft.com/cognitiveservices/v1" \
-H "Ocp-Apim-Subscription-Key: $SPEECH_KEY" -H "Content-Type: application/ssml+xml" \
-H "X-Microsoft-OutputFormat: audio-24khz-160kbitrate-mono-mp3" \
--data '<speak version="1.0" xml:lang="en-US"><voice name="en-US-Harper:MAI-Voice-2">Hello from MAI Voice.</voice></speak>' -o out.mp3
Sources: https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-voices · https://microsoft.ai/models/mai-voice-2/ · https://microsoft.ai/news/building-a-hillclimbing-machine-launching-seven-new-mai-models/
Timeline entry
- Microsoft launches seven in-house MAI models at Build 2026, led by MAI-Thinking-1 ★★★★
At Build on 2026-06-02 Microsoft AI (led by Mustafa Suleyman) launched seven first-party MAI models, including its first flagship reasoning model MAI-Thinking-1, the MAI-Code-1-Flash coding model in GitHub Copilot and VS Code, MAI-Image-2.5, MAI-Transcribe-1.5 and MAI-Voice-2 - Microsoft's…
Other Microsoft models
Phi-4-Reasoning-Vision-15B · VibeVoice (ASR, ASR-Streaming, ASR-BitNet, Realtime-0.5B TTS) · MAI-Transcribe-2 · Phi-4 (14B)