Mistral releases Voxtral TTS, an open-weight 4B text-to-speech model with 3-second voice cloning
On 2026-03-23 Mistral launched Voxtral TTS, its first text-to-speech model: a 4B-parameter model with open weights (CC BY-NC 4.0) that clones a voice from ~3 seconds of audio in 9 languages and, per Mistral, beats ElevenLabs Flash v2.5 in 68.4% of human preference tests, priced at $0.016 per 1K characters via API.
Key facts
- API id voxtral-tts-2603; HF weights mistralai/Voxtral-4B-TTS-2603 (CC BY-NC 4.0, non-commercial)
- Architecture: 3.4B transformer decoder + 390M flow-matching acoustic transformer + 300M neural codec
- 9 languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, Arabic
- ~70 ms model latency, ~9.7x real-time factor, up to 2 minutes of native audio
- 68.4% win rate vs ElevenLabs Flash v2.5 in multilingual voice-cloning preference tests (Mistral)
- Price: $0.016 per 1K characters
- Followed Voxtral Transcribe 2 (2026-02-04): Voxtral Mini Transcribe V2 ($0.003/min) and open Apache-2.0 Voxtral Realtime 4B
What happened
Mistral added speech output to its Voxtral audio family. Voxtral TTS is served on the Mistral API
(/v1/audio/speech), in Le Chat and Mistral Studio, and its weights were published on Hugging Face under a
non-commercial license. Six weeks earlier Mistral had shipped Voxtral Transcribe 2, including the open Apache-2.0
Voxtral Realtime streaming ASR model (sub-200 ms latency).
Why it matters
With both open ASR and open TTS, Mistral became one of the few frontier labs offering a full open-weight voice stack, giving European and self-hosting customers an alternative to ElevenLabs and OpenAI voices. Quality comparisons are Mistral-reported.
Changelog
- 2026-09-29: created
Models
- Voxtral TTS Mistral AI · current
- Voxtral Transcribe 2 (Mini Transcribe V2 + Voxtral Realtime) Mistral AI · current
Related events
- Fish Audio open-sources S2: expressive 80+ language TTS with inline emotion tags ★★★
- BreezeBlue releases Breeze TTS 2, the new top open-weights text-to-speech model ★★★
Sources (5)
- officialMistral AI - Speaking of Voxtral
- docsMistral docs - Voxtral TTS model card
- codeHugging Face - Voxtral-4B-TTS-2603
- officialMistral AI - Voxtral Transcribe 2
- pressSiliconANGLE - Mistral releases an open-weights 'speaking' AI model
id: 2026-03-23-mistral-voxtral-tts · updated 2026-09-29 · open in the interactive timeline