Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. Mistral releases Voxtral TTS, an open-weight 4B…

Mistral releases Voxtral TTS, an open-weight 4B text-to-speech model with 3-second voice cloning

★★★open-sourceMistral AIconfidence: high

On 2026-03-23 Mistral launched Voxtral TTS, its first text-to-speech model: a 4B-parameter model with open weights (CC BY-NC 4.0) that clones a voice from ~3 seconds of audio in 9 languages and, per Mistral, beats ElevenLabs Flash v2.5 in 68.4% of human preference tests, priced at $0.016 per 1K characters via API.

Key facts

What happened

Mistral added speech output to its Voxtral audio family. Voxtral TTS is served on the Mistral API (/v1/audio/speech), in Le Chat and Mistral Studio, and its weights were published on Hugging Face under a non-commercial license. Six weeks earlier Mistral had shipped Voxtral Transcribe 2, including the open Apache-2.0 Voxtral Realtime streaming ASR model (sub-200 ms latency).

Why it matters

With both open ASR and open TTS, Mistral became one of the few frontier labs offering a full open-weight voice stack, giving European and self-hosting customers an alternative to ElevenLabs and OpenAI voices. Quality comparisons are Mistral-reported.

Changelog

  • 2026-09-29: created

Models

Related events

  1. Fish Audio open-sources S2: expressive 80+ language TTS with inline emotion tags ★★★
  2. BreezeBlue releases Breeze TTS 2, the new top open-weights text-to-speech model ★★★

Sources (5)

id: 2026-03-23-mistral-voxtral-tts · updated 2026-09-29 · open in the interactive timeline