Post-Cutoff.com
  1. Home
  2. Models
  3. Voxtral Transcribe 2 (Mini Transcribe V2 + Voxtral Realtime)

Voxtral Transcribe 2 (Mini Transcribe V2 + Voxtral Realtime)

Mistral AIcurrentaudio/speechVoxtralopen weights

13 languages (en, zh, hi, es, ar, fr, pt, ru, de, ja, ko, it, nl). Batch model is API-only ('Premier' license); Realtime has open weights. Replaced voxtral-mini-2507 / Voxtral Mini Transcribe (deprecated 2026-02-27, retired 2026-05-31). Tech report arXiv 2602.11298. Accuracy claims are Mistral's.

Input
audio
Output
text
License
apache-2.0
Pricing
per minute batch: $0.003 · per minute realtime: $0.006 (USD per minute of audio (voxtral-mini-2602 batch / voxtral-mini-transcribe-realtime-2602)) source
Verified
2026-09-29

How to call it

ProviderModel idEndpoint / URLDocs
Mistral API (batch)voxtral-mini-2602https://api.mistral.ai/v1/audio/transcriptionsdocs
Mistral API (realtime)voxtral-mini-transcribe-realtime-2602—docs
Hugging Face (Realtime, open weights)mistralai/Voxtral-Mini-4B-Realtime-2602huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602—
Web app—chat.mistral.ai—

Notable capabilities (2)

curl https://api.mistral.ai/v1/audio/transcriptions -H "Authorization: Bearer $MISTRAL_API_KEY" \
  -F model=voxtral-mini-2602 -F file=@audio.mp3

Sources: https://mistral.ai/news/voxtral-transcribe-2 · https://docs.mistral.ai/models/overview

Timeline entry

  1. Mistral releases Voxtral TTS, an open-weight 4B text-to-speech model with 3-second voice cloning ★★★

    On 2026-03-23 Mistral launched Voxtral TTS, its first text-to-speech model: a 4B-parameter model with open weights (CC BY-NC 4.0) that clones a voice from ~3 seconds of audio in 9 languages and, per Mistral, beats ElevenLabs Flash v2.5 in 68.4% of human preference tests, priced at $0.016 per 1K…

Other Mistral AI models

Mistral OCR 4.1 · Mistral Medium 3.5 · Voxtral TTS · Mistral Small 4 · Mistral Large 3 · Codestral 25.08 · Voxtral Small · Robostral Navigate