Breeze TTS 2
Model card lists English + Chinese; the Artificial Analysis post mentions 50 languages (possibly the hosted model) — unresolved. Needs 12 GB VRAM (24 GB recommended), CUDA/Linux. Weights are NOT commercially usable without a BreezeBlue subscription. Some secondary blogs claim it is the 'first open-weight model to beat ElevenLabs' flagship' — unverified and contradicted by the AA leaderboard (Eleven v4 far ahead).
- Input
- text, audio
- Output
- audio
- License
- BreezeBlue Research and Non-Commercial License (weights); Apache-2.0 (code)
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | BreezeBlue/Breeze-TTS-2 | huggingface.co/BreezeBlue/Breeze-TTS-2 | — |
| GitHub | — | github.com/breezeblue-ai/breeze-tts | — |
| BreezeBlue (hosted / commercial license) | — | breezeblue.ai | — |
Notable capabilities (2)
- #1 open-weights TTS on Artificial Analysis (found after launch): ~1,206-1,215 Elo in the Artificial Analysis Speech Arena, ~90 points above Fish Audio S2 Pro, #6 overall at launch — the leading open-weights TTS as of Sept 2026. source
- Clone + design + direct in one 3B checkpoint, <40 ms TTFA: Voice cloning from reference audio, voice design from text descriptions, voice direction (tone/emotion keeping identity), vocal events (laughs, coughs); streaming TTFA under 40 ms on H100 with fast path, RTF 0.32. source
Sources: https://huggingface.co/BreezeBlue/Breeze-TTS-2 , https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice/open-weights
Timeline entry
- BreezeBlue releases Breeze TTS 2, the new top open-weights text-to-speech model ★★★
On 2026-08-25 BreezeBlue published weights and inference code for Breeze TTS 2, a 3B text-to-speech model with voice cloning, voice design and voice direction and under-40 ms time-to-first-audio on an H100. It became the highest-rated open-weights model on the Artificial Analysis Speech Arena…