Resemble AI Chatterbox (Turbo / Nano / Multilingual V3)
All MIT-licensed. Multilingual (23 langs) first released Sept 2025. Multilingual V3 released 2026-06-10 (Resemble post; V3 T3 weights first pushed to HF 2026-04-22): same 0.5B Llama backbone, training data up from 25.6k to 36.7k hours, 25 languages incl. 4 dialects and 6 tuned Language Pack models, PerTh watermark on by default; Resemble reports CER under 0.20% for Italian/German but ~71-75% for Korean/Vietnamese (not production-ready); NVIDIA NIM claims 2x-39x throughput. Chatterbox-Nano HF repo created 2026-04-14 (public announcement date not found). Artificial Analysis lists Chatterbox at ~1020 Elo (secondary source). Resemble's pricing page now centres on deepfake detection; hosted TTS price not verified.
- Input
- text, audio
- Output
- audio
- License
- mit
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | ResembleAI/chatterbox | huggingface.co/ResembleAI/chatterbox | — |
| Hugging Face | ResembleAI/chatterbox-turbo | huggingface.co/ResembleAI/chatterbox-turbo | — |
| pip | chatterbox-tts | github.com/resemble-ai/chatterbox | — |
| NVIDIA NIM | resembleai/chatterbox-multilingual-tts | build.nvidia.com/resembleai/chatterbox-multilingual-tts/modelcard | — |
Notable capabilities (4)
- Emotion exaggeration control: Original 0.5B Chatterbox exposes an exaggeration/intensity knob plus CFG; zero-shot cloning from ~5 s. source
- Chatterbox-Turbo: one-step decoder, paralinguistic tags (found after launch): 350M params (Dec 2025); speech-token-to-mel decoder distilled from 10 steps to 1; native [laugh], [cough], [chuckle] tags; sub-200 ms production latency. source
- Built-in PerTh watermark: Every output carries Resemble's imperceptible Perth neural watermark that survives MP3 compression and edits. source
- Multilingual V3 and Nano (found after launch): Multilingual V3 (0.5B, 23 languages, better speaker similarity, fewer hallucinations) plus single-language fine-tune packs; Chatterbox-Nano (110M, English, ~3x real time on 8 CPU cores). source
Sources: https://www.resemble.ai/resources/chatterbox-multilingual-v3-tts-with-embedded-watermarking-for-25-languages , https://huggingface.co/ResembleAI/chatterbox-nano , https://github.com/resemble-ai/chatterbox , https://www.resemble.ai/learn/models/chatterbox-multilingual , https://huggingface.co/ResembleAI/chatterbox-turbo