Boson AI Higgs Audio v3 (Higgs TTS 3 4B / Higgs STT 3)
TTS weights non-commercial; production/hosted use needs a Boson commercial license or the Boson API (pricing not found). Also mirrored as bosonai/higgs-tts-3-4b. Predecessor Higgs Audio v2 (2025, Apache-2.0-style) on the same GitHub.
- Input
- text, audio
- Output
- audio, text
- License
- Boson Higgs TTS 3 Research and Non-Commercial License (TTS weights; Creator Use Grant for attributed monetized content)
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | bosonai/higgs-audio-v3-tts-4b | huggingface.co/bosonai/higgs-audio-v3-tts-4b | — |
| GitHub | — | github.com/boson-ai/higgs-audio | — |
| SGLang-Omni | — | — | docs |
Notable capabilities (2)
- 102-language expressive TTS with zero-shot cloning: ~4B AR decoder (24 kHz, 8 codebooks); 85 languages at production quality (WER/CER <5%), 17 usable; inline control of emotion, style, prosody, pauses and sound effects; 8K-token context; sub-second TTFA streaming. source
- Higgs STT 3 (API): Speech-to-text model (2026-03-18) for 94 languages; 1.55% WER on LibriSpeech test-clean vs 2.10% for Whisper-large-v3 (company figures). No open weights found. source