Soniox TTS v2
Soniox launched TTS on 2026-04-23 (tts-rt-v1); TTS v2 (tts-rt-v2, replacing v1) was reported by audioXpress on 2026-08-10. Streaming only; regions US, EU, Japan. The v2 date is from secondary press, not a Soniox post.
- Input
- text, audio
- Output
- audio
- License
- proprietary
- Pricing
- input: $4 · output: $21.5 (USD per 1M tokens (text in / audio out); ≈ $0.70 per hour of generated speech (1 hour ≈ 30,000 audio tokens)) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Soniox API (real-time streaming, WebSocket) | tts-rt-v2 | — | docs |
Notable capabilities (2)
- 60+ languages in one model, mid-sentence switching: Single multilingual model with mixed-language text and mid-sentence language switching; Soniox claims 'hallucination-free' output (no invented or dropped words) and accurate reading of emails, phone numbers and IDs. source
- Audio tags and 20-second voice cloning (v2): TTS v2 adds expressive audio tags (whispering, laughter, hesitation, excitement), voice cloning from ~20 s of reference audio, and character-level timestamps. source
Soniox's text-to-speech, the companion to its v5 STT, aimed at multilingual voice agents.
Sources: https://soniox.com/blog/soniox-text-to-speech , https://soniox.com/pricing , https://audioxpress.com/news/soniox-tts-v2-adds-expressive-control-and-voice-cloning-to-its-multilingual-voice-ai-platform