Qwen-Audio-3.1-TTS-Next
Chinese and English only; max 3,000 input characters; output up to 240 s for podcasts, 120 s otherwise. Comparable to ByteDance Seed Audio 1.0 (Jul 2026) and StepAudio 3 Gen. Sibling TTS model Qwen-Audio-3.1-TTS (plain TTS, ~70% cheaper than 3.0) exists but its exact API id was not verified: as of 2026-09-29 the international Model Studio docs (models page, qwen-tts page) list only qwen-audio-3.0-tts-flash / -plus, and neither qwen-audio-3.1-tts-flash/-plus nor an ASR-Next id resolves on QwenCloud (404). Verified 3.1 ASR ids: qwen-audio-3.1-asr-flash(-streaming/-filetrans).
- Input
- text, audio
- Output
- audio
- License
- proprietary
- Pricing
- input: $0.848 · output: $1.696 (USD per 1M tokens (China/Beijing region price shown in docs; international price not listed)) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Alibaba Cloud Model Studio | qwen-audio-3.1-tts-next | — | docs |
Notable capabilities (1)
- One-pass speech + sound effects + ambience: 'AudioGen' model (LM + diffusion) that generates complete audio - speech, multi-speaker dialogue, podcasts, sound effects and ambient soundscapes - in a single pass from text, timestamps and up to 3 reference clips. source
Scene-level audio creation (audiobooks, podcasts, games, ads) rather than plain TTS.
Sources: https://www.alibabacloud.com/help/en/model-studio/qwen-audio-3-1-tts-next
Timeline entry
- Alibaba launches Qwen-Audio-3.1 five-model voice stack and cuts audio API prices up to 95% ★★★
Around its 2026 Apsara Conference Alibaba's Qwen team released Qwen-Audio-3.1: upgraded ASR, TTS and full-duplex Realtime models plus two new ones (ASR-Next for audio understanding, TTS-Next for one-pass speech+SFX+ambience generation), with price cuts of ~70% (TTS), ~85% (Realtime) and up to 95%…
Other Alibaba (Qwen) models
Qwen-Audio-3.1-ASR (Flash) · Qwen-Audio-3.1-Realtime (Plus) · Qwen3.8-LiveTranslate (Flash Realtime) · Qwen3.8-Omni-Flash · Qwen3.8-27B · Qwen3.8-Flash · Qwen3.8-Max · Qwen-Image-3.0 (Pro) · Qwen3.7-Plus · Qwen3-ASR (0.6B / 1.7B) + Qwen3-ForcedAligner · Qwen3-TTS (open weights 0.6B / 1.7B; API qwen3-tts-flash)