Alibaba launches Qwen-Audio-3.1 five-model voice stack and cuts audio API prices up to 95%
Around its 2026 Apsara Conference Alibaba's Qwen team released Qwen-Audio-3.1: upgraded ASR, TTS and full-duplex Realtime models plus two new ones (ASR-Next for audio understanding, TTS-Next for one-pass speech+SFX+ambience generation), with price cuts of ~70% (TTS), ~85% (Realtime) and up to 95% (ASR). Qwen3.8-LiveTranslate (60 input languages, 29 with voice output) debuted alongside.
Key facts
- Five models: Qwen-Audio-3.1-ASR, -ASR-Next, -TTS, -TTS-Next, -Realtime
- qwen-audio-3.1-realtime-plus: 262K context; $6.40 audio in / $24 audio out per 1M tokens on QwenCloud
- Realtime task success 82.0% (from 78.4%); response rate to background speech cut from 73.0% to 13.0% (arXiv 2609.25176)
- qwen-audio-3.1-tts-next (model docs dated 2026-09-22): zh/en, up to 3,000 chars, up to 240 s podcast output
- Qwen3.8-LiveTranslate (announced 2026-09-19, id qwen3.8-livetranslate-flash-realtime): LAAL latency cut from 2.8 s to 2.3 s; 60 input / 29 voice-output languages; $7.50 audio in / $30 audio out per 1M tokens; API-only
What happened
Alibaba's Qwen team shipped a complete hosted audio stack in one release: recognition (ASR, ASR-Next with diarization, emotion and sound-event detection), synthesis (TTS with cross-language voice transfer, TTS-Next that mixes speech, sound effects and ambience in one pass) and a full-duplex Realtime model with tool use and web search. It came two months after Qwen-Audio-3.0 (July 2026, see 2026-07-20-qwen-audio-3-0-tts) and alongside Qwen3.8-LiveTranslate at Apsara 2026.
Why it matters
Chinese labs (Alibaba, StepFun, ByteDance) now field voice-agent models that top or approach GPT-Live / Gemini Live on public leaderboards at a fraction of the price, turning real-time voice into a price war.
The exact API ids of the 3.1 TTS and ASR-Next models (not in the international Model Studio docs as of 2026-09-29; only qwen-audio-3.0-tts-flash/-plus and qwen-audio-3.1-asr-flash-streaming/-filetrans are listed), and Model Studio international prices were not verified.
Changelog
- 2026-09-29: created
- 2026-09-29: added verified Qwen3.8-LiveTranslate id, date, pricing and model file qwen3-8-livetranslate
- 2026-09-29: linked the Qwen-Audio-3.0-TTS and Apsara 2026 entries; recorded which 3.1 API ids are published
Models
- Qwen-Audio-3.1-ASR (Flash) Alibaba (Qwen) · current
- Qwen-Audio-3.1-Realtime (Plus) Alibaba (Qwen) · current
- Qwen-Audio-3.1-TTS-Next Alibaba (Qwen) · current
- Qwen3.8-LiveTranslate (Flash Realtime) Alibaba (Qwen) · current
Related events
- Alibaba's Qwen-Audio-3.0-TTS takes #1 on the Artificial Analysis text-to-speech leaderboard ★★★
- Apsara 2026: Alibaba says Qwen 4 is in training, targets 5-10T-parameter Qwen 4.5/5, reports self-improvement runs and unveils Zhenwu V900 chip ★★★
- StepFun releases StepAudio 3 family; its Realtime model tops Artificial Analysis full-duplex rankings ★★★
- Google ships Gemini 3.8 Live voice models and Gemini 3.8 Flash TTS with voice design and cloning ★★
- Alibaba open-sources Qwen3-TTS (voice design, 3-second cloning, 97 ms streaming) and, a week later, Qwen3-ASR ★★★
- ByteDance deploys SeedRealtime, a native audio-visual full-duplex model, in the Doubao app ★★★
Sources (8)
- officialQwen on X - Meet Qwen-Audio-3.1
- docsQwenCloud - qwen-audio-3.1-realtime-plus
- docsModel Studio - qwen-audio-3.1-tts-next
- paperQwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction
- officialQwen on X - Meet Qwen3.8-LiveTranslate (2026-09-19)
- docsQwenCloud - qwen3.8-livetranslate-flash-realtime
- pressThe Decoder - Qwen Audio 3.1 slashes prices up to 95%
- pressMarkTechPost - Qwen-Audio-3.1-Realtime
id: 2026-09-23-qwen-audio-3-1 · updated 2026-09-29 · open in the interactive timeline