Alibaba's Qwen-Audio-3.0-TTS takes #1 on the Artificial Analysis text-to-speech leaderboard
On 2026-07-20 Alibaba's Tongyi Lab released Qwen-Audio-3.0-TTS in Flash (real-time) and Plus (quality) tiers. It supports 16 languages and 20 Chinese dialect regions. The Plus tier ranked first on the independent Artificial Analysis TTS leaderboard while costing about $27.6 per 1M characters, roughly a quarter of Eleven v3's price. It was the first Chinese hosted TTS to top that arena.
Key facts
- API ids: qwen-audio-3.0-tts-flash, qwen-audio-3.0-tts-plus (Alibaba Cloud Model Studio)
- Artificial Analysis TTS arena: Plus #1 at Elo ~1,236-1,237 vs Speechify Simba 3.2 ~1,234 (press, July 2026)
- Technical report arXiv 2607.23938 (submitted 2026-07-27): 12.5 Hz tokenizer, five-stage LM + flow-matching training, SOTA claims on SEED-TTS-Eval and CV3-Eval
- 16 languages, 20 Chinese dialect regions, up to 3 minutes of one-pass long-form output, natural-language and inline-tag control, voice cloning and Voice Design
- Plus: $27.59 per 1M characters vs Eleven v3 $100 (press)
- The first-place ranking did not last: Inworld TTS-2, Cartesia Sonic 3.6 and Eleven v4 (2026-09-28) led later
What happened
Alibaba released a new generation of hosted TTS models built on a low-frame-rate tokenizer and a multi-stage training recipe, with strong control features (instructions, inline tags, dialects, long-form output). Its Plus tier topped the Artificial Analysis blind-listening arena at launch.
Why it matters
A Chinese lab led the main independent TTS leaderboard at a fraction of ElevenLabs' price, which started the summer-2026 TTS price and quality race. Alibaba followed two months later with Qwen-Audio-3.1 and price cuts of about 70%.
Changelog
- 2026-09-29: created
Models
- Qwen-Audio-3.0-TTS (Flash / Plus) Alibaba (Qwen / Tongyi Lab) · current
Related events
- Alibaba launches Qwen-Audio-3.1 five-model voice stack and cuts audio API prices up to 95% ★★★
- ElevenLabs launches Eleven v4 and Eleven v4 Turbo, #1 on Artificial Analysis TTS arena ★★★★
- Inworld Realtime TTS-2 reaches GA with audio-aware, prompt-directed speech ★★★
- Cartesia Sonic-3.6 goes GA and tops the Artificial Analysis Speech Arena ★★★
- Alibaba open-sources Qwen3-TTS (voice design, 3-second cloning, 97 ms streaming) and, a week later, Qwen3-ASR ★★★
Sources (4)
- paperarXiv 2607.23938 - Qwen-Audio-3.0-TTS technical report
- docsModel Studio - non-real-time speech synthesis (qwen-audio-3.0-tts-flash)
- pressMarkTechPost - Qwen-Audio-3.0-TTS in Flash and Plus tiers across 16 languages
- docsArtificial Analysis - text-to-speech leaderboard
id: 2026-07-20-qwen-audio-3-0-tts · updated 2026-09-29 · open in the interactive timeline