Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. Alibaba open-sources Qwen3-TTS (voice design, 3-second…

Alibaba open-sources Qwen3-TTS (voice design, 3-second cloning, 97 ms streaming) and, a week later, Qwen3-ASR

★★★open-sourceAlibabaQwenconfidence: high

On 2026-01-22 Alibaba's Qwen team released Qwen3-TTS under Apache-2.0 (0.6B and 1.7B checkpoints plus a 12 Hz tokenizer). It offers voice design from text descriptions, voice cloning from about 3 s of audio in 10 languages, and ~97 ms streaming latency. On 2026-01-29 Qwen3-ASR followed (0.6B/1.7B plus a forced aligner, 30 languages and 22 Chinese dialects). Both became among the most-downloaded open speech models of 2026.

Key facts

What happened

Qwen released a complete open TTS family with voice design, cloning and low-latency streaming, then an open ASR family with a forced aligner for timestamps a week later, both under Apache-2.0.

Why it matters

Voice cloning and voice design had mostly been proprietary (ElevenLabs and others). Qwen3-TTS made them freely self-hostable, and it became one of the most-downloaded speech models of 2026.

Changelog

  • 2026-09-29: created

Models

Related events

  1. Alibaba's Qwen-Audio-3.0-TTS takes #1 on the Artificial Analysis text-to-speech leaderboard ★★★
  2. Alibaba launches Qwen-Audio-3.1 five-model voice stack and cuts audio API prices up to 95% ★★★
  3. Fish Audio open-sources S2: expressive 80+ language TTS with inline emotion tags ★★★

Sources (5)

id: 2026-01-22-qwen3-tts-asr-open-weights · updated 2026-09-29 · open in the interactive timeline