Post-Cutoff.com
  1. Home
  2. Models
  3. Qwen-Audio-3.1-TTS-Next

Qwen-Audio-3.1-TTS-Next

Alibaba (Qwen)currentaudio/speechQwen-Audio 3.1

Chinese and English only; max 3,000 input characters; output up to 240 s for podcasts, 120 s otherwise. Comparable to ByteDance Seed Audio 1.0 (Jul 2026) and StepAudio 3 Gen. Sibling TTS model Qwen-Audio-3.1-TTS (plain TTS, ~70% cheaper than 3.0) exists but its exact API id was not verified: as of 2026-09-29 the international Model Studio docs (models page, qwen-tts page) list only qwen-audio-3.0-tts-flash / -plus, and neither qwen-audio-3.1-tts-flash/-plus nor an ASR-Next id resolves on QwenCloud (404). Verified 3.1 ASR ids: qwen-audio-3.1-asr-flash(-streaming/-filetrans).

Input
text, audio
Output
audio
License
proprietary
Pricing
input: $0.848 · output: $1.696 (USD per 1M tokens (China/Beijing region price shown in docs; international price not listed)) source
Verified
2026-09-29

How to call it

ProviderModel idEndpoint / URLDocs
Alibaba Cloud Model Studioqwen-audio-3.1-tts-next—docs

Notable capabilities (1)

Scene-level audio creation (audiobooks, podcasts, games, ads) rather than plain TTS.

Sources: https://www.alibabacloud.com/help/en/model-studio/qwen-audio-3-1-tts-next

Timeline entry

  1. Alibaba launches Qwen-Audio-3.1 five-model voice stack and cuts audio API prices up to 95% ★★★

    Around its 2026 Apsara Conference Alibaba's Qwen team released Qwen-Audio-3.1: upgraded ASR, TTS and full-duplex Realtime models plus two new ones (ASR-Next for audio understanding, TTS-Next for one-pass speech+SFX+ambience generation), with price cuts of ~70% (TTS), ~85% (Realtime) and up to 95%…

Other Alibaba (Qwen) models

Qwen-Audio-3.1-ASR (Flash) · Qwen-Audio-3.1-Realtime (Plus) · Qwen3.8-LiveTranslate (Flash Realtime) · Qwen3.8-Omni-Flash · Qwen3.8-27B · Qwen3.8-Flash · Qwen3.8-Max · Qwen-Image-3.0 (Pro) · Qwen3.7-Plus · Qwen3-ASR (0.6B / 1.7B) + Qwen3-ForcedAligner · Qwen3-TTS (open weights 0.6B / 1.7B; API qwen3-tts-flash)