Qwen-Audio-3.1-ASR (Flash)
Secondary sources report 30 languages + Chinese dialects and ~160 ms latency (unverified). Sibling Qwen-Audio-3.1-ASR-Next adds speaker diarization with timestamps, emotion and sound-event detection (API id not verified). Previous: qwen-audio-3.0-asr-flash; open-weights alternative Qwen3-ASR (see qwen3-asr). Pricing not verified on an official page.
- Input
- audio
- Output
- text
- License
- proprietary
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Alibaba Cloud Model Studio (streaming) | qwen-audio-3.1-asr-flash-streaming | — | docs |
| Alibaba Cloud Model Studio / QwenCloud (file transcription) | qwen-audio-3.1-asr-flash-filetrans | — | docs |
Notable capabilities (1)
- Multilingual + dialect ASR with disfluency cleanup: Improved multilingual and Chinese-dialect recognition that automatically removes filler words and repetitions; launched with up to 95% price cut. source
Hosted speech recognition in the Qwen-Audio 3.1 stack.
Sources: https://www.alibabacloud.com/help/en/model-studio/models · https://the-decoder.com/alibaba-launches-qwen-audio-3-1-with-five-new-models-and-slashes-ai-audio-prices-by-up-to-95-percent/
Timeline entry
- Alibaba launches Qwen-Audio-3.1 five-model voice stack and cuts audio API prices up to 95% ★★★
Around its 2026 Apsara Conference Alibaba's Qwen team released Qwen-Audio-3.1: upgraded ASR, TTS and full-duplex Realtime models plus two new ones (ASR-Next for audio understanding, TTS-Next for one-pass speech+SFX+ambience generation), with price cuts of ~70% (TTS), ~85% (Realtime) and up to 95%…
Other Alibaba (Qwen) models
Qwen-Audio-3.1-Realtime (Plus) · Qwen-Audio-3.1-TTS-Next · Qwen3.8-LiveTranslate (Flash Realtime) · Qwen3.8-Omni-Flash · Qwen3.8-27B · Qwen3.8-Flash · Qwen3.8-Max · Qwen-Image-3.0 (Pro) · Qwen3.7-Plus · Qwen3-ASR (0.6B / 1.7B) + Qwen3-ForcedAligner · Qwen3-TTS (open weights 0.6B / 1.7B; API qwen3-tts-flash)