StepFun releases StepAudio 3 family; its Realtime model tops Artificial Analysis full-duplex rankings
Chinese lab StepFun launched StepAudio 3, five audio models (Realtime, ASR Max, TTS, Gen, Music). StepAudio 3 Realtime, a "think-while-speaking" full-duplex voice model, ranked #1 on Artificial Analysis for Conversational Dynamics (98.9%) and Speech Reasoning (99.7%), and StepAudio 3 ASR ranked #1 on AA-WER (1.7%).
Key facts
- API ids: stepaudio-3-realtime-preview, stepaudio-3-chat-preview, stepaudio-3-asr-max, stepaudio-3-tts, stepaudio-3-gen-preview, stepaudio-3-music-preview
- Realtime/Gen/Music free during preview; ASR Max $0.40/hour; TTS $0.36 per 10k characters
- Realtime runs private reasoning in parallel with speech (Think-While-Speaking); 98.9 on Artificial Analysis Full-Duplex Bench
- StepAudio 3 ASR 1.7% WER on AA-WER (StepAudio 2.5 ASR: 4.7%) per Artificial Analysis
- Follows StepAudio 2.5 Realtime (2026-05-26): persona/role-play realtime model (zh/en) with million-scale persona augmentation and role-play RLHF; project page reports 80.41 human eval, 86.36 general dialogue, 79.80 spoken QA, 82.18 paralinguistics, first on all five of StepFun's own dimensions
What happened
StepFun released a full audio stack at once and made the Realtime, Gen and Music models free during a preview period. The Realtime model's technical report describes a listen-converse-think-act loop with "Deep Perception", "Seamless Duplex" and "Think-While-Speaking" components.
Why it matters
A Chinese startup's voice model led a major independent leaderboard on conversational dynamics ahead of Western frontier-lab voice models (GPT-Live-1 per StepFun's comparison), showing how fast full-duplex voice is commoditizing.
Leaderboard positions are as of launch and come from StepFun's and Artificial Analysis's X posts.
Changelog
- 2026-09-29: created
- 2026-09-29: added StepAudio 2.5 Realtime project page and its self-reported scores
Models
- StepAudio 3 ASR Max / StepAudio 3 TTS StepFun · current
- StepAudio 3 Realtime StepFun · preview
Related events
- Alibaba launches Qwen-Audio-3.1 five-model voice stack and cuts audio API prices up to 95% ★★★
- ByteDance deploys SeedRealtime, a native audio-visual full-duplex model, in the Doubao app ★★★
Sources (7)
- officialStepFun on X - Introducing StepAudio 3
- docsStepFun audio models docs
- docsStepFun pricing
- paperStepAudio 3 Realtime Technical Report
- discussionArtificial Analysis on X - StepAudio 3 ASR #1 on AA-WER
- officialStepAudio 2.5 Realtime project page
- pressDecrypt - StepFun's voice AI topped every benchmark (StepAudio 2.5)
id: 2026-09-15-stepfun-stepaudio-3 · updated 2026-09-29 · open in the interactive timeline