Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. StepFun releases StepAudio 3 family; its Realtime model…

StepFun releases StepAudio 3 family; its Realtime model tops Artificial Analysis full-duplex rankings

★★★after cutoffmodel-releaseStepFunconfidence: high

Chinese lab StepFun launched StepAudio 3, five audio models (Realtime, ASR Max, TTS, Gen, Music). StepAudio 3 Realtime, a "think-while-speaking" full-duplex voice model, ranked #1 on Artificial Analysis for Conversational Dynamics (98.9%) and Speech Reasoning (99.7%), and StepAudio 3 ASR ranked #1 on AA-WER (1.7%).

Key facts

What happened

StepFun released a full audio stack at once and made the Realtime, Gen and Music models free during a preview period. The Realtime model's technical report describes a listen-converse-think-act loop with "Deep Perception", "Seamless Duplex" and "Think-While-Speaking" components.

Why it matters

A Chinese startup's voice model led a major independent leaderboard on conversational dynamics ahead of Western frontier-lab voice models (GPT-Live-1 per StepFun's comparison), showing how fast full-duplex voice is commoditizing.

Leaderboard positions are as of launch and come from StepFun's and Artificial Analysis's X posts.

Changelog

  • 2026-09-29: created
  • 2026-09-29: added StepAudio 2.5 Realtime project page and its self-reported scores

Models

Related events

  1. Alibaba launches Qwen-Audio-3.1 five-model voice stack and cuts audio API prices up to 95% ★★★
  2. ByteDance deploys SeedRealtime, a native audio-visual full-duplex model, in the Doubao app ★★★

Sources (7)

id: 2026-09-15-stepfun-stepaudio-3 · updated 2026-09-29 · open in the interactive timeline