StepAudio 3 ASR Max / StepAudio 3 TTS
Family file for the non-realtime StepAudio 3 models. Languages: zh, en, ja, ko, fr, es (non-zh/en in preview). stepaudio-3-gen-preview (speech+SFX+ambience+BGM) and stepaudio-3-music-preview are free during preview. Previous gen: stepaudio-2.5-asr ($0.022/h), stepaudio-2.5-asr-stream ($0.18/h), stepaudio-2.5-tts ($0.85/10k chars). Exact per-model release date assumed = family launch 2026-09-15.
- Input
- audio, text
- Output
- text, audio
- License
- proprietary
- Pricing
- per hour asr: $0.4 · per 10k characters tts: $0.36 (USD: stepaudio-3-asr-max $0.40/hour of audio; stepaudio-3-tts $0.36 per 10,000 characters; voice cloning $1.50/voice) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| StepFun API | stepaudio-3-asr-max | — | docs |
| StepFun API | stepaudio-3-tts | — | docs |
| StepFun API (preview) | stepaudio-3-gen-preview | — | docs |
Notable capabilities (2)
- #1 non-streaming ASR on AA-WER: Artificial Analysis ranked StepAudio 3 ASR #1 on its AA-WER Index for non-streaming speech-to-text with 1.7% WER (StepAudio 2.5 ASR: 4.7%). source
- Context-aware streaming TTS: Natural, context-aware speech with low-latency streaming, natural-language control and voice cloning; 1,000-char input limit; wav/mp3/flac/opus/pcm. source
Sources: https://platform.stepfun.ai/docs/en/guides/models/audio · https://platform.stepfun.ai/docs/en/pricing/details
Timeline entry
- StepFun releases StepAudio 3 family; its Realtime model tops Artificial Analysis full-duplex rankings ★★★
Chinese lab StepFun launched StepAudio 3, five audio models (Realtime, ASR Max, TTS, Gen, Music). StepAudio 3 Realtime, a "think-while-speaking" full-duplex voice model, ranked #1 on Artificial Analysis for Conversational Dynamics (98.9%) and Speech Reasoning (99.7%), and StepAudio 3 ASR ranked #1…
Other StepFun models
StepFun Step-Audio-EditX · Step 5 Preview · StepAudio 3 Realtime