ByteDance deploys SeedRealtime, a native audio-visual full-duplex model, in the Doubao app
ByteDance Seed launched SeedRealtime, an end-to-end LLM that listens, watches (live video) and speaks at the same time instead of chaining ASR, vision and TTS, and rolled it out at scale in Doubao (Dola internationally). ByteDance says it halves conversational pacing problems compared with cascaded systems.
Key facts
- Announced 2026-08-05 by ByteDance Seed
- Single model over continuous audio, video and text streams; full-duplex with proactive interaction
- Uses visual context to resolve homophones and references to what the camera sees
- Available in Doubao/Dola and BytePlus Playground; no public API id or pricing announced
- Two weeks after Seed Audio 1.0 (2026-07-20), a one-pass speech+SFX+ambience model
What happened
ByteDance's Seed team shipped an audio-visual full-duplex model to Doubao, China's largest consumer chatbot, letting users hold natural video-call-style conversations with the assistant (demos include menu translation, museum guiding and walking through an espresso machine).
Why it matters
It puts end-to-end "see, hear and talk at once" interaction in front of a mass consumer audience. ByteDance published only human-evaluation claims, no quantitative benchmarks.
Changelog
- 2026-09-29: created
Models
- SeedRealtime (Doubao realtime audio-visual model) ByteDance · current
Related events
- StepFun releases StepAudio 3 family; its Realtime model tops Artificial Analysis full-duplex rankings ★★★
- Alibaba launches Qwen-Audio-3.1 five-model voice stack and cuts audio API prices up to 95% ★★★
- ByteDance Seed LiveInterpret 2.0: end-to-end Chinese-English simultaneous interpretation in your own voice, ~3 s behind ★★★
Sources (4)
- officialByteDance Seed - SeedRealtime released
- officialByteDance Seed - SeedRealtime page
- pressTechNode - ByteDance launches SeedRealtime
- officialByteDance Seed - Seed Audio 1.0
id: 2026-08-05-bytedance-seedrealtime · updated 2026-09-29 · open in the interactive timeline