SeedRealtime (Doubao realtime audio-visual model)
Deployed at scale in the Doubao app (Dola internationally). No public API model id, pricing or benchmark numbers published; Volcengine offers a separate Doubao end-to-end realtime dialogue API (/api/v3/realtime/dialogue) whose relation to SeedRealtime is unverified. Some press calls it the first model to watch, listen and speak simultaneously; not claimed by ByteDance, and Gemini Live / GPT-Realtime already accepted video.
- Input
- audio, video, text
- Output
- audio, text
- License
- proprietary
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Web app (Doubao / Dola) | — | dola.com/chat | — |
| BytePlus Playground | — | ai.byteplus.com/en/playground | — |
Notable capabilities (2)
- Native audio-visual full-duplex LLM: Single end-to-end model perceives continuous audio, video and text streams while listening and speaking (no ASR/VLM/TTS cascade); resolves homophones from visual context and temporal references to what it sees. source
- Proactive turn-taking: ByteDance says it halves audio-visual conversational pacing problems vs cascaded systems (fewer cut-offs, slow replies, false triggers) and can speak up proactively. source
Sources: https://seed.bytedance.com/en/SeedRealtime · https://seed.bytedance.com/en/blog/seedrealtime-audio-visual-full-duplex-llm-released-toward-omni-modal-natural-interaction · https://technode.com/2026/08/05/bytedance-launches-seedrealtime-full-duplex-audio-video-model/
Timeline entry
- ByteDance deploys SeedRealtime, a native audio-visual full-duplex model, in the Doubao app ★★★
ByteDance Seed launched SeedRealtime, an end-to-end LLM that listens, watches (live video) and speaks at the same time instead of chaining ASR, vision and TTS, and rolled it out at scale in Doubao (Dola internationally). ByteDance says it halves conversational pacing problems compared with…