ByteDance Seed LiveInterpret 2.0: end-to-end Chinese-English simultaneous interpretation in your own voice, ~3 s behind
On 2025-07-24 ByteDance's Seed team released Seed LiveInterpret 2.0, an end-to-end speech-to-speech simultaneous interpretation model for Chinese<->English that speaks the translation in the speaker's cloned voice about 2.5-3 s behind. In ByteDance's human evaluations it came close to professional interpreters and far ahead of other systems. It shipped on Volcano Engine as "Doubao - Simultaneous Interpretation 2.0".
Key facts
- Latency: ~2.21 s first-word (speech-to-text) and ~2.53 s (speech-to-speech), which ByteDance says is 60-70% lower than cascaded systems (down from nearly 10 s)
- Accuracy: >70% in multi-speaker and >80% in single-speaker settings; human-eval score 74.8/100 (speech-to-text) vs 47.3 for the runner-up baseline; 66.3/100 speech-to-speech
- Real-time zero-shot voice cloning of each speaker; large-scale pretraining plus reinforcement learning to trade accuracy against latency
- Paper: arXiv 2507.17527 'Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice'
- Available on Volcano Engine (Ark console, 'Doubao - Simultaneous Interpretation 2.0'); planned for ByteDance's Ola Friend earbuds from end of Aug 2025. Public API model id and pricing not verified
What happened
ByteDance replaced the usual ASR -> MT -> TTS cascade with a single end-to-end model that listens, translates and speaks at the same time, and renders each speaker's translation in their own cloned voice.
Why it matters
It was one of the first product-grade end-to-end simultaneous interpreters. It came roughly ten months before OpenAI's gpt-realtime-translate (May 2026), Google's Gemini 3.5 Live Translate (June 2026) and Alibaba's Qwen3.8-LiveTranslate (Sept 2026). All accuracy figures are ByteDance's own evaluations. It covers only Chinese and English.
Changelog
- 2026-09-29: created
Related events
- Google launches Gemini 3.5 Live Translate, voice-preserving real-time speech translation in 70+ languages ★★★
- OpenAI releases GPT-Realtime-2 (reasoning voice), GPT-Realtime-Translate and GPT-Realtime-Whisper ★★★
- ByteDance deploys SeedRealtime, a native audio-visual full-duplex model, in the Doubao app ★★★
Sources (3)
- officialByteDance Seed blog - Seed LiveInterpret 2.0 released
- paperarXiv 2507.17527 - Seed LiveInterpret 2.0 technical report
- docsVolcano Engine console - simultaneous interpretation demo
id: 2025-07-24-bytedance-seed-liveinterpret-2 · updated 2026-09-29 · open in the interactive timeline