Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2025
  4. ByteDance Seed LiveInterpret 2.0: end-to-end…

ByteDance Seed LiveInterpret 2.0: end-to-end Chinese-English simultaneous interpretation in your own voice, ~3 s behind

★★★model-releaseByteDance Seedconfidence: high

On 2025-07-24 ByteDance's Seed team released Seed LiveInterpret 2.0, an end-to-end speech-to-speech simultaneous interpretation model for Chinese<->English that speaks the translation in the speaker's cloned voice about 2.5-3 s behind. In ByteDance's human evaluations it came close to professional interpreters and far ahead of other systems. It shipped on Volcano Engine as "Doubao - Simultaneous Interpretation 2.0".

Key facts

What happened

ByteDance replaced the usual ASR -> MT -> TTS cascade with a single end-to-end model that listens, translates and speaks at the same time, and renders each speaker's translation in their own cloned voice.

Why it matters

It was one of the first product-grade end-to-end simultaneous interpreters. It came roughly ten months before OpenAI's gpt-realtime-translate (May 2026), Google's Gemini 3.5 Live Translate (June 2026) and Alibaba's Qwen3.8-LiveTranslate (Sept 2026). All accuracy figures are ByteDance's own evaluations. It covers only Chinese and English.

Changelog

  • 2026-09-29: created

Related events

  1. Google launches Gemini 3.5 Live Translate, voice-preserving real-time speech translation in 70+ languages ★★★
  2. OpenAI releases GPT-Realtime-2 (reasoning voice), GPT-Realtime-Translate and GPT-Realtime-Whisper ★★★
  3. ByteDance deploys SeedRealtime, a native audio-visual full-duplex model, in the Doubao app ★★★

Sources (3)

id: 2025-07-24-bytedance-seed-liveinterpret-2 · updated 2026-09-29 · open in the interactive timeline