Listening & Speaking with GPT-Live
OpenAI · 2026-07-08 · official · 50,206 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary This official OpenAI demonstration showcases GPT-Live-1, a full-duplex speech-to-speech model capable of simultaneous listening and speaking. OpenAI technical staff members Yuchen Zhang, Alyssa Huang, and Justin Uberti introduce the technology and demonstrate continuous, real-time multilingual translation and conversational interaction.
What is shown
- [00:00] Justin Uberti and Yuchen Zhang chat casually with GPT-Live-1 running on an iPhone.
- [00:11] Title card displays "GPT-Live-1" and "Listening & Speaking," introducing team members Yuchen Zhang, Alyssa Huang, and Justin Uberti.
- [00:41] Justin instructs the phone: "Hey Chat, I'd like you to do real-time translation for us from the language that you're hearing into English."
- [00:51] Alyssa speaks French about her favorite dish (omelettes with tomatoes and mushrooms), and the model translates concurrently into spoken English with near-zero latency.
- [01:09] Yuchen speaks Mandarin Chinese detailing his love for Cantonese dim sum (crystal shrimp dumplings, sticky rice chicken, blanched beef tripe, egg tarts), which the model translates into English in real time.
- [01:26] Justin speaks Spanish describing street tacos al pastor, which the model instantly interprets into English.
- [01:36] Yuchen prompts the model in English to summarize everyone's favorite foods; the model accurately synthesizes the foods listed across all three languages and answers a follow-up question humorously.
- [01:52] The team discusses the model's full-duplex architecture and ability to process speech every millisecond.
Claims & numbers
- Yuchen Zhang states the model "needs to think and make decision in every millisecond, understand the conversation, manage the conversation flow" to speak and listen simultaneously [00:20].
- Yuchen Zhang states that by processing in real time, the model can predict and "respond even before the user finish" to ensure natural conversational turn-taking [02:24].
Notable quotes
- [00:20] "It needs to think and make decision in every millisecond, understand the conversation, manage the conversation flow." — Yuchen Zhang
- [01:40] "Sure. Alyssa's is omelets, yours is Cantonese dim sum, and Justin is al pastor street tacos with pineapple, cilantro, spicy salsa, and lime." — GPT-Live-1
- [02:24] "If you can think in real time, then you can respond even before the user finish. That is a secret sauce for how to make it very natural." — Yuchen Zhang
Assessment This is an official OpenAI product launch demo showcasing live end-to-end full-duplex translation and conversation. While the video is cleanly produced and presented in a scripted sequence, the phone audio interface and seamless low-latency multilingual translation demonstrate genuine real-time model capabilities.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.