OpenAI launches GPT-Live, full-duplex voice models replacing ChatGPT's Advanced Voice Mode
On 2026-07-08 OpenAI released GPT-Live-1 and GPT-Live-1 mini, full-duplex voice models that listen and speak at the same time, backchannel ("mhmm") and hand hard questions to GPT-5.5 in the background without pausing the conversation. They replaced turn-based Advanced Voice Mode in ChatGPT (mini as default for everyone, GPT-Live-1 for paid tiers); the gpt-live-1 API went GA on 2026-09-10 at $0.05 per minute.
Key facts
- GPT-Live-1 default for ChatGPT Go/Plus/Pro; GPT-Live-1 mini default for Free users; iOS, Android and web
- Full-duplex: can be interrupted naturally, gives backchannels, stays quiet while the user thinks
- Delegates search, reasoning and agentic tasks to GPT-5.5 while the conversation continues
- OpenAI says 150M+ people use ChatGPT voice features (TechCrunch)
- ChatGPT desktop (macOS/Windows) got GPT-Live around 2026-07-23; voice plugins (email, calendar, Slack) followed 2026-09-23, together with Voice inside ChatGPT Work (press; see 2026-09-23-chatgpt-voice-plugins-work)
- API: gpt-live-1 on new v1/live/sessions endpoint, GA 2026-09-10, $0.05/min billed per second plus backend model
- Before GPT-Live, ChatGPT voice mode ran on a GPT-4o-era model: on 2026-04-10 Simon Willison noted it reported an April 2024 knowledge cutoff, so text and voice in the same subscription had different knowledge (see docs/cutoff-blindness case 017)
What happened
OpenAI replaced the voice stack in ChatGPT with a new model family built for simultaneous listening and speaking. Instead of waiting for the user to finish a turn, GPT-Live tracks the conversation continuously and offloads heavy reasoning or tool use to a text model (GPT-5.5 at launch) while it keeps talking.
Why it matters
ChatGPT's default voice experience moved to a full-duplex model with a separate "thinker" behind it, narrowing the gap between natural conversation and capable agents for one of the largest voice-assistant user bases. Rivals followed: Anthropic moved Claude's voice mode to Opus/Sonnet (2026-07-23) and Google shipped Gemini 3.8 Live (2026-09-15).
OpenAI's own post could not be fetched by our tools; details are from TechCrunch and the API docs.
Changelog
- 2026-09-29: created
- 2026-09-29: linked the 2026-09-23 Voice plugins / Voice-in-Work entry
- 2026-09-29: added pre-GPT-Live voice-mode knowledge-cutoff note (Willison)
Models
- GPT-Live 1 OpenAI · current
Videos (2)
Listening & Speaking with GPT-Live
OpenAI · 2026-07-08 · officialDescription by Gemini, which watched the video:
Summary This official OpenAI demonstration showcases GPT-Live-1, a full-duplex speech-to-speech model capable of simultaneous listening and speaking. OpenAI technical staff members Yuchen Zhang, Alyssa Huang, and Justin Uberti introduce the technology and demonstrate continuous, real-time multilingual translation and conversational interaction.
What is shown
- [00:00] Justin Uberti and Yuchen Zhang chat casually with GPT-Live-1 running on an iPhone.
- [00:11] Title card displays "GPT-Live-1" and "Listening & Speaking," introducing team members Yuchen Zhang, Alyssa Huang, and Justin Uberti.
- [00:41] Justin instructs the phone: "Hey Chat, I'd like you to do real-time translation for us from the language that you're hearing into English."
- [00:51] Alyssa speaks French about her favorite dish (omelettes with tomatoes and mushrooms), and the model translates concurrently into spoken English with near-zero latency.
- [01:09] Yuchen speaks Mandarin Chinese detailing his love for Cantonese dim sum (crystal shrimp dumplings, sticky rice chicken, blanched beef tripe, egg tarts), which the model translates into English in real time.
- [01:26] Justin speaks Spanish describing street tacos al pastor, which the model instantly interprets into English.
- [01:36] Yuchen prompts the model in English to summarize everyone's favorite foods; the model accurately synthesizes the foods listed across all three languages and answers a follow-up question humorously.
- [01:52] The team discusses the model's full-duplex architecture and ability to process speech every millisecond.
Claims & numbers
- Yuchen Zhang states the model "needs to think and make decision in every millisecond, understand the conversation, manage the conversation flow" to speak and listen simultaneously [00:20].
- Yuchen Zhang states that by processing in real time, the model can predict and "respond even before the user finish" to ensure natural conversational turn-taking [02:24].
Notable quotes
- [00:20] "It needs to think and make decision in every millisecond, understand the conversation, manage the conversation flow." — Yuchen Zhang
- [01:40] "Sure. Alyssa's is omelets, yours is Cantonese dim sum, and Justin is al pastor street tacos with pineapple, cilantro, spicy salsa, and lime." — GPT-Live-1
- [02:24] "If you can think in real time, then you can respond even before the user finish. That is a secret sauce for how to make it very natural." — Yuchen Zhang
Assessment This is an official OpenAI product launch demo showcasing live end-to-end full-duplex translation and conversation. While the video is cleanly produced and presented in a scripted sequence, the phone audio interface and seamless low-latency multilingual translation demonstrate genuine real-time model capabilities.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.
This is the new ChatGPT Voice, powered by GPT-Live
OpenAI · 2026-07-08 · officialDescription by Gemini, which watched the video:
Summary
OpenAI introduces the updated ChatGPT Voice powered by the GPT-Live 1 model, presented in a lighthearted studio setup by three senior women (SJ, Constance, and Lavelle). They demonstrate the system's full-duplex conversation capabilities, complex reasoning with real-time web search, and live spoken translation.
What is shown
- Full-Duplex Conversational Flow [00:00–00:44]: SJ interacts casually while knitting and then asks ChatGPT Voice to define "full duplex," showing natural conversational cadence where the model can speak and listen simultaneously.
- Web Browsing & Reasoning Fact-Check [01:29–02:22]: Constance asks ChatGPT Voice to fact-check audio history dates while concurrently checking live transit alerts for San Francisco's 16th Street Mission BART station and local weather; the model accurately reports no BART delays, predicts no rain in SF, and catches an incorrect date (correcting Edison's phonograph from 1865 to 1877).
- Live Speech-to-Speech Translation [02:34–03:09]: Lavelle negotiates buying a rare book in English, while ChatGPT Voice translates in real-time into French for SJ, culminating in an agreed price.
- Mobile App UI [00:06, 00:37, 01:08, 01:45, 02:43]: Displays the ChatGPT mobile interface with the pulsating visual orb representing the active voice session.
Claims & numbers
- The presenter states ChatGPT Voice is powered by GPT-Live 1, calling it "the most powerful voice model ever built" [00:23].
- The presenter claims the model supports true full-duplex interaction, allowing it to handle interruptions, pauses, spontaneous thoughts, and corrections naturally [00:38–00:56].
- Constance states the model can solve harder reasoning tasks and retrieve up-to-date web data during live voice sessions [01:21].
Notable quotes
- "Today, we are announcing the all-new ChatGPT Voice, powered by GPT-Live 1, a full-duplex conversational partner that is the most powerful voice model ever built." — SJ [00:19]
- "Imagine a normal call with a friend. You can listen and talk at the same time. That's full duplex." — ChatGPT Voice [00:38]
- "The new ChatGPT Voice listens while it speaks, is smarter than ever, and it knows when to jump in or get out of the way!" — SJ [03:11]
Assessment
This is an official OpenAI marketing launch video demonstrating live features in structured, pre-scripted vignettes. While the demonstrations showcase actual capabilities (multitasking web retrieval, fact correction, and live speech translation), the setting is tightly produced and rehearsed rather than an unscripted live test.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.
Related events
- ChatGPT Voice gets plugins and moves into ChatGPT Work: spoken requests can now drive agent tasks ★★
- OpenAI releases GPT-Realtime-2 (reasoning voice), GPT-Realtime-Translate and GPT-Realtime-Whisper ★★★
- OpenAI launches ChatGPT Work, a long-running agent for office work ★★★★
- OpenAI launches GPT-4o, a natively multimodal 'omni' model ★★★★
- Claude voice mode moves beyond Haiku to Opus and Sonnet, gains connectors and more languages ★★★
- NVIDIA releases NemotronLabs VoiceChat 11B, an open full-duplex voice model with tool calling ★★★
- Google ships Gemini 3.8 Live voice models and Gemini 3.8 Flash TTS with voice design and cloning ★★
Sources (7)
- officialOpenAI - Introducing GPT-Live
- pressTechCrunch - OpenAI releases new voice models for more natural live conversations
- docsgpt-live-1 model page
- docsOpenAI API changelog (GPT-Live 1 GA, 2026-09-10)
- discussionSimon Willison on X - ChatGPT voice mode reports an April 2024 cutoff
- discussionSimon Willison - ChatGPT voice mode is a weaker model (2026-04-10)
- pressPondero - GPT-Live comes to ChatGPT desktop
id: 2026-07-08-openai-gpt-live-chatgpt-voice · updated 2026-09-29 · open in the interactive timeline