Models & Assistants

GPT-Live-1 Is Here: OpenAI's Full-Duplex Voice Model Replaces Advanced Voice Mode Globally

OpenAI's GPT-Live-1 and GPT-Live-1 mini launch July 8, replacing Advanced Voice Mode with simultaneous listen-and-speak AI powered by GPT-5.5 in the background.

OpenAI GPT-Live voice model interface showing full-duplex simultaneous listen-and-speak capability rolling out in ChatGPT on iOS, Android, and web

Update, 13 July 2026: OpenAI passes GPT-5.6 Sol inference savings to subscribers — ~10% more effective usage, 500k accounts reset

GPT-Live-1 launched alongside significant backend momentum that has continued into mid-July. On July 13, OpenAI’s Thibault Sottiaux (Tibo) confirmed via X that inference optimizations for GPT-5.6 Sol are being passed directly to all subscription tiers, delivering roughly 10% more effective usage at no added cost. His framing was explicit: “No nerfing, only good stuff.”

Separately, OpenAI deposited a banked reset into approximately 500,000 ChatGPT Work and Codex accounts as a milestone reward, and fixed a bug that had prevented resets from applying correctly for a subset of users. The banked reset feature, which lets subscribers save a usage-limit reset and activate it on demand, has also expanded from desktop only to web and mobile.

OpenAI also temporarily lifted the five-hour usage-limit cap for Plus, Business, and Pro subscribers, with no confirmed return date.

These changes follow a traffic surge roughly double previous peak levels triggered by the GPT-5.6 Sol general availability rollout on July 9. For agentic Codex workloads where token burn accumulates quickly, the 10% efficiency gain meaningfully reduces cost pressure. The original post below covers the GPT-Live-1 launch details that remain current.

Update, 9 July 2026: GPT-5.6 (Sol, Terra, Luna) goes GA across ChatGPT, Codex, and the API

One day after GPT-Live-1 launched, OpenAI moved its next model family to general availability. GPT-5.6 is now live globally across ChatGPT, Codex, and the OpenAI API, and it introduces a three-tier naming system that will carry forward through future generations: Sol (frontier reasoning and agentic work), Terra (GPT-5.5-level performance at half the price), and Luna (fastest and cheapest). All three share a 1.05M-token context window, 128K max output, and a February 16, 2026 knowledge cutoff.

This is directly relevant to the GPT-Live-1 story: GPT-5.5, which powers GPT-Live-1’s backend reasoning, now has a stronger successor in Sol, available to all paying users and self-serve API developers. OpenAI has confirmed GPT-5.4 retires July 23; GPT-5.5 remains available.

New capabilities shipping at GA include ultra multi-agent mode (four parallel agents, available in ChatGPT Work and Codex), Programmatic Tool Calling in the Responses API (GPT-5.6 writes and runs JavaScript in an isolated V8 runtime to coordinate tools, returning smaller structured results), and a max reasoning effort setting beyond xhigh. API pricing runs $5/$30 per million tokens for Sol, $2.50/$15 for Terra, and $1/$6 for Luna.

Full details are in OpenAI’s GA announcement and the Help Center access guide.

OpenAI has replaced Advanced Voice Mode with two new voice models, GPT-Live-1 and GPT-Live-1 mini, rolling out to ChatGPT users globally from July 8, 2026. The short version: ChatGPT can now listen and speak at the same time, delegates genuinely hard questions to GPT-5.5 running quietly in the background, and is available to every ChatGPT tier from day one.

This is the third generation of ChatGPT’s voice technology in two years, and the architectural change is substantial enough that it warrants a proper look.

What Actually Changed Under the Hood

The original 2023 ChatGPT Voice worked in three sequential steps: Whisper transcribed your speech, GPT-4 generated a text response, and a text-to-speech model read it back. Each handoff introduced delay and lost nuance. Advanced Voice Mode, which arrived for paid users in September 2024, collapsed those steps into a single model but kept the rigid turn-by-turn structure. A brief pause or a bit of background noise could trigger an interruption mid-sentence.

GPT-Live scraps turn detection entirely. The model processes audio continuously and makes decisions many times per second: speak, keep listening, pause, say “mhmm”, or quietly hand the question off to another model. There is no silence threshold to trip over.

The second structural change is delegation. When a question needs web search, reasoning, or multi-step work, GPT-Live passes it to GPT-5.5 in the background and keeps the conversation moving while it waits. You do not sit in silence while it thinks. OpenAI has said the background model will be updated as newer frontier models are released, so this is not a fixed ceiling.

The Numbers Behind the Claims

OpenAI published head-to-head benchmarks comparing GPT-Live-1 against Advanced Voice Mode. The most striking figures are on GPQA (expert-level scientific reasoning): Advanced Voice Mode scored 45.3%, while GPT-Live-1 on High reasoning reaches 84.2%. On BrowseComp, which tests agentic web search, Advanced Voice Mode managed 0.7%, GPT-Live-1 High reaches 75.2%.

Worth being honest about what those numbers actually mean. GPQA and BrowseComp were not designed for voice models, and the gains reflect GPT-5.5 Thinking doing the heavy lifting in the background rather than the voice model itself becoming a stronger reasoner. The voice model’s job is to have the conversation; GPT-5.5’s job is to get the answer right. The architecture separates those two things cleanly, which is the point.

In conversational head-to-heads over five to ten minute sessions, GPT-Live-1 and GPT-Live-1 mini are strongly preferred over Advanced Voice Mode on turn-taking, naturalness, handling of interruptions, and overall flow.

What This Means for You

If you use ChatGPT on your phone or browser: The update is already rolling out. GPT-Live-1 becomes the default for Plus and Pro subscribers, GPT-Live-1 mini for Free users. You can choose between Instant, Medium, and High reasoning depending on how much thinking time you want ChatGPT to spend on a response.

The practical difference is in longer conversations. ChatGPT’s voice product lead has described 30 to 40 minute walking conversations as a routine use case now. The old model was fine for quick queries; this one is designed for extended back-and-forth. You can interrupt, pause to think, ask it to slow down, and it adjusts without losing the thread.

Visual cards now appear during voice conversations for things like weather, stocks, and sports scores. Search, memory, images, and file uploads continue to work. Nine existing voices have been remastered for the new architecture.

If you are building on the API: GPT-Live is not in the API yet. OpenAI has opened a waitlist signup at openai.com/form/gpt-live-1-in-the-api/ and describes access as “coming soon.” Business, Enterprise, and Edu workspaces are also excluded from this launch.

That is a meaningful gap. A voice agent built on this architecture could hold a natural customer conversation while simultaneously querying a database or running a search, with no dead air. The use case is obvious; the access is not there yet.

If you are a parent or have a teen using ChatGPT: Age-appropriate behaviour has been trained directly into the model for teen users. Parents can control voice access through Parental Controls, and in situations involving signs of potential self-harm or suicidal intent, linked parents may be notified. That is a meaningful addition given how voice interactions tend to be more personal in tone than text.

What Is Still Missing

A few gaps worth flagging before you update your expectations. Video and screen sharing are not supported at launch, which limits the visual context use cases that competitors like Google Gemini Live already cover. OpenAI says these capabilities are coming.

Multilingual quality is uneven. OpenAI’s own documentation acknowledges that some languages produce non-native accents. The Hindi demo at launch drew criticism for sounding heavily American-accented and bookish in tone. That matters if you are evaluating this for non-English speaking markets.

The API delay is the most significant constraint for anyone hoping to build on this architecture immediately. There is no timeline beyond “soon.”

Safety and Voice Cloning

OpenAI’s system card for GPT-Live covers a few areas worth knowing. The model was tested with audio-native safety evaluations covering self-harm, emotional reliance, violence, and sexual content, and red-teamed specifically for voice-unique risks. Real-time intervention is built in: if a potentially unsafe output is detected, the system can guide the model mid-speech or cut the audio entirely in high-risk situations.

One firm design decision: GPT-Live uses only preset voices and will not mimic real people’s voices. Given the state of voice cloning concerns, that is a sensible line to draw clearly.

OpenAI’s Safety Advisory Group reviewed both models and concluded that neither, operating without delegation, reaches the High threshold in any of the company’s Preparedness Framework categories covering biological and chemical risk, AI self-improvement, or cybersecurity.

The Broader Picture

150 million people use ChatGPT voice features every week, according to OpenAI. Shipping full-duplex architecture to that entire base on day one, with a free tier included, is a wide surface area for a genuinely new interaction model to land on.

The free-to-paid tier split also encodes something about OpenAI’s product thinking: let everyone experience the new voice mode, reserve the more capable model for subscribers. That is a conversion funnel built into the feature itself, and it is worth recognising it as such.

For now, the practical advice is straightforward. If you use ChatGPT Voice regularly, update your app and spend some time with a longer conversation. If you are a developer planning to build on voice capabilities, sign up for the API waitlist and keep an eye on what comes through when enterprise access opens.