Models & Assistants

GPT-Live voice gets file uploads, Projects integration, and becomes the default for Enterprise, Edu and Healthcare

ChatGPT's GPT-Live voice mode now supports file uploads mid-conversation, works inside Projects, and is the default voice experience for enterprise workspaces.

models assistants category

OpenAI has pushed a meaningful update to GPT-Live, the voice model family it introduced on 8 July 2026. Two new capabilities have arrived for all users, and for anyone working inside an Enterprise, Edu, or Healthcare workspace, GPT-Live is now the automatic default voice experience when an admin enables voice.

Here is what has changed and what it means for you.

You can now upload files mid-voice-conversation

This is the one that will change how people actually use voice mode day to day. Previously, GPT-Live was a purely spoken medium. If you wanted to discuss a document, a spreadsheet, or a report, you had to switch to a text chat. That limitation is gone.

You can now upload a file during a live voice conversation and ask questions about it, talk through its contents, or have GPT-Live summarise and explain it, all without breaking the conversational flow. The file sits in the chat alongside the spoken exchange, so the context carries through naturally.

In practice, this means you could pull up a PDF contract while talking through its terms, share a data export mid-call and ask GPT-Live to identify anomalies, or upload a draft document and discuss revisions verbally. The voice interface stops being a lightweight alternative to text chat and becomes a genuinely capable working mode.

Voice now works inside Projects

GPT-Live was previously restricted to standard chats. Users had been asking for Projects support, and OpenAI has now delivered it.

When you open a Project and switch to voice, GPT-Live can reference your recent project chats, any sources you have added to the project, and the custom instructions you have set. That means your voice conversations carry the same context and configuration as your text conversations within that project. If you have set up a project for a specific client, a research area, or a recurring workflow, voice is now a full participant in that space rather than a workaround outside it.

This is a meaningful quality-of-life change. Projects exist precisely to keep related context together, and voice mode working without that context was an obvious friction point.

Enterprise, Edu and Healthcare: GPT-Live is now the default, no extra toggle required

For workspace administrators, this is the most operationally significant part of the update.

Previously, rolling out GPT-Live to an Enterprise or Edu workspace required enabling both the Voice setting and a separate Early Model Access toggle. That two-step requirement has been removed. From now on, when a workspace owner enables Voice, GPT-Live is the experience their users get, without any additional configuration.

If you want voice available but prefer not to offer GPT-Live, you can disable voice entirely through workspace settings. There is no partial option described in the release notes, so the practical choice for admins is on or off.

This matters because it removes a common point of confusion during rollouts. Early Model Access toggles are easy to miss, and support queries about voice not appearing in settings were a predictable consequence. Removing that dependency simplifies deployment considerably.

A quick recap of what GPT-Live actually is

GPT-Live uses a full-duplex architecture, which means it listens and speaks simultaneously rather than waiting for you to finish before it starts processing. Turn-taking, interruptions, and changes in pace feel natural rather than mechanical. The model can acknowledge what you are saying mid-sentence with brief responses like “mhmm” or “yeah”, and it stays quiet when you need a moment to think.

For questions that need web search, deeper reasoning, or more complex work, GPT-Live delegates to a frontier model running in the background and brings the result back into the conversation when it is ready. While that work is happening, GPT-Live keeps the conversation going rather than going silent.

ChatGPT Voice settings let you switch between Live, Advanced, and Standard modes via Settings, then Voice.

Two voice modes for business workspaces

It is worth being clear on how voice is structured for business users, because there are now two distinct experiences:

Voice in Chat is powered by GPT-Live and is available on Desktop Chat, supported web, iOS, and Android. It is designed for natural real-time conversation, questions, brainstorming, and quick follow-ups. Business workspaces include one hour of Voice in Chat, with additional usage billed at 5 credits per minute.

Voice in Work and Codex is a separate mode designed for starting and monitoring agent tasks, checking progress, and coordinating multiple agents through a single conversation. It is available in the ChatGPT desktop app on macOS and Windows, with paired iOS remote access. It is not available on web or standalone mobile. Usage runs at approximately 6 credits per minute.

The file upload and Projects features described above apply to Voice in Chat, where GPT-Live operates.

SynthID watermarking on all GPT-Live audio

From 31 July 2026, all audio generated through GPT-Live via ChatGPT Voice and the OpenAI API includes SynthID watermarking. OpenAI’s public verification tool can detect the provenance signal in supported audio files, and an API is available so developers and compliance teams can build automated provenance checks into their own workflows.

OpenAI uses a dual-layer approach, combining SynthID’s embedded signal with C2PA verifiable metadata. For organisations with content governance obligations, this gives you a documented technical basis for AI provenance claims in audio output.

The short version

GPT-Live voice mode has moved from a useful novelty to a more complete working tool. File uploads remove a significant limitation, Projects integration means your voice conversations now share context with your text workflows, and the simplified Enterprise rollout reduces admin friction. If you have been holding off on voice mode because it felt disconnected from your actual work, these changes are worth revisiting.