OTHER

OpenAI Introduces New Voice Models for Improved Natural Live Conversations

OpenAI has launched new conversational models, called GPT-Live-1 and GPT-Live-1 mini, claiming they deliver more natural interactions and excel in turn-taking. These full-duplex models enable simultaneous speaking and listening, allowing users to interrupt organically and supporting features like live translation.

The company will replace the current Advanced Voice Mode in ChatGPT with GPT-Live-1 mini as the default setting. Users on paid subscriptions will gain access to the more sophisticated GPT-Live-1 model. The previous version combined a speech-to-text component for transcribing, a large language model for generating responses, and a text-to-speech model for answering queries.

During a recent press briefing, OpenAI highlighted that the new models tackle challenges such as unintentionally interrupting speakers and lacking adequate intelligence for responding to questions. The updated models will route queries to newer text models like GPT-5.5 for tasks involving search, reasoning, or agent-like functions while ensuring smooth conversation flow.

OpenAI also showcased the models’ ability to remain silent for long periods, allowing them to absorb context until prompted. Furthermore, the new voice mode is capable of visually presenting some information, made possible by access to the latest GPT models. Other startups, like Monogram—having secured $40 million in seed funding from DST and Lux Capital—are also exploring visual responses to enhance interactivity in digital assistants.

The refreshed voice mode in ChatGPT is designed for longer conversations. Atty Eleti, the product lead for ChatGPT Voice, shared that he enjoyed conversations lasting between 30 and 40 minutes using the voice feature during walks.

OpenAI envisions voice as the key interface for complex computing tasks. There is speculation about the potential release of AI-capable earbuds this year, although no details about hardware products have been disclosed.

“We believe this will eventually enable the use of voice as the main interface for managing increasingly complex, long-term agentic tasks. We see voice as the future interface for a range of tasks, akin to what is currently being accomplished with Codex and ChatGPT,” Eleti noted.

OpenAI has been continuously improving voice features for years to enhance the naturalness of ChatGPT’s voice mode. The company reported that over 150 million users engage with ChatGPT through features like Voice and Dictation.

Competitors are also working to boost the expressiveness of their assistants.

Both Apple and Amazon have upgraded their assistants to enable more conversational interactions with improved context management. Startups like Sesame, co-founded by Oculus co-founder Brendan Iribe and Ankit Kumar, have developed AI assistants that engage in more fluid conversations while performing background tasks.

OpenAI is following a similar direction, aiming for users to interact hands-free with its assistant for extended periods. While the new voice mode is reported to sound more natural, the company clarified that it is not intended to create an AI companion. It stressed that the new models include safeguards to provide age-appropriate responses for teenagers and to offer resources if conversations turn toward sensitive subjects such as self-harm.

The new voice mode still needs fine-tuning. In a demonstration showcasing the live translation feature in Hindi, the assistant displayed a pronounced American accent and produced Hindi that sounded unnatural and somewhat formal. The company mentioned that the new mode is optimized for “most spoken languages,” but did not indicate which ones specifically.

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.