Last Updated: 8 May 2026

OpenAI is giving developers new ways to build apps that can talk with people, understand speech, translate live conversations, and turn spoken words into text as they happen. The company added three new audio models to its Realtime API: GPT Realtime 2, GPT Realtime Translate, and GPT Realtime Whisper.
The simple idea is clear: apps are starting to listen and respond more like a person would. A travel app could help with a booking by voice. A support app could answer a customer in their own language. A meeting tool could write notes while people are still talking.
OpenAI’s new voice intelligence update is built around three tools. GPT Realtime 2 is made for spoken AI agents that can follow longer talks, handle harder requests, and use tools during a conversation. GPT Realtime Translate is made for live speech translation. GPT Realtime Whisper is made for live speech to text.
For regular users, this means future apps may feel less like typing into a box and more like asking for help out loud. A person could ask for a schedule change, a product answer, a travel update, or a short summary without switching to text.
For developers, the update gives more building blocks for voice apps. These tools can be used in customer service, education, events, media, creator platforms, sales, healthcare, recruiting, and other places where spoken conversations happen often.
This means OpenAI is pushing deeper into voice AI for apps, businesses, and daily users.
GPT Realtime 2 is the main voice model in this release. It is designed to keep a conversation going while the AI works through a request, checks tools, handles corrections, and responds in a natural way. OpenAI says it also supports longer context, which helps the model remember more during longer sessions.
That can make a big difference in voice apps. People interrupt. They change their mind. They forget details. They add new information halfway through a sentence. A useful voice agent needs to follow that without falling apart.
OpenAI also says the model can use short phrases such as “let me check that” before giving the main answer. That may sound small, but it helps users know the app is still working instead of frozen or lost.
Features of GPT Realtime 2
GPT Realtime Translate is built for live multilingual conversations. It supports more than 70 input languages and 13 output languages, which means a person can speak in one language and hear the answer in another during the same conversation.
This could be useful in many simple situations. A customer could call support in their preferred language. A teacher could explain a topic to a mixed class. A creator could share a video with people in other countries without waiting for a separate version.
The hard part is not only changing words from one language to another. Real speech is messy. People pause, change topics, use local words, and speak with different accents. The model is meant to keep up with that kind of speech while keeping the meaning clear.
Features of GPT Realtime Translate
GPT Realtime Whisper is OpenAI’s new streaming speech to text model. It writes down speech while a person is still speaking, instead of waiting until the full conversation ends.
This can help with live captions, meeting notes, classroom tools, event transcripts, and support records. It can also help voice agents understand what a user is saying during longer talks.
The main value is speed. When speech becomes text right away, apps can use that text for summaries, follow ups, records, or other work. A sales call could produce notes while the call is still active. A meeting app could capture decisions before people leave.
Features of GPT Realtime Whisper
OpenAI says the Realtime API includes safety systems to help stop misuse. The company says some conversations can be stopped when harmful content is detected, and developers can add their own safety rules. OpenAI also says developers must make it clear when users are speaking with AI unless the situation already makes that obvious.
The new models are already being tested by companies such as Zillow, Priceline, and Deutsche Telekom. Reuters also named these three as customers testing the models.
Pricing is split by model. GPT Realtime 2 starts at $32 per million audio input tokens and $64 per million audio output tokens. GPT Realtime Translate costs $0.034 per minute. GPT Realtime Whisper costs $0.017 per minute.
The models are available in OpenAI’s Realtime API and can be tested in the developer Playground.
This update pushes OpenAI deeper into voice AI, an area where apps need to feel fast, clear, and useful. Text chat is still important, but voice can make AI easier to use for people who do not want to type or read long answers.
It also gives OpenAI a stronger place in business software. Customer support, travel, sales, education, and meetings all depend on spoken conversations. If developers build strong products with these models, OpenAI becomes part of more daily workflows.
OpenAI wants its API to power AI that can listen, speak, translate, and help in real time. That makes voice one of its most important bets for the future.
OpenAI’s new voice intelligence features show where AI apps are moving. The focus is no longer only on chat boxes. Developers can now build apps that listen, speak, translate, transcribe, and complete tasks during live conversations.
For users, the change could be simple: fewer forms, fewer clicks, and more natural voice help. For developers and businesses, the update opens more ways to build voice tools for support, travel, learning, media, meetings, and daily work.
OpenAI launched three new audio models for its API: GPT Realtime 2, GPT Realtime Translate, and GPT Realtime Whisper. These models are made for apps that use voice in real time. GPT Realtime 2 helps voice agents talk and complete tasks. GPT Realtime Translate helps translate speech during live conversations. GPT Realtime Whisper turns speech into text while people are speaking. Together, they give developers more ways to build voice based apps for support, travel, education, media, events, and business use.
GPT Realtime 2 is OpenAI’s new voice model for live AI agents. It is designed to understand spoken requests, keep track of context, handle corrections, and use tools during a conversation. For example, a user might ask a travel app to change a booking or ask a real estate app to search for homes and schedule a visit. The model can respond by voice while working through the task. The goal is to make voice apps more useful during normal, changing conversations.
GPT Realtime Translate is a live voice translation model. It can take speech from more than 70 input languages and produce speech in 13 output languages. This can help people talk across languages in customer support, education, sales, events, media, and creator tools. A user could speak in their own language and hear the response in another language during the same exchange. It is built for real speech, including changes in topic, accents, and local terms.
GPT Realtime Whisper is a live speech to text model. It writes down what people say while they are still speaking. This can be used for captions, meeting notes, classroom transcripts, event records, support calls, and business workflows. It can also help AI voice agents understand users during longer conversations. The main benefit is that spoken information becomes usable right away. Apps can save it, summarize it, search it, or send it into another workflow without waiting for the full call to end.
The new voice intelligence models are made for developers using OpenAI’s Realtime API. Businesses can use them to build customer support agents, travel assistants, meeting tools, learning apps, media products, and live translation services. OpenAI says the models are available in the Realtime API, and developers can test them in the Playground. The tools are not only for large companies. Smaller teams can also use them if they are building apps that need speech, translation, or live transcription.