GPT Realtime Whisper, Explained Simply
What GPT Realtime Whisper is, what it does well, and what it costs โ released 2026-05-07 by OpenAI.
- It is a new OpenAI tool that turns spoken audio into written text.
- It works in real time, meaning it writes out the words as you are still speaking.
- It is designed for developers who want to build apps with live captions or transcriptions.
Want a heads-up when OpenAI Audio updates?
We email you the moment a new OpenAI Audio version drops โ plus plain-English release notes like these. No spam.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
GPT Realtime Whisper is a streaming speech-to-text model. That means it listens to someone talking and instantly types out what they are saying. Developers would want this to build apps that need live captions, voice notes, or instant text records of a conversation.
What it does well
Its main strength is "streaming" speech-to-text. Instead of making you wait until you finish a long recording to get the text, it writes the words out live as you speak them.
How it compares
OpenAI released this alongside two other new audio models. One is called GPT-Realtime-2, which is built for having full back-and-forth voice conversations. Another is GPT-Realtime-Translate, which translates spoken languages on the fly. GPT Realtime Whisper is different because its only job is to listen to speech and instantly turn it into written text, rather than talking back or translating.
This is our plain-language summary. Read the complete, official notes on the OpenAI Audio official changelog โ.
