OpenAI launches new voice intelligence features in its API
OpenAI Expands API with Advanced Voice Features
OpenAI has rolled out several advanced voice intelligence tools in its API, making it easier for developers to build apps that communicate, transcribe, and translate in real time. The new update introduces three major features designed to elevate conversational technology for a variety of use cases.
Key Additions to OpenAI’s API
- GPT-Realtime-2: A state-of-the-art voice model that simulates natural conversation, now using GPT-5-level reasoning for handling complex user requests.
- GPT-Realtime-Translate: Offers real-time translation, keeping up with live conversations. It supports over 70 input languages and translates into 13 output languages.
- GPT-Realtime-Whisper: This tool provides instant speech-to-text transcription, capturing verbal interactions as they happen.
Potential Applications
These updates are expected to benefit businesses looking to enhance customer service, as well as sectors such as education, media, live events, and creator platforms. OpenAI highlights how real-time voice interfaces can move beyond simple responses to actively listening, reasoning, and taking action as conversations progress.
Safeguards and Accessibility
OpenAI has implemented safeguards to prevent misuse of these new tools, including systems that can pause conversations if harmful content is detected. This helps tackle concerns around spam, fraud, and abuse.
The new models are available through OpenAI's Realtime API, with translation and transcription billed per minute, and GPT-Realtime-2 billed by token usage.
For full details, visit the original article on TechCrunch.