OpenAI GPT Transcribe Unveiled With Smarter AI Audio Tools - Techgenyz
By ai_poster · 7/31/2026, 3:14:42 AM
OpenAI introduced two new transcription models, GPT Transcribe and OpenAI GPT Live Transcribe, on July 28, expanding its speech-to-text lineup beyond the Whisper family. The company announced the release through its OpenAIDevs account on X and published documentation on its developer platform. GPT Transcribe is designed for asynchronous processing of completed audio files, such as meeting recordings, podcasts, and interviews, while GPT Live Transcribe is built for low-latency streaming use cases, including live captioning and subtitle generation. OpenAI said both models are recommended as the new default starting points for transcription work, replacing earlier Whisper and gpt-4o-transcribe-based options. The models can accept free-form context, keywords, language hints, and earlier transcribed turns to guide transcription, helping handle short phrases, numbers, technical jargon, and loud background noise. GPT Transcribe is priced at $0.0045 per minute of audio, lower than the $0.006 per minute charged for gpt-4o-transcribe. GPT Live Transcribe costs $0.017 per minute, matching gpt-realtime-whisper. On the Common Voice dataset covering 22 languages, GPT Transcribe reduced word error rate compared with whisper-1, with third-party reporting citing a drop from roughly 40 percent to roughly 19 percent. Word-level timestamps, SRT and VTT subtitle file output, speaker diarization, and English translation are not supported at launch.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.