According to monitoring by Dongcha Beating, OpenAI has released two speech-to-text models: GPT Transcribe and GPT Live Transcribe. The former handles recorded audio files and batch tasks, while the latter is designed for real-time scenarios such as live captions and phone calls. Both models incorporate context, keywords, and language hints to enhance recognition accuracy for short phrases, numbers, technical terms, multiple accents, and noisy environments. Artificial Analysis measured a word error rate of 3.31% for GPT Transcribe, 0.7 percentage points lower than the previous GPT-4o Transcribe model. The price has also been reduced by 25%, with a charge of $4.5 per 1,000 minutes of audio.
All Comments