OpenAI is expanding its audio capabilities by releasing two new specialized models: GPT-transcribe for high-precision audio-to-text conversion and GPT-live-transcribe for real-time processing with minimal latency.

image
image

What Happened

OpenAI announced new API models for Speech-to-Text (STT) tasks. GPT-transcribe is focused on batch processing and high-precision archiving of audio files at a price of $0.0045 per minute. GPT-live-transcribe is designed for interactive streaming and low-latency interfaces at a rate of $0.017 per minute. Both models support multilingual capabilities and context prompting.

Context

Previously, developers had to rely on general-purpose or heavy models that were not always optimal in terms of the cost-to-speed ratio. Segmenting the market into high-precision archiving and ultra-fast real-time streaming allows for a more flexible approach to voice service architecture.

Why It Matters for the Industry

For the industry, this means the ability to optimize costs and latency depending on the specific use case. Developers can implement specialized API calls instead of universal models, reducing architectural complexity and increasing the economic efficiency of mass transcription services and voice AI agents.

Why It Matters for Users

Users will gain access to higher quality and cheaper tools for automatic subtitling, meeting recordings, and interacting with voice assistants. The introduction of models with separate pricing policies makes high-quality speech recognition more accessible for a wide range of applications.

What Is Not Yet Known / Limitations

Technical details regarding the differences in accuracy between the models depending on noise levels or specific languages are not specified; the focus has been shifted toward separation by use case (latency vs. accuracy).

Sources

Author

Look at AI, Editorial Team