Reasoning

Realtime speech-to-speech model with configurable reasoning, tool use, and robust voice-agent behavior

openai/gpt-realtime-2.1 2026-07-06 128K context $4/M input $24/M output
2 providers

Streaming speech-to-text model for low-latency transcript deltas from live audio

openai/gpt-realtime-whisper 2026-05-07 Not documented context Input not listed Output not listed
1 provider

The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced...

openai/gpt-audio 128K context $2.5/M input $10/M output

A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.60 per million...

openai/gpt-audio-mini 128K context $0.6/M input $2.4/M output