Grok Voice STT 1.0 is xAI's speech-to-text model. It supports transcription with word-level timestamps, optional speaker diarization, and multichannel audio.
Capability badges appear only where the provider catalog explicitly lists support. A blank cell means the source is silent, not that the feature is absent.
Specification
Capabilities
Recorded from the source catalog and provider listings.
Every result stays attached to its source, version, metric and harness. Scores from different versions are never merged, and the profile below plots each benchmark against its own population rather than on a shared scale.
ModelBench does not infer quality from price, context size, or model name. When a source publishes a comparable result, it appears here with its version and link.
Answered directly from the stored record — nothing here is generated beyond the catalog's own fields.
What is Grok Voice STT 1.0?
Grok Voice STT 1.0 is xAI's speech-to-text model. It supports transcription with word-level timestamps, optional speaker diarization, and multichannel audio. It is published by xAI and catalogued here from Models.dev.
What is the context length of Grok Voice STT 1.0?
Grok Voice STT 1.0 accepts up to 15K tokens of context and returns up to 15K output tokens.
Which providers serve Grok Voice STT 1.0?
1 provider list this model: ZenMux.
Are the weights for Grok Voice STT 1.0 open?
No. This model is served through hosted APIs only.
When was Grok Voice STT 1.0 released?
The catalog records a release date of 2026-08-04, last verified Sep 19, 2026.