xAI logo

xAI: Grok Voice STT 1.0

Grok Voice STT 1.0 is xAI's speech-to-text model. It supports transcription with word-level timestamps, optional speaker diarization, and multichannel audio.

Source-linked Closed weights Released 2026-08-04
API record Report
Input
OutputT
Input priceNot listed
Output priceNot listed
Context15K
Max output15K
Providers1
Inference availability

Providers

Provider-specific identifiers, limits, and listed prices per million tokens. Every row links back to the provider's own documentation.

Report a price
Providers offering Grok Voice STT 1.0
ProviderProvider model IDContextMax outputInputOutputCache readCapabilitiesDocs
ZenMux x-ai/grok-voice-stt-1.0 15K 15K Not listed Docs ↗

Capability badges appear only where the provider catalog explicitly lists support. A blank cell means the source is silent, not that the feature is absent.

Specification

Capabilities

Recorded from the source catalog and provider listings.

× Reasoning No
× Tool calling No
? Structured output Unknown
Attachments Yes
× Vision input No
× Open weights No
Creator
xAI
Model family
Not documented
Knowledge cutoff
Not documented
License
Not documented
Release date
2026-08-04
Model ID
xai/grok-voice-stt-1.0
Published evaluations

Benchmarks

Every result stays attached to its source, version, metric and harness. Scores from different versions are never merged, and the profile below plots each benchmark against its own population rather than on a shared scale.

Benchmark registry
No published benchmark results

ModelBench does not infer quality from price, context size, or model name. When a source publishes a comparable result, it appears here with its version and link.

Read the methodology
Catalog activity

Change log

Field-level changes detected between successful source imports.

Full change log
No changes recorded

This record has not changed within the retained import history.

Public API

Use this record

Fetch the complete source-linked model record. No key, no account, no rate-limited tier.

API documentation
Endpoint
GET https://model.kyssta.lol/api/v1/models/xai/grok-voice-stt-1.0
curl
curl "https://model.kyssta.lol/api/v1/models/xai/grok-voice-stt-1.0"
Common questions

Frequently asked questions

Answered directly from the stored record — nothing here is generated beyond the catalog's own fields.

What is Grok Voice STT 1.0?

Grok Voice STT 1.0 is xAI's speech-to-text model. It supports transcription with word-level timestamps, optional speaker diarization, and multichannel audio. It is published by xAI and catalogued here from Models.dev.

What is the context length of Grok Voice STT 1.0?

Grok Voice STT 1.0 accepts up to 15K tokens of context and returns up to 15K output tokens.

Which providers serve Grok Voice STT 1.0?

1 provider list this model: ZenMux.

Are the weights for Grok Voice STT 1.0 open?

No. This model is served through hosted APIs only.

When was Grok Voice STT 1.0 released?

The catalog records a release date of 2026-08-04, last verified Sep 19, 2026.

More models from xAI

View all →