docs
models / xAI / Grok STT

Grok STT

audioxai/grok-sttZDR supported

xAI's Grok speech-to-text model. Transcribes audio files into text across 25 languages with word-level timestamps, multichannel transcription, speaker diarization, and key-term biasing.

$0.00167list priceper audio minute
pricing
unitlist price
per audio minute$0.00167
per audio minute (streaming)$0.00333

Everyone pays the list price. Reaching Elite (hold or stake 1,337 $XAVA, or hold 1,337,000 $XINF) unlocks the Elite price and a bigger buyback of your token. How it works

curl
curl https://zinf.ai/v1/media/generations \
  -H "Authorization: Bearer $XINF_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"xai/grok-stt","input":{"text":"Hello from one key."}}'
limits
providerxAI
typeaudio
context window—
max output—
prompt storagenone
provider retentionnone ZDR supported
livesoon
requests 24h—
p50 latency—
uptime 7d—