Capability update · Speech to text

Grok Transcribe 2.0: fewer word errors at the same price

An independent benchmark reports word error rates of 4.0% for the old model and 2.3% for the new one. The practical choice also depends on speed, language and access through the product you use.

In this article3 chapters

After capture, the words still need to be right

On September 18, SpaceXAI released Grok Voice Transcribe 2.0, an update to its speech-to-text model. It handles recorded files and live audio, with timestamps and optional speaker labels to help locate words in the recording.

For meeting notes and video captions, the useful question is whether names, numbers and conversations require fewer corrections. The announcement highlights these cases, and an independent benchmark offers a comparison with the previous version.

Official Grok Voice Transcribe 2.0 announcement image.
Official SpaceXAI announcement artwork identifying the release, not a benchmark chart. · Open full-size image
Sources and further reading

Fewer errors, but not a simultaneous speed gain

As checked on September 20, Artificial Analysis reports a 2.3% word error rate for Grok 2.0 and 4.0% for 1.0 in the same non-streaming benchmark. Word error rate counts substitutions, deletions and insertions; lower is better. It does not promise that result for every recording.

The same table gives speed factors of 242.1 for 1.0 and 155.1 for 2.0: audio seconds processed per second, with higher values meaning faster processing. These are seven-day medians, not live-caption latency. Recorded interviews may put more weight on correction work; live captions need a separate responsiveness check.

Same non-streaming benchmark · Artificial Analysis · Checked September 20, 2026
VersionWord error rate ↓Speed factor ↑
Grok 1.04.0%242.1
Grok 2.02.3%155.1
Sources and further reading

Start with your recording and the product you use

API pricing is unchanged: $0.10 per hour of recorded audio, or $0.20 for streaming. The developer console provides a trial entry, and the API accepts formats including MP3 and M4A. These are service prices based on audio duration; a recording app or device chooses its own model and charges.

For English interviews or video, try a short recording you can check yourself and inspect names, numbers and speaker attribution. The current documentation’s language and text-formatting list does not list Chinese, and we have not tested Chinese meetings. If Chinese is your main need, compare again when product access and Chinese results are clearer.

Sources and further reading