Capability update · Speech to text
Grok Transcribe 2.0: fewer word errors at the same price
An independent benchmark reports word error rates of 4.0% for the old model and 2.3% for the new one. The practical choice also depends on speed, language and access through the product you use.
After capture, the words still need to be right
On September 18, SpaceXAI released Grok Voice Transcribe 2.0, an update to its speech-to-text model. It handles recorded files and live audio, with timestamps and optional speaker labels to help locate words in the recording.
For meeting notes and video captions, the useful question is whether names, numbers and conversations require fewer corrections. The announcement highlights these cases, and an independent benchmark offers a comparison with the previous version.

Sources and further reading
- Grok Voice Transcribe 2.0 announcementOfficial · September 18The update, audio-duration pricing and official features.
- Speech to Text documentationOfficial documentation · Read September 20Model selection, file formats, language information and access.
Fewer errors, but not a simultaneous speed gain
As checked on September 20, Artificial Analysis reports a 2.3% word error rate for Grok 2.0 and 4.0% for 1.0 in the same non-streaming benchmark. Word error rate counts substitutions, deletions and insertions; lower is better. It does not promise that result for every recording.
The same table gives speed factors of 242.1 for 1.0 and 155.1 for 2.0: audio seconds processed per second, with higher values meaning faster processing. These are seven-day medians, not live-caption latency. Recorded interviews may put more weight on correction work; live captions need a separate responsiveness check.
| Version | Word error rate ↓ | Speed factor ↑ |
|---|---|---|
| Grok 1.0 | 4.0% | 242.1 |
| Grok 2.0 | 2.3% | 155.1 |
Sources and further reading
- Artificial Analysis non-streaming benchmarkIndependent benchmark · Read September 20Errors and speed for 1.0 and 2.0 in the same table, not our measurements.
Start with your recording and the product you use
API pricing is unchanged: $0.10 per hour of recorded audio, or $0.20 for streaming. The developer console provides a trial entry, and the API accepts formats including MP3 and M4A. These are service prices based on audio duration; a recording app or device chooses its own model and charges.
For English interviews or video, try a short recording you can check yourself and inspect names, numbers and speaker attribution. The current documentation’s language and text-formatting list does not list Chinese, and we have not tested Chinese meetings. If Chinese is your main need, compare again when product access and Chinese results are clearer.
- Open the official transcription trialDeveloper console; account required. Available credit is shown in your account.
- Compare the recording card: capture and transcription allowancesCapture hardware and transcription are separate choices; this update does not establish an upgrade to that device.
Sources and further reading
- Grok Voice Transcribe 2.0 announcementOfficial · September 18The update, audio-duration pricing and official features.
- Speech to Text documentationOfficial documentation · Read September 20Model selection, file formats, language information and access.