From model to experience · Live speech

MAI adds live transcription: how soon do the words become stable?

Early text can change as context arrives; a final result confirms a segment. Compare MAI and Grok on errors, waiting after speech ends and audio cost in the same AA streaming test, then see how this connects to live captions and voice interaction.

This is an API for developers to integrate into products. This reading uses official documentation and AA's external test; BitShovel has not tested Chinese meetings or a complete voice assistant.

Microsoft's MAI streaming transcription announcement illustration: pink punctuation forms a receding pattern against a blue background.
Official Microsoft AI announcement illustration identifying this speech-model update; it is not a live-caption screenshot or latency test. · Open full-size image
In this article4 chapters

中文

Why the words appearing as you speak can change

Microsoft released MAI-Transcribe-2-Streaming on October 1. Audio arrives continuously and text returns incrementally, across 60 languages with continuous automatic language detection. It sends partial transcripts, revises them as more context arrives, then confirms stable segments. Words that appear in live captions and later change are part of this process.

Live captions need text to keep up with speech and settle promptly when an utterance ends. A recorded interview can be submitted as a complete file, with processing speed and correction work assessed afterwards. A voice assistant must also interpret the request, use tools and generate a reply, so transcription finalization is only one part of the wait for a spoken response.

Sources and further reading

中文

Stream the same way and compare errors, waiting and price

In Artificial Analysis's streaming benchmark read on October 2, MAI's final word error rate is about 2.5%, versus 2.7% for Grok Voice Transcribe 2.0 Streaming. Mean waiting from speech end to final transcription is about 0.13 and 0.49 seconds, respectively. WER counts substitutions, deletions and insertions; lower is better. MAI has fewer errors and a shorter finalization wait in this test, at a higher audio price.

This timer starts when SileroVAD detects the end of speech and includes network waiting in the test request. The benchmark uses about eight hours of audio, weighted 50% AA-AgentTalk, 25% VoxPopuli and 25% Earnings22. Microsoft's just over 100 ms to the first partial after receiving audio has different start and end points. The 0.13 seconds is also not the full delay from speaking to visible captions, and these results do not establish performance on Chinese meetings.

Same streaming test · Artificial Analysis · Checked October 2, 2026; lower is better for all three metrics
ModelFinal WERSpeech end → finalPer 1,000 audio minutes
MAI-Transcribe-2-Streaming2.5%0.13 s$9.00
Grok Voice Transcribe 2.0 (Streaming)2.7%0.49 s$3.33
Same AA streaming test: MAI final word error rate 2.5%, about 0.13 seconds from speech end to final text, and $9 per 1,000 minutes; Grok 2.7%, 0.49 seconds and $3.33. Lower is better for all three.
Drawn by BitShovel from Artificial Analysis data read on October 2, 2026, with rounded values. Final WER is final word error rate; Speech end → final measures waiting after detected speech end. This compares an external benchmark, not BitShovel model measurements. · Open full-size image
Sources and further reading

中文

Audio-based pricing and actual access

MAI Streaming's introductory price is $0.54 per audio hour through the end of 2026. That converts to $9 per 1,000 minutes, matching AA's table. Cost accumulates with the audio duration sent to the service. Grok Streaming is $3.33 per 1,000 minutes in the same table; assess waiting, correction work and the cost of your audio together.

Microsoft provides access through Foundry and Azure Speech. Vercel has a dedicated MAI Streaming page with a microphone-based live transcription playground. It bills playground usage to the team at API rates and lists $5 of credit every 30 days for users who have not made a payment; check the account's actual balance. These are service and gateway prices; an app using the model sets its own charges.

Sources and further reading

中文

Watch an utterance appear, then connect it to a voice

To explore the difference, try a short utterance you can check, including a name or number, then pause and watch how the partial text changes and when it settles. Microsoft's Chatter Beta combines transcription and voice models into an assistant. It offers a way to explore conversational pacing, while the full reply spans several stages and cannot reproduce AA's transcription-finalization measure.

For recorded interviews, continue with our Grok 2.0 non-streaming comparison of full-file errors and processing speed. For a character or short film, our Gemini 3.8 TTS reading covers the text-to-voice stage: voice identity, pauses and emotion. Distinguishing speech recognition from speaking text helps explain which capabilities and human choices make up a voice experience.

Sources and further reading

See how these changes connect