From model to experience · Live speech
MAI adds live transcription: how soon do the words become stable?
Early text can change as context arrives; a final result confirms a segment. Compare MAI and Grok on errors, waiting after speech ends and audio cost in the same AA streaming test, then see how this connects to live captions and voice interaction.
This is an API for developers to integrate into products. This reading uses official documentation and AA's external test; BitShovel has not tested Chinese meetings or a complete voice assistant.

Why the words appearing as you speak can change
Microsoft released MAI-Transcribe-2-Streaming on October 1. Audio arrives continuously and text returns incrementally, across 60 languages with continuous automatic language detection. It sends partial transcripts, revises them as more context arrives, then confirms stable segments. Words that appear in live captions and later change are part of this process.
Live captions need text to keep up with speech and settle promptly when an utterance ends. A recorded interview can be submitted as a complete file, with processing speed and correction work assessed afterwards. A voice assistant must also interpret the request, use tools and generate a reply, so transcription finalization is only one part of the wait for a spoken response.
Sources and further reading
- Microsoft AI: its first streaming transcription modelOfficial release · October 160 languages, partial and stable transcripts, the vendor's first-hypothesis timing, introductory pricing and Chatter.
- Microsoft Foundry: new MAI voice modelsOfficial integration notes · October 1The distinction between recorded-audio and live-interaction uses, plus Foundry / Azure Speech access and pricing.
Stream the same way and compare errors, waiting and price
In Artificial Analysis's streaming benchmark read on October 2, MAI's final word error rate is about 2.5%, versus 2.7% for Grok Voice Transcribe 2.0 Streaming. Mean waiting from speech end to final transcription is about 0.13 and 0.49 seconds, respectively. WER counts substitutions, deletions and insertions; lower is better. MAI has fewer errors and a shorter finalization wait in this test, at a higher audio price.
This timer starts when SileroVAD detects the end of speech and includes network waiting in the test request. The benchmark uses about eight hours of audio, weighted 50% AA-AgentTalk, 25% VoxPopuli and 25% Earnings22. Microsoft's just over 100 ms to the first partial after receiving audio has different start and end points. The 0.13 seconds is also not the full delay from speaking to visible captions, and these results do not establish performance on Chinese meetings.
| Model | Final WER | Speech end → final | Per 1,000 audio minutes |
|---|---|---|---|
| MAI-Transcribe-2-Streaming | 2.5% | 0.13 s | $9.00 |
| Grok Voice Transcribe 2.0 (Streaming) | 2.7% | 0.49 s | $3.33 |
Sources and further reading
- Microsoft AI: its first streaming transcription modelOfficial release · October 160 languages, partial and stable transcripts, the vendor's first-hypothesis timing, introductory pricing and Chatter.
- Artificial Analysis: streaming transcription benchmarkIndependent benchmark · Checked October 2MAI / Grok word error rate, speech-end-to-final latency and price per 1,000 minutes, with datasets and timing definitions.
Audio-based pricing and actual access
MAI Streaming's introductory price is $0.54 per audio hour through the end of 2026. That converts to $9 per 1,000 minutes, matching AA's table. Cost accumulates with the audio duration sent to the service. Grok Streaming is $3.33 per 1,000 minutes in the same table; assess waiting, correction work and the cost of your audio together.
Microsoft provides access through Foundry and Azure Speech. Vercel has a dedicated MAI Streaming page with a microphone-based live transcription playground. It bills playground usage to the team at API rates and lists $5 of credit every 30 days for users who have not made a payment; check the account's actual balance. These are service and gateway prices; an app using the model sets its own charges.
- Open Vercel's live transcription playgroundCheck account billing and credits before starting a microphone session.
- Read Foundry's model access notes
Sources and further reading
- Microsoft AI: its first streaming transcription modelOfficial release · October 160 languages, partial and stable transcripts, the vendor's first-hypothesis timing, introductory pricing and Chatter.
- Microsoft Foundry: new MAI voice modelsOfficial integration notes · October 1The distinction between recorded-audio and live-interaction uses, plus Foundry / Azure Speech access and pricing.
- Artificial Analysis: streaming transcription benchmarkIndependent benchmark · Checked October 2MAI / Grok word error rate, speech-end-to-final latency and price per 1,000 minutes, with datasets and timing definitions.
- Vercel: MAI streaming transcription and playgroundProvider page · Checked October 2Live transcription interface, hourly audio pricing, team billing and playground credits.
Watch an utterance appear, then connect it to a voice
To explore the difference, try a short utterance you can check, including a name or number, then pause and watch how the partial text changes and when it settles. Microsoft's Chatter Beta combines transcription and voice models into an assistant. It offers a way to explore conversational pacing, while the full reply spans several stages and cannot reproduce AA's transcription-finalization measure.
For recorded interviews, continue with our Grok 2.0 non-streaming comparison of full-file errors and processing speed. For a character or short film, our Gemini 3.8 TTS reading covers the text-to-voice stage: voice identity, pauses and emotion. Distinguishing speech recognition from speaking text helps explain which capabilities and human choices make up a voice experience.
- Open the official Chatter Beta demo
- Transcribe after recording: Grok 2.0's non-streaming comparison
- Text to voice: direct a character's voice and each line
Sources and further reading
- Microsoft AI: its first streaming transcription modelOfficial release · October 160 languages, partial and stable transcripts, the vendor's first-hypothesis timing, introductory pricing and Chatter.
- MAI Playground: Chatter BetaOfficial demo entry · Checked October 2A public interface combining transcription and voice models into an assistant; this check inspected the entry without conducting a voice conversation.
See how these changes connect
How do timing and correction differ between a recorded file and live captions?
Speaking the recognized text is another stage: explore voice identity, pauses and each line's delivery.

