AI that turns into text before you finish speaking: What changes with Microsoft's new voice AI?
Microsoft announced on October 1, 2026 a new suite of voice AI tools that can return text while a speaker is still talking. The transcription service, called MAI-Transcribe-2-Streaming, shows tentative text about 100 milliseconds after audio arrives and refines it as the utterance continues. It supports 60 languages, including Japanese, while the companion text‑to‑speech models – MAI-Voice‑2.1…
What you can do with it
AI at Work →Generate live captions with MAI-Transcribe-2-Streaming
Turns spoken words into text in real time.
- Who for
- marketers
- Cost
- paid from $0.54 per hour
- Effort
- minutes
Use it for
- Subtitle meetings
- Capture lecture notes
- Generate live captions
Watch out Only in public preview, not for production
MAI-Transcribe-2-Streaming in the tool directory · 2 changes covered
Key points from the news
- Microsoft’s MAI-Transcribe-2-Streaming returns tentative text within 100 ms of audio input
- Pricing: $0.54 per hour for transcription; $22/$15 per million characters for standard/Flash voice generation
- Transcription supports 60 languages (including Japanese); voice generation supports 23 languages, currently no Japanese voice option
Pricing is published in US dollars: transcription costs $0.54 per hour of audio (introductory rate through the end of 2026), standard voice generation $22 per million characters, and Flash voice generation $15 per million characters. The services are in public preview and Microsoft warns they are not recommended for production use. Developers can try them via the MAI Playground or integrate them through Microsoft Foundry, but they must verify language support and cost implications for their applications.
Potential use cases include live subtitles for meetings or lectures, real‑time note taking, and quick voice responses in interactive apps. Accuracy for Japanese proper nouns, dialects, or noisy environments has not been independently verified, so adopters should test the tools before committing to production workloads.
Model page: MAI-Transcribe-2-Streaming →
The story so far
2 episodes →- AI that turns into text before you finish speaking: What changes with Microsoft's new voice AI?this story
AI that turns into text before you finish speaking: What changes with Microsoft's new voice AI?
note.com · 4 October 2026
Loading the full article…
This text was published by note.com and written by ひぃ / AIエージェント開発. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Agents & Tools
All →- OpenAI admits AI agent leaked 53 user images to public site · 1 src
- RemoveMacAI tool disables Apple Intelligence and removes models on macOS 27 · 1 src
- OpenAI's GPT-6 Astra cheats by downloading human bot Stardust in StarSkirmish · 1 src
- Google researchers find way to prevent self-improving AI agents from memorizing tests · 1 src
- SCM launches local AI search for photos and videos on macOS · 1 src
Comments
via GitHub Discussions