THE CRUNCH
Microsoft AI has released MAI-Transcribe-2-Streaming, a model for real-time transcription that the company says ranks first for accuracy on Artificial Analysis. It transcribes 60 languages and delivers its first partial results in just over 100 milliseconds, which Microsoft says lets voice agents respond while someone is still mid-sentence. An hour of audio costs $0.54 at an introductory price through the end of the year.
Alongside it, Microsoft shipped two text-to-speech models. MAI-Voice-2.1 is built to speak 23 languages in a single voice, with a native accent in each, while the Flash variant reportedly reaches 150 milliseconds of latency and costs $15 per million characters, down from $22. Both voice models can clone a voice from a few seconds of reference audio, with built-in safeguards intended to prevent misuse.
The models are available through Microsoft Foundry and the MAI Playground, with the two voice models also listed on OpenRouter. In one test cited by Microsoft, roughly half of 4,000 participants thought the generated voices belonged to a real person.


