Microsoft launches MAI-Transcribe-2-Streaming, its first real-time transcription model, and announces two new voice models

In 30 seconds
Microsoft AI introduced MAI-Transcribe-2-Streaming, a model that turns speech into text while the person is speaking. It works in 60 languages, including Spanish, and returns its first results in just over 100 milliseconds. It costs $0.54 per hour of audio as an introductory price through the end of the year, and it is in public preview in Microsoft Foundry.
On October 1, Microsoft AI, the team that builds Microsoft's own models (MAI), introduced MAI-Transcribe-2-Streaming, a model that turns speech into text while the person is speaking, without waiting for the sentence to end. It works in 60 languages, including Spanish, and automatically detects the language being spoken.
According to Microsoft, it returns its first results in just over 100 milliseconds and ranks number 1 for accuracy on Artificial Analysis, a website that compares AI models. This means a voice agent can start reasoning or using tools before the customer has finished speaking.
It comes with two new voices (models that turn text into speech): MAI-Voice-2.1, which speaks 23 languages with the same voice and a native accent in each one, and MAI-Voice-2.1-Flash, the version built for high volume and fast responses. Both can clone a voice from a few seconds of audio and include consent guardrails to prevent misuse.
Published prices: $0.54 per hour of audio for transcription, an introductory price through the end of the year; $22 per million characters for MAI-Voice-2.1 and $15 for the Flash version.
All three models are available in Microsoft Foundry, Microsoft's platform for building AI applications. The transcription model is in public preview: it can be used from any country, it is served from the Sweden Central region and Microsoft does not yet recommend it for production.
We believe this is a useful building block for small businesses that handle a lot of calls: Microsoft gives the example of customer service agents that understand the request while the customer is speaking and reply in natural speech. Since it is a preview, it is worth testing before using it in a live service.
Why it matters
A voice agent can start working out the request before the customer finishes speaking.
Official source: Microsoft AI


