Microsoft Launches MAI-Transcribe-2, Undercutting Rivals on Speech-to-Text Pricing

Microsoft AI has introduced MAI-Transcribe-2, a speech-recognition model the company says beats rival offerings from OpenAI, Google, and ElevenLabs on both speed and accuracy — while costing dramatically less to run. The model is launching at an introductory price of $0.10 per audio hour, a cut of roughly 72 percent from what Microsoft charged for the first version of the model just five months earlier.

The new system supports 60 languages, up from 43 in the prior release, and adds features aimed squarely at professional transcription workflows: speaker diarization, configurable output styles, and word-level timestamps. Microsoft says the model tops the FLEURS multilingual benchmark with a 5.2 percent average word-error rate, and that independent testing shows it running many times faster than comparable systems from OpenAI and ElevenLabs.

The release is the latest sign of aggressive price competition in the speech-AI market, where cost per audio hour has become a key battleground as companies build voice assistants, meeting-transcription tools, and accessibility features on top of these models. Cheaper, faster transcription could accelerate adoption in call centers, media production, and enterprise software that relies on converting speech into searchable, structured text.

For now, Microsoft is betting that a combination of aggressive pricing and benchmark-topping accuracy will pull developers away from established players — a strategy that mirrors the broader price war playing out across large language models more generally.

Leave a Reply

Your email address will not be published. Required fields are marked *