Microsoft released two in-house models into public preview on Wednesday: MAI-Image-2.5-Pro, its highest-fidelity image generator, and MAI-Voice-2-Flash, a speech model targeting high-volume enterprise workloads. The numbers attached to these launches are the real story. MAI-Image-2.5 cuts GPU costs by 84% in PowerPoint versus GPT-Image-2. MAI-Voice-2-Flash reduces GPU costs by 89% in Dynamics 365 Contact Center. In OneDrive, the image model is now the default for key editing scenarios, where it produced a 26% increase in save rates and 2.5 times greater efficiency under production load.
The deployment footprint is no longer theoretical. MAI-Image-2.5 runs Bing Image Creator end to end, the first time that consumer tool is fully in-house. Dragon Copilot, used by 170,000 medical providers processing 28 million patient encounters last quarter, now runs on MAI-Transcribe-1.5 across 58 languages, with a reported 50% relative reduction in transcription error rates. MAI-Code-1-Flash in GitHub Copilot achieves a 10% higher code accept rate than GPT-5.4 Mini and Claude Haiku 4.5 in VS Code, while using 10% fewer median tokens. Then Microsoft fine-tuned that same checkpoint inside an Excel reinforcement learning environment and matched GPT-5.6 on common spreadsheet tasks.
What makes the full piece worth reading is the hardware argument buried at the end. That Excel model runs on H100 and A100 GPUs, not the latest-generation accelerators. A frontier-adjacent model that runs on two-generation-old silicon changes deployment economics at scale and frees newer hardware, including Microsoft's GB200 cluster, for training. Microsoft is not just replacing OpenAI models in its products. It is building a cost structure that third-party model providers cannot easily match inside Microsoft's own stack.
[READ ORIGINAL →]