Zendoric
← Back to the day · July 28, 2026

Microsoft launches in-house AI models it says cut costs by up to 89% versus OpenAI

🕒 Published on Zendoric: July 28, 2026 · 00:38

Microsoft AI unveiled two new internally developed models on Wednesday, July 23, 2026, and released them in public preview: MAI-Image-2.5-Pro, its highest-fidelity image generator to date, and MAI-Voice-2-Flash, a voice model designed for high-volume enterprise workloads.

Microsoft AI on Wednesday 23 July 2026 unveiled two new models developed in-house and released them in public preview: MAI-Image-2.5-Pro, its highest-fidelity image generator to date, and MAI-Voice-2-Flash, a voice model built for high-volume enterprise workloads. The announcement, signed by Microsoft AI's Superintelligence team, comes roughly a year after the company committed to building its own «bespoke» models, and arrives with an unusual level of detail about where they are already running: Bing, PowerPoint, OneDrive, Dynamics 365, Excel, GitHub Copilot and Azure. The message, aimed at enterprise customers and, implicitly, at OpenAI, is that these homegrown models are no longer research projects but production infrastructure serving millions of users. As the company itself puts it in its blog: «Each of these improvements is a step toward the same goal: Microsoft products, powered by Microsoft models».

The two launches sit at opposite ends of what Microsoft calls the «quality-speed-cost curve», and that positioning is deliberate. MAI-Image-2.5-Pro targets the premium segment: visually striking images, detailed editing and accurate rendering of text within the image, a historic weak point of image generators. Microsoft priced it at $5 per million input text tokens, $8 per million input image tokens and $106 per million output image tokens. The base model, MAI-Image-2.5, had recently ranked number 2 in image editing on Arena, the community leaderboard that serves as the de facto scoreboard in multimedia generative AI. The creative industry appears to be taking note: Rob Reilly, global chief creative officer at advertising agency WPP, called the Pro model «a major leap forward for GenMedia tools», adding that «Microsoft has firmly established itself among the leaders in generative AI».

MAI-Voice-2-Flash goes the other way. First shown at Microsoft's Build conference, Flash runs twice as fast as MAI-Voice-2 and costs 32% less, priced at $15 per million characters. It is designed for the unglamorous but enormous high-volume voice market: call centers, voice agents and real-time speech applications, where latency and cost per call matter more than marginal gains in expressiveness. Together, the two models reflect a strategy of building families of models rather than a single flagship, on the premise that a creative studio chasing maximum fidelity has very different needs from a customer-service operation handling millions of calls a day.

Beyond the launches themselves, the most striking element is the deployment metrics Microsoft attached, which read as a systematic argument for replacing third-party frontier models across its product portfolio. Bing Image Creator now runs entirely on MAI-Image-2.5, end to end, the first time this consumer image tool has been completely in-house. In PowerPoint, Microsoft says MAI-Image-2.5 cuts GPU costs by up to 84% compared with GPT-Image-2, OpenAI's image model. In OneDrive, where MAI-Image-2.5 is now the default for key image-editing scenarios, the company reports a 26% increase in save rates, roughly 25% lower P95 latency and 2.5 times greater efficiency under medium-utilization production loads.

On the voice side, MAI-Voice-2-Flash now powers Dynamics 365 Contact Center, the platform used by customers such as T-Mobile and EasyJet, where Microsoft claims GPU cost reductions of up to 89%. The model is also integrated into Azure Voice Live, aimed at developers building speech-to-speech agents. Perhaps the most significant deployment is in healthcare: Dragon Copilot, used by 170,000 medical professionals and which processed 28 million patient encounters last quarter, now runs on MAI-Transcribe-1.5 for its multilingual workflow across 58 languages. Microsoft says its internal evaluations show a 50% relative reduction in both transcription and language-identification errors in most languages, a notable claim in a field where transcription errors can feed directly into clinical notes.

In a companion post the same day, Microsoft detailed the methodology behind these results, which it calls its «hill-climbing machine»: an integrated flywheel of data, models and the product «harness» around them. The clearest example is MAI-Code-1-Flash, the lightweight coding model launched in GitHub Copilot in June. Microsoft says it achieves a code acceptance rate roughly 10% higher than GPT-5.4 Mini and Claude Haiku 4.5 in VS Code, while using 10% fewer average tokens. Developer retention tells a similar story: users were 6% more likely to return across multiple days than with GPT-5.4 Mini, and 11% more likely than with Claude Haiku 4.5.

Microsoft went a step further: it took the MAI-Code-1-Flash checkpoint and kept training it inside an Excel-specific reinforcement learning environment, teaching a coding model the tools and workflows of spreadsheet knowledge. The result, according to user feedback in production, is a model on par with GPT-5.6 on the most common Excel tasks, but small enough to run on older Nvidia GPUs such as H100s and even A100s, rather than requiring the latest accelerators. That hardware detail is worth underlining: every major AI company is competing for access to the most advanced chips, and a model that delivers near-frontier quality on two-generation-old silicon substantially changes the economics of deployment. It also frees up the newest hardware — including Microsoft's GB200 cluster, already up and running — for training rather than serving traffic.

Microsoft chief executive Satya Nadella framed the announcements in a lengthy post on X titled «Frontier Diffusion & Control», which reads almost as a strategic manifesto. «We can now take saturated frontier capabilities and deliver them at scale and at lower cost with models optimized for high-usage products, while continuing to use frontier models for frontier needs», Nadella wrote, adding that Microsoft is «beginning to route traffic through our own surfaces to MAI whenever our models match or beat the frontier alternatives». Translated out of executive prose: capabilities that were state of the art a year ago are now table stakes, and Microsoft believes it can replicate them cheaply for the specific, repetitive tasks that dominate real product usage. Why pay frontier prices for a frontier model when a user only wants to reformat a spreadsheet column?

Nadella was careful to qualify that «OpenAI's and Anthropic's frontier models are part of the orchestration system alongside MAI», but he also articulated a sharp principle of model independence, arguing that the company's evaluations «must keep scaling even when any given model is retired». Keeping the harness, memory, context and skills outside the model, he argued, is what gives Microsoft control. The backdrop is hard to miss: Reuters reported in April that Microsoft's exclusive license to OpenAI's technology had been revised into a non-exclusive agreement, and The Information reported in September that Microsoft had begun incorporating Anthropic models into some products. Wednesday's announcement completes the triangle: Microsoft as orchestrator, with its partners' frontier models as interchangeable components and its own models absorbing an ever-larger share of routine traffic.

The reaction on social media reflected both the appeal of and the skepticism about the strategy. «I love it when people use small models for specific tasks», wrote X user @mavihsk, replying to Nadella's post. «Why do I have to use the know-it-all model just to change a field in Excel?». Another user, @nabu_lines, summed up the argument precisely: «cost and performance improve when you stop overusing the biggest model». Others were less forgiving about Microsoft's execution record: designer @designedbyabin wrote that «Microsoft is the worst at listening to user feedback», arguing that the company «will lose the AI race because it repeatedly fails to understand user needs». And one user, @tokenoverflow, offered a drier critique of the model-independence argument: «I want it to keep scaling after Microsoft is retired».

The skepticism has grounds: the metrics Microsoft reports — acceptance rates, save rates, GPU savings — come from its own internal evaluations, not from independent benchmarks, and it is the company itself that chooses which comparisons to publish. But the logic of the strategy does not hinge on any single figure. Nadella's point that software now has «a real marginal cost for the first time» explains why Microsoft is obsessed with tokens, GPUs and serving costs: when AI features run on every keystroke across a product portfolio with a billion users, an 84% cut in GPU costs is not a mere optimization, it is the difference between a viable business and a bottomless pit.

The final piece of the strategy is selling the method itself, not just the models. Nadella explicitly positioned the «hill-climbing» approach as «a template for any other AI-native, SaaS or enterprise company», and Microsoft is packaging that toolchain through Foundry and what it calls Frontier Tuning, which lets companies train specialized models against their own evaluations and reinforcement learning environments. That turns an internal cost-cutting exercise into an Azure product, and gives enterprise customers a reason to run their AI workloads on Microsoft's cloud even if the models come from somewhere else. The company's emphasis that its models are trained «on clean, traceable, enterprise-grade data, with no distillation from third-party models» serves the same commercial end: in an industry under growing scrutiny over the provenance of training data, Microsoft is betting that enterprise buyers — and courts — will care where those capabilities came from.

Microsoft says it is now extending the hill-climbing approach to Copilot Chat, Outlook and PowerPoint, and both new models are available in public preview through Microsoft Foundry and MAI Playground. «None of this is an endpoint», the company wrote. «We're just getting started». Seven years ago, Microsoft bet more than $13 billion that OpenAI would build the future of AI. Wednesday's announcement suggests the company has since learned a cheaper lesson: the future of AI may belong to whoever builds the frontier, but the profits belong to whoever makes it ordinary.

🔗 Related on Zendoric

Sources & references