Microsoft Builds Its Own AI Models to Challenge OpenAI and Google

Quick Reads
- Microsoft has launched three new AI models, one for speech transcription, one for voice generation, and one for image creation, developed entirely in-house.
- The models were built by Microsoft’s MAI Superintelligence team, led by Mustafa Suleyman, CEO of Microsoft AI, a unit formed in November 2025.
- All three models are now available to developers on Microsoft Foundry, the company’s AI development platform, and are priced to undercut rival offerings from OpenAI and Google.
- MAI-Transcribe-1 can convert spoken language into text across 25 languages and runs at two and a half times the speed of Microsoft’s own previous transcription service.
- The move signals Microsoft is no longer content to rely solely on its $13 billion partnership with OpenAI; it is now building and competing with its own models.
The Three Models
The three models all carry the MAI label, shorthand for Microsoft AI, and each targets a specific capability that businesses and developers commonly need.
MAI-Transcribe-1 is a speech-to-text model that converts spoken audio into written text across the 25 most widely used languages globally. Think of it as a highly accurate, high-speed captioning engine. According to Microsoft’s own press release, it runs at two and a half times the speed of the company’s existing Azure Fast transcription service, the tool developers have relied on until now. In benchmark testing across the top 25 languages, MAI-Transcribe-1 ranked first in 11 core languages and outperformed OpenAI’s Whisper model across all remaining 14.
MAI-Voice-1 is a voice generation model designed to produce natural, realistic speech that preserves a speaker’s identity even across long-form content.What makes it notable is both its speed and a new personalisation feature: the model can generate 60 seconds of audio in a single second, and developers can now create a custom voice clone using just a few seconds of audio. That means a business could, in theory, build a customer service voice agent that sounds like a specific person with that person’s consent built into the platform’s guardrails, Microsoft says.
MAI-Image-2 is an image generation model built specifically with photographers, designers, and visual storytellers in mind, optimised for natural lighting, accurate skin tones, and clear in-image text for graphics and layouts. It had a soft launch on Microsoft’s MAI Playground in March before this wider release. WPP, one of the world’s largest advertising and communications groups, is among the first enterprise partners already building with MAI-Image-2 at scale.
On price, the part Microsoft is leaning into most is MAI-Transcribe-1, which starts at $0.36 per hour; MAI-Voice-1 starts at $22 per one million characters, and MAI-Image-2 starts at $5 per one million tokens for text input and $33 per one million tokens for image output.Microsoft claims this puts it at the best price-to-performance ratio among major cloud providers, a dig at Google and OpenAI without naming them outright.
The Bigger Picture
All three models were developed by Microsoft’s MAI Superintelligence team, an AI research unit led by Mustafa Suleyman, CEO of Microsoft AI, that was formed and announced in November 2025. In the official blog post announcing the launch, Suleyman described the philosophy behind the models: “At Microsoft AI, we’re building Humanist AI. We have a distinct view when creating our AI models, putting humans at the centre, optimising for how people actually communicate, and training for practical use.“
Despite launching its own AI models, Suleyman used a VentureBeat interview to reaffirm Microsoft’s commitment to its OpenAI partnership, though that partnership was renegotiated recently, reportedly giving Microsoft the space to pursue this kind of independent superintelligence research. Microsoft has invested more than $13 billion into OpenAI and continues to host its models across its products through a multi-year deal. The company is clearly running a dual strategy: partner with the best external labs while quietly building its own stack in parallel. It does the same with chips bought from Nvidia and AMD while developing its own silicon.





