16 Sources
[1]
Microsoft takes on AI rivals with three new foundational models | TechCrunch
Microsoft AI, the tech giant's research lab, announced the release of three foundational AI models on Thursday that can generate text, voice, and images. The release signals Microsoft's continued push to build out its own stack of multimodal AI models -- and compete with rival AI labs -- even
[2]
Microsoft's New AI Models Go Beyond Just Text
Microsoft is doubling down on AI models that aren't large language models. The company announced on Thursday that it is releasing three new models: brand new models for voice and text transcription and the second generation of its in-house image model. The voice and text transcription models are
[3]
Microsoft shivs OpenAI with new AI models for speech, images
Microsoft on Thursday unveiled public preview versions of three home-baked machine learning models focused on speech recognition, speech synthesis, and image generation. The release makes the Windows biz look more like a direct competitor to OpenAI than an investor - Redmond held an OpenAI stake
[4]
Microsoft launches three in-house AI models in direct challenge to OpenAI
Six months after renegotiating the contract that once barred it from independently pursuing frontier AI, Microsoft has released three in-house models that directly challenge the partner it spent $13 billion cultivating. MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2 are now available in Microsoft
[5]
Microsoft releases new AI models to expand further beyond OpenAI
Microsoft is expanding its roster of in-house AI models, releasing a new speech-to-text system and making two existing models broadly available to developers for the first time. The moves by Microsoft AI (MAI) are part of a broader effort by the company to expand its proprietary AI capabilities
[6]
Microsoft takes on Google and OpenAI with its own AI models
Microsoft just shipped its own AI models, and they're coming for OpenAI and Google. The company has publicly released three proprietary models: MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2. The models are available via the Microsoft Foundry platform and the MAI Playground. So, what can
[7]
Microsoft releases foundational AI models targeting enterprises
Microsoft wants to offer the 'most complete AI and app agent factory'. Microsoft has released three new AI foundational models, created in-house, in a move that places the company in direct competition with enterprise AI rivals, despite its deep ties with OpenAI. The new foundational models
[8]
Microsoft launches new high-speed voice and image models - SiliconANGLE
Microsoft Corp. today introduced a trio of artificial intelligence models optimized to process images and audio. The algorithms are available through Microsoft Foundry, an Azure service that developers can use to build AI applications. The tech giant has also started rolling out the models to a
[9]
Microsoft launches MAI-Transcribe-1 for speech recognition in 25 languages
Microsoft has launched "MAI-Transcribe-1," an AI model designed for accurate speech-to-text transcription across 25 widely spoken languages. This model is intended for applications such as meetings, closed captioning, and dictation. MAI-Transcribe-1 will be available on Microsoft Foundry alongside
[10]
Microsoft's Three New AI Models Said to Rival OpenAI and Google
Voice-1 can generate realistic speech with an emotional range Microsoft released three specialised artificial intelligence (AI) models on Thursday, focusing on image generation, voice generation, and speech-to-text transcription. The Redmond-based tech giant claims that these models outperform
[11]
Microsoft launches 3 AI models for transcription, image, and speech generation - The Economic Times
Microsoft on Thursday announced three new models from its Microsoft AI (MAI) model family for transcription, image, and speech generation. This includes MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2, as Microsoft aims to expand its push into multimodal artificial intelligence (AI) capabilities
[12]
Microsoft Enters Next AI Phase with Three New Foundational Models
The newly introduced models include MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2, each focused on different aspects of multimodal AI capabilities. MAI-Transcribe-1 is designed for speech-to-text conversion across multiple languages and is claimed to be significantly faster than Microsoft's
[13]
Microsoft rolls out MAI-Transcribe-1, MAI-Voice-1 and MAI-Image-2 in Foundry public preview
Microsoft has announced three new AI models -- MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2 -- in public preview through its AI development platform Microsoft Foundry. The update is part of Microsoft's broader approach to building a unified AI and application agent platform that provides
[14]
Microsoft Seeks Self Sufficiency in AI Race, Launches Three Foundational Models
The three models aim to transcribe audio, have spoken conversations and create images and offer cheaper options to users compared to OpenAI and Google Microsoft cocked-a-snook at its big tech and startup AI rivals by releasing three new foundational models that it has trained internally, thus
[15]
Microsoft Releases AI Models for Transcription, Voice and Image Generation
Microsoft unveiled three new artificial intelligence models offering speech-to-text transcription as well as voice and image generation. The software giant said Thursday it's working to deploy the models to power its consumer and commercial products, and they are now available for its Foundry
[16]
Microsoft unveils three new AI models for speech and imaging: What they can do
The new models include MAI-Transcribe-1, MAI-Voice-1 and MAI-Image-2. Microsoft has introduced three new AI models that can generate text, voice and images. The new models- MAI-Transcribe-1, MAI-Voice-1 and MAI-Image-2- are now available through Microsoft Foundry and MAI Playground. The tech giant
Share
Copy Link
Microsoft unveiled three foundational AI models—MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2—marking its first major independent release since renegotiating its OpenAI partnership. The models handle speech-to-text transcription, voice generation, and image creation, positioning Microsoft as a direct competitor to its $13 billion investment partner while expanding its proprietary capabilities in the crowded AI market.
Microsoft has released three foundational AI models that generate text, voice, and images, signaling a strategic shift toward independence from its longstanding OpenAI partnership
1
. The announcement marks the first publicly released output from the MAI Superintelligence team, formed in November 2025 under Mustafa Suleyman, CEO of Microsoft AI4
. Six months after renegotiating a contract that previously barred independent frontier AI development, Microsoft now competes directly with the partner it spent $13 billion cultivating4
.
Source: CXOToday
The three models—MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2—are available through Microsoft Foundry and MAI Playground, with future plans to integrate MAI-Image-2 into Bing and PowerPoint . These in-house AI models do not carry OpenAI's name anywhere on the label, representing a clear departure from Microsoft's previous reliance on its partner's technology
4
.
Source: Gadgets 360
The speech-to-text model claims the lowest word error rate across 25 languages on the FLEURS benchmark, averaging 3.8 percent
4
. Microsoft reports that MAI-Transcribe-1 outperforms OpenAI's Whisper-large-v3 on all 25 languages, Google's Gemini 3.1 Flash on 22 of 25, and ElevenLabs' Scribe v2 on 15 of 254
. The model runs 2.5 times faster than Microsoft's Azure Fast offering and operates at approximately 50 percent lower GPU cost than leading alternatives3
.Pricing starts at $0.36 per hour of audio, positioning it competitively in the LLM market
1
. Suleyman told The Verge that the transcription model runs at "half the GPU cost of the other state-of-the-art models" and was built by a team of just 10 people5
. The model handles noisy real-world conditions such as call centers and conference rooms, with Microsoft testing integrations with Copilot and Teams5
.MAI-Voice-1 generates 60 seconds of natural-sounding audio in under one second on a single GPU and supports custom voice creation from a few seconds of sample audio
4
. The text-to-speech model is priced at $22 per 1 million characters1
. Combined with MAI-Transcribe-1 and a large language model of the customer's choosing, it forms a complete voice pipeline that runs entirely on Microsoft infrastructure without any dependency on OpenAI's technology4
.MAI-Image-2, originally released on MAI Playground on March 19, debuted at number three on the Arena.ai text-to-image leaderboard, placing behind only Google's Gemini 3.1 Flash and OpenAI's GPT Image 1.5
4
. The model was developed in collaboration with photographers, designers, and visual storytellers, with WPP, one of the world's largest marketing groups, among the first enterprise partners building with it at scale4
. Pricing starts at $5 for 1 million tokens for text input and $33 for 1 million tokens for image output1
.Related Stories
Until the September 2025 renegotiation, Microsoft's original partnership agreement with OpenAI contractually prevented the company from independently pursuing general AI development
4
. The revised memorandum of understanding changed that calculus fundamentally—Microsoft retained licensing rights to everything OpenAI builds through 2032, gained $250 billion in new Azure cloud business commitments, and crucially won the freedom to build competing models4
.
Source: GeekWire
Suleyman acknowledged the pivot directly, stating that the contract renegotiation enabled Microsoft to independently pursue its own superintelligence
4
. In a March internal memo first reported by Business Insider, Suleyman wrote that he intended to focus all of his energy on superintelligence and deliver world-class models for Microsoft over the next five years4
. He told VentureBeat that Microsoft plans to eventually build a frontier large language model to be "completely independent" if needed5
.In an increasingly crowded market, Microsoft hopes a selling point for these models is that they are cheaper than those from Google and OpenAI
1
. "At Microsoft AI, we're building Humanist AI. We have a distinct view when creating our AI models—putting humans at the center, optimizing for how people actually communicate, training for practical use," Suleyman wrote in a blog post1
.Naomi Moneypenny, who leads the Microsoft Azure AI Foundry Models product team, noted that "these are the same models already powering our own products such as Copilot, Bing, PowerPoint, and Azure Speech"
3
. Copilot's Audio Expressions runs on MAI-Voice-1 while Copilot's Voice Mode transcription service uses MAI-Transcribe-13
.Microsoft Foundry, the platform formerly known as Azure AI Foundry and before that Azure AI Studio, now serves developers at more than 80,000 enterprises including 80 percent of Fortune 500 companies
4
. That distribution advantage makes the MAI model family strategically significant—Microsoft does not need to beat OpenAI on every benchmark to shift enterprise spending4
. The company also recently hired former Allen Institute for AI CEO Ali Farhadi and other top AI researchers to further bolster Suleyman's team5
.Summarized by
Navi
[3]
29 Aug 2025•Technology

15 Apr 2026•Technology
14 Oct 2025•Technology

1
Technology

2
Technology

3
Policy and Regulation
