Microsoft picked a crowded week to start a price war in voice AI. On July 23, 2026, its consumer AI division shipped two models into public preview: MAI-Image-2.5-Pro, its highest-fidelity image generator to date, and the one this article cares about, MAI-Voice-2-Flash, a text-to-speech model priced at $15 per million characters. Mustafa Suleyman announced it the same day, noting the model is twice as fast as MAI-Voice-2, 32 percent cheaper, and already powering Dynamics 365 Contact Center, where Microsoft says it cuts GPU serving costs by up to 89 percent compared with routing the same traffic through OpenAI models.
That number, $15 per million characters, is the story. The reference point for premium AI voice is ElevenLabs, whose Creator plan gives you 100,000 credits for $22 a month; on its flagship Multilingual v2 model, where one character costs one credit, that works out to roughly $220 per million characters. Microsoft just published a price around fourteen times lower and attached it to a model it trusts in its own enterprise call-center product. If you generate speech at any real volume, this launch deserves twenty minutes of your attention this week.
Should you switch from ElevenLabs to MAI-Voice-2-Flash? Switch if you generate high volumes of functional speech: voice agents, support lines, e-learning, app narration. At $15 per million characters, Flash costs 7-15x less than ElevenLabs’ effective rates. Stay with ElevenLabs for voice cloning, emotional range, and its voice library, which Microsoft has not matched.
How we assess: research/news-based analysis, not hands-on testing.

Key takeaways
- $15 per 1M characters is MAI-Voice-2-Flash’s public-preview price, announced July 23, 2026 — against an effective ~$220 per 1M on ElevenLabs’ Creator plan (Multilingual v2).
- 2x faster, 32% cheaper than MAI-Voice-2, according to Microsoft, with the same prosody and acoustic quality.
- Up to 89% GPU cost reduction is Microsoft’s claim for Dynamics 365 Contact Center running on Flash instead of OpenAI models.
- 2 places to use it today: Microsoft Foundry (public preview, MAI Playground) and OpenRouter pay-as-you-go.
- 3 things ElevenLabs still does better: professional voice cloning, emotional delivery, and a library of thousands of ready voices.
- A 10M-character/month workload costs about $150 on Flash versus roughly $2,000 on ElevenLabs’ effective per-character rates.
What exactly did Microsoft launch on July 23?
Microsoft launched two in-house models into public preview: MAI-Image-2.5-Pro and MAI-Voice-2-Flash, both accessible through Microsoft Foundry and the MAI Playground. The voice model is the volume play. Per Microsoft’s announcement, Flash keeps “the natural prosody and high acoustic quality found in MAI-Voice-2” while generating speech twice as fast at 32 percent lower cost, and it is priced at a flat $15 per million characters. It is not a lab demo: it powers Azure Voice Live for building voice agents and has already replaced heavier models inside Dynamics 365 Contact Center, Microsoft’s enterprise call-center platform.
The image sibling matters for context because it shows the strategy. MAI-Image-2.5-Pro now runs Bing Image Creator fully in-house by default, and Microsoft says moving PowerPoint’s image generation to it cut GPU costs by up to 84 percent. VentureBeat’s read on the launch was blunt: Microsoft built these models to stop paying OpenAI rates for everyday workloads, with claimed savings up to 89 percent. When a hyperscaler productizes that cost advantage and hands it to developers at $15 per million characters, every per-character vendor in the market feels it, and the loudest incumbent in that market is ElevenLabs. We reviewed ElevenLabs in depth earlier this year in our ElevenLabs 2026 review, where the verdict hinged on exactly the variable Microsoft just attacked: price per usable minute of audio.
How does the pricing actually compare, character for character?

The math favors Microsoft by an order of magnitude for raw generation. ElevenLabs sells subscriptions with credit allowances; on the Multilingual v2 model, one character consumes one credit, and its Flash/Turbo models consume about half a credit per character. Convert the plans to per-million-character rates and the picture is stark:
| Provider / plan | Monthly price | Included characters | Effective cost per 1M chars | Notes |
|---|---|---|---|---|
| MAI-Voice-2-Flash | none (usage-based) | pay-as-you-go | $15 | Public preview, Foundry + OpenRouter |
| MAI-Voice-2 | none (usage-based) | pay-as-you-go | ~$22 | Flash is 32% cheaper per Microsoft |
| ElevenLabs Starter | $5 | 30,000 | ~$167 | Multilingual v2, 1 credit/char |
| ElevenLabs Creator | $22 | 100,000 | ~$220 | Pro voice cloning unlocks here |
| ElevenLabs Pro | $99 | 500,000 | ~$198 | High-volume creator tier |
| ElevenLabs Scale | $330 | 2,000,000 | ~$165 | Teams and agencies |
Even granting ElevenLabs its cheaper Flash v2.5 model at half a credit per character, the Creator plan still lands near $110 per million characters, more than seven times Microsoft’s rate. In US dollar terms, a 10-million-character monthly workload, roughly 200 hours of finished audio, costs about $150 on MAI-Voice-2-Flash versus around $1,100-$2,200 on ElevenLabs depending on model and tier (that is roughly £120 versus £880-£1,760, or AU$225 versus AU$1,650-$3,300). Overages and annual discounts move the numbers at the margins; they do not close a 7-15x gap. One caveat cuts the other way: ElevenLabs subscriptions bundle a product, not just an API, including its studio editor, dubbing tools, and commercial licensing terms baked into paid tiers.
Is the quality gap real, or is cheap finally good enough?
The honest answer in week one: Microsoft is not claiming parity with ElevenLabs’ best, and nobody serious should either. Microsoft’s own framing positions Flash for “high-volume voice experiences prioritizing responsiveness,” which is corporate for: it reads fast, clean, and natural, and it is not trying to win an Oscar. ElevenLabs’ Multilingual v2 remains the reference for emotional contour, and its ecosystem of thousands of community and professionally cloned voices has no MAI equivalent. Independent, benchmark-style comparisons of Flash against ElevenLabs do not exist yet; the model is a week old, so treat any confident quality verdict you read this month, including ours, as provisional.
What is verifiable is where Microsoft dares to run it: live customer calls in Dynamics 365 Contact Center. Call-center speech is unforgiving in one specific way, latency, and forgiving in another, theatrical range. That deployment tells you Flash’s floor is high enough for paying enterprise customers to hear all day, and its 2x speed claim tells you where the engineering effort went. The pattern rhymes with what happened in transcription, where cheap fast models ate the commodity middle while premium held the top, a dynamic we covered in Atlas-1 vs Deepgram. Voice generation is now running the same play, and the same week brought Qwen-Audio-3.0-TTS at roughly a third of ElevenLabs’ price, so the squeeze is coming from two directions at once.
Who should switch, and who should stay?
Switch-or-pilot candidates are teams whose speech is functional and voluminous. If you’re running IVR or support voice agents, the Dynamics deployment is your proof point; pilot Flash through Azure Voice Live this quarter. If you’re generating e-learning modules, product walkthroughs, or accessibility read-aloud by the hour, your cost per finished hour drops from roughly $11 (ElevenLabs Creator effective rate) to about 75 cents. If you’re a developer adding voice output to an app, $15 per million characters makes voice a rounding error instead of a line item, and OpenRouter gives you a no-Azure on-ramp. If you’re localizing at scale, pair Flash for the bulk read with a premium model for the hero content and cut the blended rate by 80 percent or more.
Stay-with-ElevenLabs candidates are creators whose voice is the product. If you’re producing audiobooks, character-driven YouTube content, or ad reads that need to land an emotion, ElevenLabs’ expressive ceiling and cloning workflow remain worth its premium, the same calculus that keeps Suno v5 subscribers paying for commercial-grade music and Runway Gen-4 users paying for controllable video. If your brand voice is a professional clone of a real person, there is no MAI migration path today. And if you need contractual commercial-use clarity in a mature product with an editor, dubbing, and versioned voices, a public-preview API is not yet your home. The clean mental model: MAI-Voice-2-Flash is infrastructure pricing for speech; ElevenLabs is a creative studio with an API attached.
Why is Microsoft doing this, and why now?
Because every character MAI speaks is a character OpenAI does not bill. Microsoft’s MAI program has spent 2026 systematically replacing OpenAI inference inside its own products: Bing Image Creator now defaults to in-house generation, PowerPoint’s imagery runs up to 84 percent cheaper, OneDrive’s photo features got faster and stickier, and Dynamics 365’s voice stack now runs on a model Microsoft owns end to end. VentureBeat framed the July 23 launch as Microsoft cutting costs “up to 89% versus OpenAI,” and that is the through-line: the models are cost weapons first, products second. Shipping them into public preview at aggressive prices does double duty, stress-testing them on external traffic while undercutting rivals who must make money on inference alone.
The timing also tracks the market’s brutal July. Seven-plus models shipped in the week of July 17-23 alone, including Ant Group’s Ling-3.0-flash free through August 3 and Qwen’s cut-rate audio stack; Anthropic answered with Claude Opus 5 on July 24 at half its predecessor’s price. Efficiency, not capability, is the axis of competition this summer, and voice is simply the loudest place it shows. For buyers, the strategy question writes itself: when a platform vendor sells the commodity tier at near-cost, paying premium rates makes sense only for work the commodity tier demonstrably cannot do. That test now applies to every ElevenLabs invoice, and it applies fresh every quarter as the previews mature; our NotebookLM review made a similar point about Google bundling “good enough” AI into tools people already pay for.
How do you actually try MAI-Voice-2-Flash this week?
The fastest route is OpenRouter: both MAI-Voice-2-Flash and MAI-Voice-2 are listed for pay-as-you-go access with a standard API key, no Azure account required. The official route is Microsoft Foundry, where the model sits in public preview alongside the MAI Playground for browser testing, and Azure Voice Live wraps it in a real-time agent pipeline with interruption handling and telephony hooks. Budget an afternoon: generate the same 90-second script you last produced in ElevenLabs, once neutral, once conversational, and A/B them with a colleague who does not know which is which.
Structure the pilot around three measurements. Latency to first audio byte, because that is Flash’s headline advantage and it decides whether a voice agent feels alive. Pronunciation of your domain vocabulary, product names, acronyms, and mixed-language phrases, since preview models stumble on precisely the words your users notice. And listener preference at your actual use case, not in the abstract; a support-line greeting and an audiobook chapter are different sports. Keep your ElevenLabs subscription live during the pilot. The goal in August is not to switch; it is to know your number, the per-million-character price at which you would, before your next renewal makes the decision for you.
What are the risks of building on a public-preview model?
The risks are real and they are mostly contractual, not technical. Public preview means Microsoft reserves the right to change pricing, rate limits, regional availability, and even model behavior before general availability, and preview services typically carry reduced or no SLA. If your product speaks to customers every minute of the day, a silent model revision that shifts pronunciation or pacing is a support ticket generator, and a price change erases the spreadsheet you built your migration case on. History says hyperscaler preview prices more often fall than rise, but “more often” is not a guarantee your finance team can book.
There are also softer risks worth naming. Voice selection in a young model catalog is thinner than ElevenLabs’ library, so brands with a settled audio identity may not find a close match on day one. Regional voice coverage and accent quality vary before GA, which matters if your audience spans US, UK, and Australian English and expects each to sound native. And OpenRouter access, convenient as it is, adds a middleman to your compliance story; regulated industries will want the Azure route with its enterprise agreements. None of this argues against piloting now. It argues against deleting your fallback. The teams that win pricing transitions run dual providers for a quarter, route ten percent of traffic to the challenger, and move the dial only when the error logs stay boring.
How does Flash fit into Microsoft’s wider MAI lineup?
Flash is the fourth visible piece of a deliberate stack. Microsoft’s MAI family now spans text, image, voice, and speech recognition inside Microsoft Foundry, and each release has followed the same recipe: build in-house, deploy inside a Microsoft product at scale, publish the cost savings, then open a public preview. MAI-Image-2.5-Pro proved it on Bing Image Creator and PowerPoint. MAI-Voice-2 established the quality bar, and Flash now productizes the speed-and-price variant, with Microsoft’s Community Hub post describing the pair as covering both the fidelity and throughput ends of voice workloads. The pattern to watch: previews that survive Dynamics-scale traffic tend to reach general availability within months, usually with enterprise SLAs attached.
For buyers, the strategic read is that voice is becoming a bundled capability rather than a standalone product category. When speech synthesis ships as a near-free ingredient inside the cloud you already rent, the standalone vendors must justify their margin with things a platform will not build: cloning consent workflows, creative direction tools, community voice marketplaces, per-project licensing. ElevenLabs has those today, which is why the sane 2026 posture is a split stack rather than a loyalty decision. Let Microsoft carry the bulk minutes, let the specialist carry the brand voice, and re-run the math each quarter as previews graduate and prices settle. The vendors are competing on your renewal calendar now; use it.
Which option fits your exact situation?
If you run a support or sales voice line, pilot MAI-Voice-2-Flash through Azure Voice Live now; latency and cost both move in your favor, and Microsoft is running the same model on its own call traffic. If you produce e-learning or corporate training by the hour, switch the bulk narration first and keep a premium voice for course intros. If you are an indie developer, start on OpenRouter with a $10 credit and you will struggle to spend it. If you narrate audiobooks or run a cloned brand voice, stay on ElevenLabs and let this launch pressure your renewal price instead. If you are an agency billing clients for voice production, split the stack: commodity reads on Flash, hero reads on ElevenLabs, margin on both. And if you are all-in on Google or AWS rather than Azure, watch Qwen’s audio line and the incumbents’ inevitable response before signing anything annual; this market is repricing monthly right now.
MAI-Voice-2-Flash vs ElevenLabs: FAQ
What is MAI-Voice-2-Flash and when did it launch?
MAI-Voice-2-Flash is Microsoft’s newest text-to-speech model, announced July 23, 2026 alongside MAI-Image-2.5-Pro. It generates speech twice as fast as MAI-Voice-2 at a 32% lower price, $15 per million characters, while keeping the same prosody and acoustic quality. It is in public preview in Microsoft Foundry and already powers Dynamics 365 Contact Center and Azure Voice Live.
How much cheaper is MAI-Voice-2-Flash than ElevenLabs?
Roughly 7 to 15 times cheaper per character, depending on the ElevenLabs plan and model. ElevenLabs’ Creator plan works out to about $220 per million characters on Multilingual v2, and its half-rate Flash model still lands near $110. MAI-Voice-2-Flash charges a flat $15 per million characters with no monthly subscription attached.
Is MAI-Voice-2-Flash better quality than ElevenLabs?
No, not for expressive work. Microsoft positions Flash for high-volume, latency-sensitive workloads like call centers and voice agents, and says it preserves MAI-Voice-2’s natural prosody. ElevenLabs still leads on emotional range, voice cloning fidelity, and its library of thousands of community voices. For narration-grade drama, ElevenLabs keeps the crown; for bulk speech, the gap narrows sharply.
Does MAI-Voice-2-Flash support voice cloning?
Microsoft has not made professional voice cloning a headline feature of the Flash preview, and this is ElevenLabs’ clearest remaining moat. ElevenLabs offers instant cloning from Starter and professional cloning from its $22 Creator tier. If your workflow depends on a cloned brand voice or your own voice, ElevenLabs remains the practical choice today.
Where can I try MAI-Voice-2-Flash?
Two official routes: Microsoft Foundry (with the MAI Playground) for API access in public preview, and OpenRouter, which lists both MAI-Voice-2 and MAI-Voice-2-Flash for pay-as-you-go use. Dynamics 365 Contact Center customers are already hearing it in production without doing anything.
Who should switch from ElevenLabs to MAI-Voice-2-Flash?
Teams generating large volumes of functional speech: IVR and support lines, voice agents, e-learning at scale, accessibility read-aloud, and app narration where cost per minute rules. A workload of 10 million characters a month drops from roughly $2,000 on ElevenLabs’ effective rates to about $150. Creators making character-driven audio, audiobooks with cloned voices, or emotional ad reads should stay put.
Does this launch threaten ElevenLabs’ business?
It pressures the commodity end. Microsoft says the model cuts its own GPU serving costs by up to 89% in Dynamics 365, and Qwen’s audio models launched the same week at roughly a third of ElevenLabs’ pricing, so the floor is falling fast. ElevenLabs’ defense is quality, cloning, and its creator ecosystem, which is why its $99 Pro tier still sells.
Is MAI-Voice-2-Flash production-ready right now?
Treat it as preview software. Public preview means Microsoft can change pricing, limits, and behavior before general availability, and SLAs are thinner than GA services. Microsoft is dogfooding it in Dynamics 365 Contact Center, which is a strong signal, but risk-averse teams should pilot it alongside their current provider rather than cut over in a weekend.
Verdict: the $15 number changes the default
MAI-Voice-2-Flash does not dethrone ElevenLabs on quality, and Microsoft is not pretending it does. What it changes is the default question. Until July 23, teams asked “which voice AI is best?” and ElevenLabs usually won. Now the question is “what part of my speech workload actually needs $220-per-million voice when $15-per-million exists from a vendor running it in production call centers?” For most functional speech, the answer will embarrass a lot of invoices. Pilot it now, keep the premium tier for the work that earns it, and reread your renewal terms before autumn. Price moves this large do not stay secret from your CFO.
Sources
- Microsoft AI — Introducing MAI-Image-2.5-Pro and MAI-Voice-2-Flash · Verified July 30, 2026
- VentureBeat — Microsoft launches in-house AI models it says cut costs up to 89% vs OpenAI · Verified July 30, 2026
- OpenRouter — MAI-Voice-2-Flash pricing and providers · Verified July 30, 2026
- Microsoft Community Hub — New MAI models in Microsoft Foundry · Verified July 30, 2026
- BIGVU — ElevenLabs pricing 2026 plan breakdown · Verified July 30, 2026
- Digital Applied — Seven days, seven model releases (July 2026) · Verified July 30, 2026