Blog Post

Microsoft Foundry Blog
3 MIN READ

MAI-Image-2.6 and MAI-Image-2.6-Flash: Quality and speed at production scale

Naomi Moneypenny's avatar
Sep 04, 2026

Image generation has moved beyond the demo stage. Marketing teams are shipping campaign assets, retailers are building catalogs, and product teams are generating visuals inside the tools they already use. That shift changes what matters. A model must do more than create a striking one-off image. It needs to render legible text, maintain consistency across formats, and follow creative direction precisely.

Today, we’re introducing MAI-Image-2.6 and MAI-Image-2.6-Flash in Microsoft Foundry, the newest image generation models in the Microsoft AI family.

MAI-Image-2.6 delivers a substantial quality improvement over MAI-Image-2.5, along with multi-reference editing, web grounding, and greater control over format and resolution. MAI-Image-2.6-Flash brings these advances to speed- and efficiency-sensitive production workloads. It is more than 2x faster GPT-Image-2-Medium and 78% more efficient than GPT-Image-2. 

Another step up in quality

MAI-Image-2.6 arrives with strong third-party benchmark results across both Arena leaderboards.

  • No. 2 on the Arena text-to-image leaderboard at launch.
  • No. 3 on Arena's image-editing leaderboard, ahead of Google's Nano Banana family and Meta's Muse Image.
Bar chart of MAI-Image-2 amongst other text-to-image models on Arena.ai rankings

What's new in MAI-Image-2.6

  • Multi-reference editing. MAI-Image-2.6 accepts up to five reference images in a single request, so a locked creative — a product, a character, a brand mark — stays consistent as you revise and resize it. Combined with adaptive aspect ratios and 1.5K resolution output, one approved concept can travel across channel formats without drifting.
  • Web grounding. The model can pull in real-world context rather than relying only on what it learned in training, producing imagery that reflects accurate details of places, objects, and subjects.
  • Greater control over format and resolution. Developers get explicit levers over how much the model deliberates before generating, what shape the output takes, and how much detail it resolves — so the same model can serve a fast iteration loop and a final production render.
Multi-reference feature: combine your brief, brand assets, and up to 5 references into one coherent direction

Faster production with MAI-Image-2.6-Flash

MAI-Image-2.6-Flash is optimized for applications that need to generate many variations, respond quickly, or process content at scale.

It more than 2x faster than GPT-Image-2-Medium and 72% more efficient than GPT-Image-2. This makes it well suited to interactive applications, automated pipelines, personalization, and high-volume creative production.

Choosing the right model

Choose MAI-Image-2.6 when precision and the highest output quality are the priority, including final campaign assets, complex edits, detailed commercial design, and high-fidelity text rendering.

Choose MAI-Image-2.6-Flash when latency, throughput, and efficiency matter most, including interactive experiences, rapid iteration, personalization, and high-volume generation.

Teams can also use Flash to explore concepts and variations quickly, then use MAI-Image-2.6 for final assets that require greater precision.

Get started today in Microsoft Foundry

MAI-Image-2.6 and MAI-Image-2.6-Flash are available in public preview in Foundry. MAI-Image-2.6 pricing starts at $5 USD per 1M tokens for text input, $8 USD per 1M tokens for image input, and $38 USD per 1M tokens for image output. MAI-Image-2.6-Flash pricing starts at $1.75 USD per 1M tokens for text input, $2.50 USD per 1M tokens for image input, and $19 USD per 1M tokens for image output.

Updated Sep 04, 2026
Version 1.0