<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>Azure Architecture Blog articles</title>
    <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/bg-p/AzureArchitectureBlog</link>
    <description>Azure Architecture Blog articles</description>
    <pubDate>Mon, 31 Aug 2026 21:44:15 GMT</pubDate>
    <dc:creator>AzureArchitectureBlog</dc:creator>
    <dc:date>2026-08-31T21:44:15Z</dc:date>
    <item>
      <title>Choosing the Right Agent in Microsoft Foundry</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/choosing-the-right-agent-in-microsoft-foundry/ba-p/4547827</link>
      <description>&lt;P&gt;Many discussions about Microsoft Foundry Agent Service eventually arrive at the same question: should this workload be implemented as a Prompt Agent or a Hosted Agent? While the documentation explains both options well, the architectural decision usually comes down to something much simpler: where do you want the orchestration logic to live?&lt;/P&gt;
&lt;H4 aria-level="1"&gt;&lt;STRONG&gt;First, what actually makes something an agent?&lt;/STRONG&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;201341983&amp;quot;:0,&amp;quot;335559738&amp;quot;:480,&amp;quot;335559739&amp;quot;:0,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H4&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;A basic AI assistant generates an answer. An agent can also decide what to do next, call tools, access data, maintain context and complete work across multiple steps.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;At the center of most agents are three building blocks:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:360,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;singleLevel&amp;quot;}" data-aria-posinset="3" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;Model: provides language understanding, generation, and reasoning.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:360,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;singleLevel&amp;quot;}" data-aria-posinset="4" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;Instructions: define the job, boundaries, role, and expected behaviour.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:360,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;singleLevel&amp;quot;}" data-aria-posinset="5" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;Tools: connect the agent to knowledge and actions such as search, APIs, databases, code execution, MCP servers or business systems.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;For enterprise use, that is only the starting point. You also need identity, authorization, network controls, content safety, session management, evaluation, tracing, versioning, rollback, and cost controls. Foundry Agent Service provides the surrounding platform capabilities, while letting you choose how much runtime logic your team owns.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H4&gt;&lt;STRONG&gt;Where should orchestration logic live and who should own the runtime?&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P&gt;&lt;STRONG&gt;Understanding the Runtime Boundary&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;When evaluating Foundry Agent Service, many teams focus on models. In practice, models are rarely the architectural differentiator.&lt;/P&gt;
&lt;P&gt;Most architecture reviews eventually come down to three questions:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Who owns orchestration?&lt;/LI&gt;
&lt;LI&gt;Who owns state?&lt;/LI&gt;
&lt;LI&gt;Who owns operations?&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Foundry Agent Service provides a managed platform for these concerns, but the amount of control retained by engineering teams depends on the selected agent type.&lt;/P&gt;
&lt;P&gt;For most teams, the architectural decision usually comes down to one of two operating models&lt;/P&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:360,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;singleLevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;Prompt agents: declarative agents defined by a model, instructions, and tools, with a managed runtime.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:360,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;singleLevel&amp;quot;}" data-aria-posinset="2" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;Hosted agents: code-based agents that you package and run in Foundry, while the service manages the endpoint, identity, scaling, sessions, and observability.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H4&gt;&lt;STRONG&gt;Prompt Agents&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P&gt;With a Prompt Agent, engineering teams focus primarily on defining the model, instructions, tools, knowledge sources and identity configuration, while Foundry takes responsibility for the surrounding runtime.&lt;/P&gt;
&lt;H6 aria-level="2"&gt;&lt;STRONG&gt;Why teams start here&lt;/STRONG&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;201341983&amp;quot;:0,&amp;quot;335559738&amp;quot;:200,&amp;quot;335559739&amp;quot;:0,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H6&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Prompt agents are usually the fastest route from an idea to a working, governed agent. They are a good fit when the behaviour can be expressed clearly through instructions and supported tools.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:360,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;singleLevel&amp;quot;}" data-aria-posinset="6" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;You need to deliver quickly.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:360,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;singleLevel&amp;quot;}" data-aria-posinset="7" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;The agent follows a fairly straightforward reasoning and tool-use loop.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:360,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;singleLevel&amp;quot;}" data-aria-posinset="8" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;Foundry-supported tools cover the required integrations.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:360,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;singleLevel&amp;quot;}" data-aria-posinset="9" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;You do not need custom libraries, middleware, or orchestration code.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:360,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;singleLevel&amp;quot;}" data-aria-posinset="10" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;You want Foundry to own compute, scaling, and patching.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:360,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;singleLevel&amp;quot;}" data-aria-posinset="11" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;Reviewers need an &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;agent&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt; definition that is easy to inspect.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H6 aria-level="2"&gt;&lt;STRONG&gt;Good examples&amp;nbsp;&lt;/STRONG&gt;&lt;/H6&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Enterprise knowledge assistant. Employees ask about policies, engineering standards, procedures, or product information. The agent retrieves approved content and cites its sources.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Document review assistant. The agent checks a proposal or design against an approved rubric and returns structured findings, while a human keeps responsibility for the final decision.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Employee self-service agent. The agent answers questions and performs a small number of tightly scoped actions, such as checking request status or creating a support case.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;H6 aria-level="2"&gt;&lt;STRONG&gt;A useful warning sign &lt;/STRONG&gt;&lt;SPAN data-contrast="auto"&gt;A Prompt agent is probably becoming the wrong fit when the prompt starts looking like application code. Large branching instructions, retry logic written in prose, state-machine behaviour, custom payload handling, framework middleware or real-time media are all signs that runtime logic belongs in code instead.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H6&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H4&gt;&lt;STRONG&gt;Hosted Agents&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P&gt;Hosted Agents move the responsibility boundary. Instead of defining behaviour through configuration alone, engineers deploy an actual application into Foundry Agent Service. Hosted Agents are framework-agnostic. Whether your team builds with Agent Framework, LangGraph, Semantic Kernel, OpenAI Agents SDK, or a custom runtime, Foundry can host the application while managing the surrounding operational services.&lt;/P&gt;
&lt;H6 aria-level="2"&gt;&lt;STRONG&gt;When Hosted agents make sense&lt;/STRONG&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;201341983&amp;quot;:0,&amp;quot;335559738&amp;quot;:200,&amp;quot;335559739&amp;quot;:0,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H6&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:360,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;singleLevel&amp;quot;}" data-aria-posinset="12" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;You need a particular agent framework or custom orchestration engine.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:360,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;singleLevel&amp;quot;}" data-aria-posinset="13" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;The flow includes branching, parallel work, fan-out and fan-in, or human approvals.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:360,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;singleLevel&amp;quot;}" data-aria-posinset="14" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;Business rules require a deterministic state machine around model reasoning.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:360,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;singleLevel&amp;quot;}" data-aria-posinset="15" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;You need custom packages, middleware, algorithms, retries, caching, or error handling.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:360,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;singleLevel&amp;quot;}" data-aria-posinset="16" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;The client sends custom payloads or webhooks.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:360,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;singleLevel&amp;quot;}" data-aria-posinset="17" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;The session needs persistent files or custom state.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:360,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;singleLevel&amp;quot;}" data-aria-posinset="18" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;The design includes multi-agent orchestration or real-time voice.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H6 aria-level="2"&gt;&lt;STRONG&gt;Good examples&lt;/STRONG&gt;&lt;/H6&gt;
&lt;P&gt;A bank onboarding workflow where uploaded documents must be validated, checked against multiple systems, and routed to a human when confidence drops below a threshold.&lt;/P&gt;
&lt;P&gt;A fraud investigation agent that gathers transaction history, enriches data from multiple internal systems, applies bank-specific risk rules, requests additional evidence when required, and generates a recommended outcome for an investigator. The process involves long-running workflows, branching logic and audit requirements that are better suited to code-based orchestration.&lt;/P&gt;
&lt;P&gt;A lending workflow that coordinates document collection, credit bureau checks, income verification, affordability assessments, policy exceptions, and approval routing. The process spans multiple systems and often requires deterministic decision paths that extend beyond prompt-driven orchestration.&lt;/P&gt;
&lt;P&gt;A security operations agent that aggregates alerts from SIEM platforms, enriches incidents with threat intelligence, executes automated containment actions, opens tickets, requests approvals for high-impact remediation steps, and maintains a complete audit trail of decisions and actions.&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;H6 aria-level="2"&gt;&lt;STRONG&gt;The trade-off&amp;nbsp;&lt;/STRONG&gt;&lt;/H6&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;More control also means more ownership. Your team must secure and patch the code and dependencies, test the runtime, manage supply-chain risk, and think about compute sizing, cold starts, session lifecycle, and cost. Hosted agents reduce platform plumbing, but they do not remove application engineering.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H4&gt;&lt;EM&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:140,&amp;quot;335559740&amp;quot;:269}"&gt;Choosing Prompt Agent/Hosted Agents&amp;nbsp;&lt;/SPAN&gt;&lt;/EM&gt;&lt;/H4&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;STRONG&gt;1)&amp;nbsp; Runtime Control Is Usually the Real Requirement&lt;/STRONG&gt;&lt;/DIV&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;A pattern I see quite often is teams arriving at the solution before they've fully articulated the requirement. The conversation usually starts with "We need a Hosted Agent," but after digging into the workload, the real requirements turn out to be things like persistent state, webhook processing, custom orchestration, background execution, framework-specific capabilities, or human approval workflows.&lt;/P&gt;
&lt;P&gt;These are runtime concerns, not agent concerns and they're usually the factors that determine whether a Hosted Agent is necessary. Hosted Agents are valuable because they give engineering teams control over those aspects of execution while still offloading much of the operational infrastructure to Foundry.&lt;/P&gt;
&lt;P&gt;This is also where teams most commonly choose the wrong agent type. A frequent assumption is that existing investments in frameworks such as LangGraph or Semantic Kernel automatically imply a Hosted Agent architecture. In practice, many of these workloads are relatively simple orchestration scenarios that can be implemented effectively as Prompt Agents, with lower operational overhead and less infrastructure to manage.&lt;/P&gt;
&lt;P&gt;My advice is usually to start by identifying the runtime requirements rather than selecting an agent type. Once those requirements are clear, the right architecture often becomes obvious.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;STRONG&gt;2)&amp;nbsp; When Hosted Agents become mandatory&lt;/STRONG&gt;&lt;/DIV&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;BR /&gt;The moment you need custom Python packages, long-running workflows, external SDKs, deterministic orchestration or framework-specific capabilities, the conversation shifts from Prompt Agents to Hosted Agents.&lt;/DIV&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&amp;nbsp;&lt;/DIV&gt;
&lt;H3 class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;STRONG&gt;What I would choose today&lt;/STRONG&gt;&lt;/H3&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;BR /&gt;
&lt;P&gt;If I were starting a new project today, I'd begin with a Prompt Agent unless there was a clear reason not to. In my experience, Prompt Agents cover far more enterprise use cases than many teams initially expect. The best projects tend to start simple, prove value, learn where the limitations are, and then introduce Hosted Agents only when runtime customization becomes a genuine requirement. That progression is usually far less risky than leading with a fully custom solution.&lt;/P&gt;
&lt;P&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/foundry/agents/quickstarts/prompt-agent?tabs=python" target="_blank"&gt;Quickstart: Create a prompt agent - Microsoft Foundry | Microsoft Learn&lt;/A&gt;&lt;/P&gt;
&lt;/DIV&gt;
&lt;P&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/hosted-agents" target="_blank"&gt;Hosted agents in Foundry Agent Service - Microsoft Foundry | Microsoft Learn&lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Wed, 19 Aug 2026 22:52:13 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/choosing-the-right-agent-in-microsoft-foundry/ba-p/4547827</guid>
      <dc:creator>supriyas</dc:creator>
      <dc:date>2026-08-19T22:52:13Z</dc:date>
    </item>
    <item>
      <title>From Features to Flow: How Real-World Adoption Reshaped the Azure Architecture Diagram Builder</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/from-features-to-flow-how-real-world-adoption-reshaped-the-azure/ba-p/4546817</link>
      <description>&lt;P&gt;In May, I introduced the open-source &lt;A href="https://aka.ms/diagram-builder" target="_blank"&gt;Azure Architecture Diagram Builder&lt;/A&gt; as a way to move from a natural-language prompt to an Azure architecture diagram, cost estimate, Well-Architected assessment, and deployment guidance. In July, I shared how the project had become &lt;A href="https://techcommunity.microsoft.com/blog/azurearchitectureblog/beyond-the-canvas-the-azure-architecture-diagram-builder-becomes-agent-ready/4534590" target="_blank"&gt;agent-ready through Model Context Protocol (MCP)&lt;/A&gt;.&lt;/P&gt;
&lt;P&gt;Those posts described what the tool could do. The more interesting story came next: what happened when people actually used it.&lt;/P&gt;
&lt;P&gt;As adoption grew, the central product question changed. It was no longer simply, &lt;EM&gt;Can AI generate an Azure architecture?&lt;/EM&gt; It became:&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;How do we help an architect choose how to begin, improve a result without losing their work, validate it responsibly, and turn it into something another person can use?&lt;/STRONG&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;That question reshaped the Azure Architecture Diagram Builder from a collection of capabilities into a guided workflow:&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Create → Refine → Validate &amp;amp; Improve → Share or Build&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;This post explains what we learned, what changed in the product, and why the hardest part of AI-assisted architecture is not the first diagram. It is everything that comes after it.&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;TL;DR.&lt;/STRONG&gt; Growing adoption created a feedback loop. Aggregate usage showed that people moved beyond generation into validation, recommendations, exports, and deployment guidance. Privacy-safe feedback revealed recurring problems with diagram integrity, preservation of human edits, cost credibility, export quality, and validation continuity. Those signals led to a four-stage architecture journey that keeps human judgment and professional review at the center. The same lesson now shapes agent access and the next product boundary: distinguish logical proposals from evidence-backed physical architecture.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;HR /&gt;
&lt;H2 id="adoption-created-a-product-feedback-loop"&gt;Adoption created a product feedback loop&lt;/H2&gt;
&lt;P&gt;As of August 5, 2026, the first two Azure Architecture Blog articles had accumulated approximately &lt;STRONG&gt;12,100 combined views&lt;/STRONG&gt;. A refreshed view of deduplicated application telemetry through August 13 recorded:&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table class="lia-border-color-custom-d6dee6 lia-border-style-solid" border="1" style="width: 100%; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th class="lia-border-color-custom-0078d4 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Activity&lt;/th&gt;&lt;th class="lia-border-color-custom-0078d4 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Aggregate count&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Architecture generation and refinement events&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;5,023&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Well-Architected validations&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;960&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Recommendations applied&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;175&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Diagram exports&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;2,020&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Deployment guides generated&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;212&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;As of August 13, the public repository had reached &lt;STRONG&gt;45 stars and 14 forks&lt;/STRONG&gt;. In GitHub’s current rolling 14-day window, the repository recorded &lt;STRONG&gt;277 unique visitors and 67 unique cloners&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P&gt;These numbers measure different things and should not be added together. Article views are not unique readers. Application activity uses anonymous telemetry identifiers, not verified people. GitHub traffic is a rolling aggregate window. The signals are useful because of the pattern they reveal, not because they can be combined into one headline user count.&lt;/P&gt;
&lt;P&gt;Activity also accelerated during the period following the second article. Compared with the May 19–July 9 baseline, daily activity from July 10 through August 13 was approximately &lt;STRONG&gt;7.0 times higher&lt;/STRONG&gt; for architecture generation and refinement, &lt;STRONG&gt;5.6 times higher&lt;/STRONG&gt; for Well-Architected validation, and &lt;STRONG&gt;5.9 times higher&lt;/STRONG&gt; for recommendation application. The timing coincided with publication; it does not prove that the article alone caused the growth.&lt;/P&gt;
&lt;P&gt;The important product lesson was simpler: people were not stopping after the first diagram.&lt;/P&gt;
&lt;P&gt;They were testing alternatives, validating designs, applying recommendations, exporting artifacts, and asking how to move toward implementation. Generation was the entry point, not the complete job.&lt;/P&gt;
&lt;P&gt;The first guided-journey signals reinforce the need for more than one starting path. Through August 13, the new journey instrumentation recorded &lt;STRONG&gt;880 interactions from 174 anonymous identifiers across 241 sessions&lt;/STRONG&gt;. At first start, structured brief/image generation and Guided Chat were selected at almost the same frequency (158 and 156 events), while template and live-Azure import added another 68 selections. These are interaction counts, not unique people or conversion rates, and the window is still too early to claim that the journey improves completion. They are enough to show that architecture work does not begin in one uniform way.&lt;/P&gt;
&lt;FIGURE&gt;&lt;BR /&gt;
&lt;FIGCAPTION aria-hidden="true"&gt;In-product Start Here panel showing the four-stage Azure Architecture Diagram Builder journey: Create, Refine, Validate and Improve, and Share or Build.&lt;/FIGCAPTION&gt;
&lt;/FIGURE&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 1.&lt;/STRONG&gt; The in-product Start Here panel explains one complete architecture loop. The stages are recommendations, not gates, and direct access to every tool remains available.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;HR /&gt;
&lt;H2 id="stage-1-create-make-the-starting-choice-explicit"&gt;Stage 1: Create — make the starting choice explicit&lt;/H2&gt;
&lt;P&gt;As capabilities accumulated, the first screen became harder to interpret. Architecture Chat and structured generation were both useful, but they competed for attention. Importing an existing architecture was available, yet easy to miss.&lt;/P&gt;
&lt;P&gt;The new starting experience makes three paths explicit:&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table class="lia-border-color-custom-d6dee6 lia-border-style-solid" border="1" style="width: 100%; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th class="lia-border-color-custom-0078d4 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Starting path&lt;/th&gt;&lt;th class="lia-border-color-custom-0078d4 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Best suited for&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;STRONG&gt;Guided Chat&lt;/STRONG&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Exploring requirements conversationally and refining them over multiple turns&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;STRONG&gt;Generate Diagram&lt;/STRONG&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Providing a structured brief or image and producing a first architecture quickly&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;STRONG&gt;Import Existing&lt;/STRONG&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Opening an existing architecture or infrastructure artifact for analysis and editing&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;This is not a marketing landing page placed in front of the tool. It is a small decision point inside the authoring experience. Once a path is selected, the user lands on the real canvas.&lt;/P&gt;
&lt;P&gt;The distinction matters because different architecture tasks begin with different levels of certainty. Sometimes the architect knows the target services. Sometimes the problem needs discovery. Sometimes the architecture already exists and the work is to understand or improve it.&lt;/P&gt;
&lt;P&gt;The product should acknowledge those differences instead of pretending every design starts with a perfect prompt.&lt;/P&gt;
&lt;FIGURE&gt;&lt;BR /&gt;
&lt;FIGCAPTION aria-hidden="true"&gt;Start chooser presenting Guided Chat, Generate Diagram, and Import Existing as three equal entry paths.&lt;/FIGCAPTION&gt;
&lt;/FIGURE&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 2.&lt;/STRONG&gt; Three starting paths reflect three different architecture situations: discovery, structured generation, and analysis of an existing design.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;HR /&gt;
&lt;H2 id="stage-2-refine-preserve-human-work"&gt;Stage 2: Refine — preserve human work&lt;/H2&gt;
&lt;P&gt;One of the clearest feedback themes was not about adding another AI capability. It was about preventing AI from casually undoing human effort.&lt;/P&gt;
&lt;P&gt;An architect might spend time arranging a one-page diagram for a review, resizing groups, moving labels, or emphasizing a specific boundary. A subsequent AI refinement could improve the service selection while disrupting that carefully prepared layout.&lt;/P&gt;
&lt;P&gt;The design principle that emerged was straightforward:&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;AI acceleration should preserve deliberate human work by default.&lt;/STRONG&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;Refinement now retains existing node positions, group geometry, sizes, and viewport context whenever possible. The model can change the architecture without treating every turn as permission to redraw the entire document.&lt;/P&gt;
&lt;P&gt;The same principle applies beyond geometry:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Preserve the prior validation result when recommendations change the architecture.&lt;/LI&gt;
&lt;LI&gt;Preserve the active light or dark theme in exported artifacts.&lt;/LI&gt;
&lt;LI&gt;Preserve the distinction between the authoring canvas and the presentation deliverable.&lt;/LI&gt;
&lt;LI&gt;Preserve user-configured pricing assumptions rather than replacing them with one fixed estimate.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;This is a broader lesson for AI-assisted tools. A generated result is not the only source of value. The edits, judgments, and communication choices a person adds afterward are part of the artifact too.&lt;/P&gt;
&lt;FIGURE&gt;&lt;BR /&gt;
&lt;FIGCAPTION aria-hidden="true"&gt;Before-and-after AADB canvases showing an AI refinement that adds Azure Front Door and WAF while retaining the positions of eight existing services and the anchors of four existing groups.&lt;/FIGCAPTION&gt;
&lt;/FIGURE&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 3.&lt;/STRONG&gt; In this controlled synthetic refinement, all eight existing service positions and four group anchors remained unchanged. The containing Application group expanded to accommodate the new edge tier, so preservation does not imply that every group dimension stays fixed.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;HR /&gt;
&lt;H2 id="quality-is-structural-not-only-visual"&gt;Quality is structural, not only visual&lt;/H2&gt;
&lt;P&gt;A diagram can look polished while still being architecturally confusing. Early feedback exposed cases where a generated service appeared disconnected because a model referenced a display name instead of the service identifier used by the canvas.&lt;/P&gt;
&lt;P&gt;The correction was not another prompt instruction alone. The application now resolves connection endpoints across identifiers, normalized service names, and service-type aliases. It repairs valid edges, drops invalid or self-referential edges, detects remaining orphan nodes, and records aggregate integrity signals.&lt;/P&gt;
&lt;P&gt;That creates a more useful definition of diagram quality:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Are the services connected as intended?&lt;/LI&gt;
&lt;LI&gt;Were any generated edges repaired or dropped?&lt;/LI&gt;
&lt;LI&gt;Are there orphaned nodes?&lt;/LI&gt;
&lt;LI&gt;Did refinement preserve the existing layout?&lt;/LI&gt;
&lt;LI&gt;Did an architecture change receive a fresh validation?&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Visual polish still matters, especially when an artifact leaves the editor. But structural integrity gives the product something deterministic to test and monitor.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="stage-3-validate-improve-treat-validation-as-a-lifecycle"&gt;Stage 3: Validate &amp;amp; Improve — treat validation as a lifecycle&lt;/H2&gt;
&lt;P&gt;The Azure Well-Architected Framework is most useful when validation becomes iterative rather than ceremonial.&lt;/P&gt;
&lt;P&gt;The Diagram Builder can assess a proposed design across the five Well-Architected pillars, surface findings, and apply selected recommendations. But that workflow exposed an important state-management problem: when the architecture changed, the prior validation result disappeared along with the obvious route back to revalidation.&lt;/P&gt;
&lt;P&gt;The updated experience keeps the previous report, marks it &lt;STRONG&gt;Revalidate Needed&lt;/STRONG&gt;, and makes clear that the score describes an earlier state of the architecture. A new validation replaces it only after the updated design has been assessed.&lt;/P&gt;
&lt;P&gt;This distinction prevents a stale score from looking current.&lt;/P&gt;
&lt;P&gt;It also clarifies what an architecture-level assessment can and cannot prove. A diagram may show that a WAF, cache, backup service, or secondary region exists. It usually cannot prove that purge protection, diagnostic routing, encryption settings, role assignments, health probes, or failover policies are configured correctly.&lt;/P&gt;
&lt;P&gt;That is why validation findings need to distinguish between:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Pattern-level gaps&lt;/STRONG&gt; — missing or misplaced architectural components&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Configuration-level gaps&lt;/STRONG&gt; — required settings that must be verified in Infrastructure as Code or the deployed environment&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Generated scores and recommendations help architects review a design; they do not replace an Azure Well-Architected Review, security review, deployment validation, or professional judgment.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;FIGURE&gt;
&lt;FIGCAPTION aria-hidden="true"&gt;Validation result retained after architecture recommendations are applied, with a Revalidate Needed status and action.&lt;/FIGCAPTION&gt;
&lt;/FIGURE&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 4.&lt;/STRONG&gt; Architecture changes make a previous validation historical, not useless. The result remains available while the interface clearly asks for a fresh validation.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;HR /&gt;
&lt;H2 id="stage-4-share-or-build-design-for-the-artifacts-destination"&gt;Stage 4: Share or Build — design for the artifact’s destination&lt;/H2&gt;
&lt;P&gt;The editing canvas and the final deliverable serve different purposes.&lt;/P&gt;
&lt;P&gt;Canvas dots, handles, navigation controls, and selection states help during authoring. They can make an exported diagram feel unfinished. The Diagram Builder now separates those concerns with &lt;STRONG&gt;Plain, Dots, and Grid&lt;/STRONG&gt; export backgrounds while preserving the active light or dark theme.&lt;/P&gt;
&lt;FIGURE&gt;&lt;BR /&gt;
&lt;FIGCAPTION aria-hidden="true"&gt;The same AADB architecture shown first on the editing canvas with the export menu open and then as the resulting Plain PNG without authoring controls.&lt;/FIGCAPTION&gt;
&lt;/FIGURE&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 5.&lt;/STRONG&gt; Authoring and delivery are different contexts. The upper view shows the editable canvas and its real export controls; the lower view is the Plain PNG produced from that same canvas, without editing chrome.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;Cost language is deliberately qualified. Azure services often combine fixed, usage-based, and configuration-dependent charges. A baseline that includes six numerically priced services but excludes 20 usage-based items is not the total cost of the architecture. The output identifies those exclusions rather than treating missing values as zero.&lt;/P&gt;
&lt;P&gt;The final stage also includes deployment guides and Infrastructure as Code. Here, honesty about artifact coverage is essential. A generated Bicep file may be a useful starter while still omitting private endpoints, diagnostic settings, failover configuration, or service-specific resources. The artifact should state what it implements, what remains conceptual, and whether Azure Resource Manager validation passed.&lt;/P&gt;
&lt;P&gt;AI-generated diagrams, costs, validation results, deployment guides, and Infrastructure as Code should all be reviewed and validated before production use.&lt;/P&gt;
&lt;H2 id="the-same-journey-now-extends-to-agents"&gt;The same journey now extends to agents&lt;/H2&gt;
&lt;P&gt;The MCP server introduced in the previous article makes the Diagram Builder available to agent experiences such as Microsoft Scout. The four-stage journey provides a useful way to think about agent orchestration too:&lt;/P&gt;
&lt;OL type="1"&gt;
&lt;LI&gt;Import or create one canonical architecture.&lt;/LI&gt;
&lt;LI&gt;Refine it without silently changing the intended topology.&lt;/LI&gt;
&lt;LI&gt;Validate it, apply supported improvements, and revalidate.&lt;/LI&gt;
&lt;LI&gt;Render or generate artifacts with explicit coverage and limitations.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;The current MCP surface exposes &lt;STRONG&gt;12 tools, three resources, and three reusable prompts&lt;/STRONG&gt;. It can normalize an existing architecture, validate and harden it deterministically, estimate regional costs from a dated pricing snapshot, render presentation/technical/cost views, and generate Bicep, Terraform, and deployment guidance. The calling agent still owns orchestration and reasoning; the MCP server is intended to remain a deterministic architecture capability, not a second hidden agent.&lt;/P&gt;
&lt;P&gt;The native MCP renderer can project one canonical architecture into three communication profiles:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Presentation&lt;/STRONG&gt; emphasizes the primary request path, reduces supporting labels, and removes pricing.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Technical&lt;/STRONG&gt; preserves complete connection detail for engineering inspection.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Cost&lt;/STRONG&gt; retains the focused composition while adding service-level pricing assumptions, a fixed-priced baseline, and explicit exclusions.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;These are MCP-generated SVG views, not Blueprint diagrams or screenshots of the editable web canvas. The services, connections, and groups remain the same; only the information treatment changes.&lt;/P&gt;
&lt;FIGURE&gt;&lt;BR /&gt;
&lt;FIGCAPTION aria-hidden="true"&gt;The AADB MCP renderer projecting the same canonical architecture into presentation, technical, and cost SVG profiles.&lt;/FIGCAPTION&gt;
&lt;/FIGURE&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 6.&lt;/STRONG&gt; Native AADB MCP output from one 8-service, 9-connection, 4-group architecture. Presentation prioritizes the story, Technical exposes connection detail, and Cost foregrounds pricing assumptions and exclusions.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;Recent work on the MCP renderer added purpose-built presentation, technical, and cost profiles. More importantly, testing agent-generated artifacts reinforced an accountability principle: a polished diagram and a compiled Bicep file do not prove deployability.&lt;/P&gt;
&lt;P&gt;An agent workflow should report whether topology changed, whether validation improved, which services are represented only conceptually, and whether the generated IaC passed Azure preflight. That is more useful than an unsupported claim that a design is production-ready.&lt;/P&gt;
&lt;P&gt;Trust also includes the tool boundary itself. The hosted MCP endpoints now require a bearer token for real session operations; missing or incorrect credentials are rejected. A shared token is appropriate for the current controlled integration, but it is not the end state for enterprise multi-user access. Entra ID/OAuth, per-client authorization, rotation, and revocation remain future hardening work.&lt;/P&gt;
&lt;FIGURE&gt;&lt;BR /&gt;
&lt;FIGCAPTION aria-hidden="true"&gt;Microsoft Scout response after an authenticated Azure Architecture Diagram Builder MCP workflow, showing the tools used, initial and final validation scores, cost scope, Bicep classification, rendered architecture, artifact links, coverage gaps, and no-deployment warning.&lt;/FIGCAPTION&gt;
&lt;/FIGURE&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 7.&lt;/STRONG&gt; The guided lifecycle extends beyond the web application. In this synthetic Scout run with GPT-5.6 Sol, the agent used authenticated AADB MCP tools to validate, harden, cost, render, and generate starter artifacts while explicitly reporting coverage gaps and that nothing was deployed.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;HR /&gt;
&lt;H2 id="learning-from-adoption-without-identifying-people"&gt;Learning from adoption without identifying people&lt;/H2&gt;
&lt;P&gt;Product learning does not require reconstructing individual identities.&lt;/P&gt;
&lt;P&gt;The findings behind this article use aggregate, deduplicated application telemetry, public article counters, public repository totals, and paraphrased feedback themes. They do not correlate Application Insights identifiers, feedback records, GitHub accounts, or email addresses.&lt;/P&gt;
&lt;P&gt;Written feedback remains submittable without contact information. When someone explicitly opts into follow-up, the email address is stored with the feedback record in Cosmos DB and is not sent to normal product telemetry. The current 180-day expiry field is a retention marker; automated deletion must be implemented and verified before describing that retention period as enforced.&lt;/P&gt;
&lt;P&gt;Those boundaries matter for both product design and public writing:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Aggregate activity rather than profiling individuals.&lt;/LI&gt;
&lt;LI&gt;Paraphrase themes rather than publishing comments without permission.&lt;/LI&gt;
&lt;LI&gt;Keep optional contact consent separate from telemetry.&lt;/LI&gt;
&lt;LI&gt;Avoid presenting anonymous identifiers as confirmed people.&lt;/LI&gt;
&lt;LI&gt;Avoid claiming that publication timing proves acquisition causality.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;This is not a claim of legal compliance. It is a product discipline: collect less, preserve user agency, and make only the claims the evidence supports.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="what-changed"&gt;What changed&lt;/H2&gt;
&lt;P&gt;The guided journey is the visible result, but the deeper change is how the project now evaluates progress.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table class="lia-border-color-custom-d6dee6 lia-border-style-solid" border="1" style="width: 100%; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th class="lia-border-color-custom-0078d4 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Earlier question&lt;/th&gt;&lt;th class="lia-border-color-custom-0078d4 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Better question&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Did the model generate a diagram?&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Did it generate a connected and understandable architecture?&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Did the user click Validate?&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Was the current architecture validated, and was it revalidated after changes?&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Did export start?&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Did a professional artifact finish generating successfully?&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Does the IaC compile?&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;What does it actually implement, and does Azure preflight pass?&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;How many features exist?&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Can an architect understand the next useful step?&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;The model portfolio continued to evolve as well. The production selector now contains 15 configured entries, including &lt;STRONG&gt;MAI-Thinking-1 (Public Preview)&lt;/STRONG&gt;. But the more consequential changes in this article are deliberately model-independent: preserve human work, keep state and provenance explicit, qualify generated artifacts, and authenticate the tools agents can call.&lt;/P&gt;
&lt;P&gt;The goal is not to remove flexibility. Architects can still open any tool directly, rearrange the canvas, reject recommendations, change pricing assumptions, or export at any point.&lt;/P&gt;
&lt;P&gt;The goal is to make the workflow coherent without pretending architecture itself is linear.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="the-next-boundary-logical-versus-physical-architecture"&gt;The next boundary: logical versus physical architecture&lt;/H2&gt;
&lt;P&gt;Recent feedback points to a harder problem than adding another model or export format. Architects working with private Azure AI landing zones need to distinguish shared platform resources from project-owned resources, preserve VNet and subnet boundaries, and reason about CIDRs, NSGs, route tables, private endpoints, DNS, and managed identities.&lt;/P&gt;
&lt;P&gt;The current Topology mode can show services and relationships, but it should not imply exact physical fidelity when those facts are absent. A useful logical diagram answers &lt;EM&gt;what exists and how it interacts&lt;/EM&gt;. A physical or low-level design must answer &lt;EM&gt;where it is deployed, how it is isolated, and which values came from evidence&lt;/EM&gt;.&lt;/P&gt;
&lt;P&gt;That is the next technical direction I am exploring: an evidence-aware Physical Architecture view backed by deterministic reconstruction from Terraform plan/state, ARM, or a live Azure inventory. Exact fields would be labeled as &lt;STRONG&gt;observed&lt;/STRONG&gt; or &lt;STRONG&gt;resolved&lt;/STRONG&gt;; AI suggestions would remain explicitly &lt;STRONG&gt;proposed&lt;/STRONG&gt;; unsupported or missing inputs would be reported instead of silently invented.&lt;/P&gt;
&lt;P&gt;This capability is not shipped today, and it will require its own schema, validation rules, layout, security review, and evaluation set. That distinction matters. The lesson from adoption is not to put every architecture concern into one crowded canvas. It is to make each artifact’s purpose and evidence boundary clear.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="try-it-challenge-it-help-shape-what-comes-next"&gt;Try it, challenge it, help shape what comes next&lt;/H2&gt;
&lt;P&gt;The Azure Architecture Diagram Builder remains open source, and the live experience is available today:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Live app:&lt;/STRONG&gt; &lt;A href="https://aka.ms/diagram-builder" target="_blank"&gt;https://aka.ms/diagram-builder&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Source code:&lt;/STRONG&gt; &lt;A href="https://github.com/Arturo-Quiroga-MSFT/azure-architecture-diagram-builder" target="_blank"&gt;github.com/Arturo-Quiroga-MSFT/azure-architecture-diagram-builder&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Getting started:&lt;/STRONG&gt; &lt;A href="https://github.com/Arturo-Quiroga-MSFT/azure-architecture-diagram-builder/blob/main/DOCS/getting-started-guide.md" target="_blank"&gt;Documentation and deployment guidance&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The next phase is to measure whether the guided journey helps people complete the full loop, especially recommendation-to-revalidation and artifact-generation success. In parallel, I am beginning the narrower physical-architecture investigation described above. Both efforts will use aggregate signals, reviewed fixtures, and sufficiently large cohorts rather than individual journey reconstruction.&lt;/P&gt;
&lt;P&gt;Try the workflow with a real architecture problem. Tell me where the handoffs are unclear, where the diagram loses intent, or where an artifact claims more than it implements. Those are the gaps worth fixing next.&lt;/P&gt;
&lt;HR /&gt;
&lt;P&gt;&lt;STRONG&gt;Measurement note:&lt;/STRONG&gt; Article views are rounded public counters observed August 5, 2026. Application figures use deduplicated retained telemetry through August 13 and anonymous identifiers. GitHub totals and rolling 14-day traffic were observed August 13. The comparison windows are May 19–July 9 and July 10–August 13. These signals have different populations and must not be added together. Timing comparisons show concurrent activity, not causal attribution.&lt;/P&gt;</description>
      <pubDate>Thu, 13 Aug 2026 22:10:46 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/from-features-to-flow-how-real-world-adoption-reshaped-the-azure/ba-p/4546817</guid>
      <dc:creator>arturoqu</dc:creator>
      <dc:date>2026-08-13T22:10:46Z</dc:date>
    </item>
    <item>
      <title>Skill or Sub-Agent. Choosing AI Capabilities You Will Actually Reuse</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/skill-or-sub-agent-choosing-ai-capabilities-you-will-actually/ba-p/4542099</link>
      <description>&lt;P&gt;&lt;STRONG&gt;Audience:&lt;/STRONG&gt; Cloud architects, platform engineers, engineering leaders&lt;/P&gt;
&lt;HR /&gt;
&lt;H2&gt;The wrong first question&lt;/H2&gt;
&lt;P&gt;Most teams building AI capabilities start with the wrong question. They ask which model to use.&lt;/P&gt;
&lt;P&gt;The model matters less than the shape of the capability around it. The first real fork is this. Are you building a skill or a sub-agent? Get that wrong and no model choice will save you. A skill and a sub-agent are two different delivery shapes, and each one fails at the other one's job.&lt;/P&gt;
&lt;H2&gt;The insight&lt;/H2&gt;
&lt;P&gt;The choice between a skill and a sub-agent is not about model power. It comes down to four checks. How the work iterates, whether the output carries a voice, how far an early wrong turn spreads, and how often it recurs. Score each, count which way they lean, and the shape falls out. An even split means build both, and let the skill drive the sub-agent.&lt;/P&gt;
&lt;P&gt;The three sections below take the checks worth a pause. Frequency is the plain one: a one-off craft piece leans to a skill, a repeatable batch job to a sub-agent.&lt;/P&gt;
&lt;P&gt;A skill lives inside the conversation. It reads files, asks a question, refines with the author, and keeps a human in the loop mid-flight. A sub-agent takes one prompt, runs to completion, and returns one report. Both are useful, for different work.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;th&gt;Dimension&lt;/th&gt;&lt;th&gt;Skill&lt;/th&gt;&lt;th&gt;Sub-agent&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Iteration&lt;/td&gt;&lt;td&gt;Conversation, many turns&lt;/td&gt;&lt;td&gt;One hand-off, one pass&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Voice&lt;/td&gt;&lt;td&gt;Holds a style profile and applies it&lt;/td&gt;&lt;td&gt;Drifts toward generic by design&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Human gate&lt;/td&gt;&lt;td&gt;Every turn&lt;/td&gt;&lt;td&gt;Once, at the end&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Best for&lt;/td&gt;&lt;td&gt;Craft, subjective output&lt;/td&gt;&lt;td&gt;Bounded, structured output&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;&lt;EM&gt;Table 1. The same three dimensions decide the shape every time.&lt;/EM&gt;&lt;/P&gt;
&lt;H2&gt;1. Decide the iteration model first&lt;/H2&gt;
&lt;P&gt;Before anything else, architects should ask how the work actually happens. Is it a conversation or a hand-off? That single answer removes most of the ambiguity.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Craft work needs back and forth&lt;/LI&gt;
&lt;LI&gt;Batch work needs one clean pass&lt;/LI&gt;
&lt;LI&gt;Conversations need memory of the thread&lt;/LI&gt;
&lt;LI&gt;Hand-offs need a bounded input and a clear output&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;A skill is right when the value comes from iteration. A blog post, a design review, a tricky refactor. A sub-agent is right when the work is well defined and the output is the deliverable.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: use a skill when the team expects three or four rounds of "close, but change this". The trade-off: a skill costs more attention per run because a human stays involved. The trap to avoid: forcing iterative craft into a one-shot agent and then editing the output by hand every time.&lt;/P&gt;
&lt;H2&gt;2. Voice fidelity decides craft work&lt;/H2&gt;
&lt;P&gt;Some outputs have a voice. An article, a customer email, an architecture narrative. Others do not. A query result, a data export, a status summary. The line between them is not cosmetic. It decides which shape survives review.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Voice-heavy work favours a skill&lt;/LI&gt;
&lt;LI&gt;Voice-neutral work favours a sub-agent&lt;/LI&gt;
&lt;LI&gt;Skills can hold a style profile and apply it&lt;/LI&gt;
&lt;LI&gt;Sub-agents drift toward generic by design&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;When the output carries a name, fidelity is the whole game. A capable model with no voice anchor produces text that reads like it came from a committee.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: give a skill an explicit voice profile with banned phrases and cadence rules. Let it self-check before it shows anyone anything. The trade-off: the profile takes real effort to write once. The trap to avoid: expecting a stateless agent to match a personal style from a single prompt.&lt;/P&gt;
&lt;H3&gt;Implementation note&lt;/H3&gt;
&lt;P&gt;A voice profile is not documentation. It lives in the skill definition, an executable contract the skill checks itself against before a draft is ever shown. A small profile goes a long way.&lt;/P&gt;
&lt;LI-CODE lang="yaml"&gt;# voice-profile (excerpt)
banned_phrases: [seamless, robust, game-changing, leverage the power]
forbid: [em-dash, semicolon, exclamation in body]
max_avg_sentence_words: 20
require:
  - one "In practice" block per section
  - a closing discussion question
self_check: run before any draft is shown to a human
&lt;/LI-CODE&gt;
&lt;H2&gt;3. Put the human gate where the risk is&lt;/H2&gt;
&lt;P&gt;Every AI capability needs a human review gate. The design question is where that gate sits. Placement is the difference between catching a problem early and unpicking it later.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Skills gate continuously, turn by turn&lt;/LI&gt;
&lt;LI&gt;Sub-agents gate once, at the end&lt;/LI&gt;
&lt;LI&gt;Continuous gates catch drift early&lt;/LI&gt;
&lt;LI&gt;End gates are cheaper but riskier for craft&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;If a wrong turn early corrupts everything after it, the team wants a skill. If the work is bounded and a bad output is easy to spot and discard, an end gate is fine.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: match the gate to the blast radius. High blast radius and subjective quality point to a skill. Low blast radius and objective output point to a sub-agent. The trap to avoid: a one-shot agent doing forty minutes of unattended work that a human then has to unpick.&lt;/P&gt;
&lt;H2&gt;4. The pattern that scales is both&lt;/H2&gt;
&lt;P&gt;The mature answer is not one or the other. It is a skill on top of a sub-agent. The two shapes compose cleanly when each one keeps to its own job.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;The skill orchestrates and holds the voice&lt;/LI&gt;
&lt;LI&gt;The sub-agent executes bounded sub-tasks&lt;/LI&gt;
&lt;LI&gt;The human reviews at the skill layer&lt;/LI&gt;
&lt;LI&gt;Each layer does what it is good at&lt;/LI&gt;
&lt;/UL&gt;
&lt;img /&gt;
&lt;P&gt;&lt;EM&gt;Figure 1. In the combined pattern the human reviews at the skill layer, and the sub-agent only touches the bounded task.&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;The skill runs the conversation and keeps quality. When it needs a bounded, repeatable job done, it delegates to a sub-agent. The result is iteration where craft lives and automation where the work is mechanical.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: a skill drafts and refines an article with the author, and calls a sub-agent to fetch and summarise reference material. The trade-off: two layers are more to build than one. The trap to avoid: collapsing both into a single agent and losing either the voice or the automation.&lt;/P&gt;
&lt;H2&gt;The operational trade-offs&lt;/H2&gt;
&lt;P&gt;Shape is not only a design choice. It shows up in cost, latency, and how you debug a bad run. Architects should price these in before committing to a pattern.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;A skill spends more tokens and more human minutes per run&lt;/LI&gt;
&lt;LI&gt;A sub-agent spends compute once and returns fast&lt;/LI&gt;
&lt;LI&gt;A skill fails in small, visible steps you can correct&lt;/LI&gt;
&lt;LI&gt;A sub-agent fails as one block you inspect after the fact&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The cost of a skill is attention. Someone stays in the loop and that time is real. The cost of a sub-agent is rework. When a one-shot run goes wrong, the whole output is suspect and someone redoes it. Observability follows the same split. A skill leaves a turn-by-turn trail you can read. A sub-agent leaves one input and one output. You instrument the boundary and log the prompt and the result. Pick the shape whose failure mode your team can afford. The wrong shape does not announce itself. It shows up later as a cost line or a rewrite.&lt;/P&gt;
&lt;H2&gt;Two capabilities, one team&lt;/H2&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Consider a team standardising its engineering work with AI. Two capabilities land on the backlog in the same week. The first is recurring status queries. Well defined input, structured output, no voice. A stateless sub-agent fits. One prompt in, one report out, gate at the end. It works on day one and keeps working.&lt;/P&gt;
&lt;P&gt;The second is authored technical content. Subjective, voice-heavy, many rounds of refinement. The reflex is to reuse the sub-agent that just shipped. That reflex is the mistake. The queries stay clean. The content reads flat and generic, and every draft needs a heavy human rewrite. Rebuilt as a skill with a voice profile and a turn-by-turn gate, the same work compounds instead of fighting back. Same team, same models, two different shapes of work, and only one right tool for each.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;What teams get wrong&lt;/H2&gt;
&lt;P&gt;The common pattern is defaulting to whichever shape the team built first. A team ships one sub-agent, likes it, and forces every new problem into a sub-agent. Or it builds one skill and runs everything as a conversation, including batch work that should be automated.&lt;/P&gt;
&lt;P&gt;It looks like consistency. It feels like reuse. But it leads to craft work that reads generic and batch work that needs babysitting. The fix is not a better model. It is naming the shape of the work before picking the tool.&lt;/P&gt;
&lt;P&gt;The three shapes to watch for in your own stack:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;A voiced deliverable coming out of a one-shot agent, rewritten by hand every run. A skill wearing a sub-agent costume.&lt;/LI&gt;
&lt;LI&gt;A batch job run as a conversation and babysat turn by turn. A sub-agent wearing a skill costume.&lt;/LI&gt;
&lt;LI&gt;A large workflow forced into one agent that holds neither the voice nor the automation. Two shapes collapsed into one.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Name which one you are looking at, and the fix picks itself.&lt;/P&gt;
&lt;H2&gt;A quick way to decide&lt;/H2&gt;
&lt;P&gt;When a new capability lands on the backlog, run four checks before picking a tool.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Iteration: conversation or one hand-off&lt;/LI&gt;
&lt;LI&gt;Output: subjective and voiced, or structured and neutral&lt;/LI&gt;
&lt;LI&gt;Blast radius: does an early wrong turn corrupt the rest&lt;/LI&gt;
&lt;LI&gt;Frequency: a one-off craft piece, or a repeatable batch job&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Three or more answers leaning subjective and iterative point to a skill. Three or more leaning structured and repeatable point to a sub-agent. A split answer usually means a skill orchestrating a sub-agent underneath.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;EM&gt;Figure 2. Score the four checks and count the leanings. Three or four one way pick the shape. An even split means a skill orchestrating a sub-agent.&lt;/EM&gt;&lt;/P&gt;
&lt;H2&gt;Where to start depends on what you have already built&lt;/H2&gt;
&lt;P&gt;The framework is the destination. Where you start depends on what your team has shipped so far. Find your stage and take the one first move for it this week.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;th&gt;Stage&lt;/th&gt;&lt;th&gt;First move, this week&lt;/th&gt;&lt;th&gt;Watch out for&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Just starting, nothing built yet&lt;/td&gt;&lt;td&gt;Pick the single task you repeat most and write a one-paragraph capability brief for it, iteration, output, blast radius, and frequency, before you build. The brief names the shape, and the shape names the tool.&lt;/td&gt;&lt;td&gt;Building a general assistant before you have named one concrete job.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;One capability, reused for everything&lt;/td&gt;&lt;td&gt;List every job you push through the one tool, find the one whose shape does not match, and rebuild just that one in the right shape. You do not need to replace what works.&lt;/td&gt;&lt;td&gt;Forcing new work into the tool you already have.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;A small fleet, a handful of capabilities&lt;/td&gt;&lt;td&gt;Take your largest layered workflow and split it, a skill that holds the voice and the human gate on top, a sub-agent that does the bounded work underneath.&lt;/td&gt;&lt;td&gt;Capabilities that duplicate each other with no composition between them.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;&lt;EM&gt;Table 2. Same framework, different first move. What you have already built decides where the leverage is this week.&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;Your setup also shapes the answer. A solo builder should optimise for their own voice and iteration speed, where one strong skill beats three thin ones. A platform team should standardise the capability brief and a shared voice profile, so the fleet stays consistent as more people add to it, and a new capability inherits the house style instead of drifting from it.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;EM&gt;Figure 3. Whatever you have built so far, the first move has the same shape. Name the work before the tool, then match the shape to the tool.&lt;/EM&gt;&lt;/P&gt;
&lt;H2&gt;The shift&lt;/H2&gt;
&lt;P&gt;The shift is from "what can the model do" to "what shape is the work". Model capability is table stakes now. The advantage is in matching the capability to the work. Our own capability fleet is built this way, interactive skills and autonomous sub-agents in separate places with an orchestrator on top, and that split is what keeps it maintainable as it grows.&lt;/P&gt;
&lt;P&gt;Iterative and voice-heavy points to a skill. Bounded and mechanical points to a sub-agent. Large and layered points to a skill orchestrating sub-agents. Decide that first, and the model becomes a detail the team can change later without rebuilding anything.&lt;/P&gt;
&lt;P&gt;Most teams collapse both ideas into "automation" and end up with neither. The teams that separate them build capabilities they actually reuse.&lt;/P&gt;
&lt;HR /&gt;
&lt;P&gt;&lt;STRONG&gt;Want to discuss?&lt;/STRONG&gt; Drop a comment with patterns you have seen in your environment. I read every reply.&lt;/P&gt;
&lt;!--
Taxonomy suggestions:
  Primary product: Azure
  Secondary: .NET, AI + Machine Learning
  Tags: AI agents, developer tools, engineering leadership, platform engineering, software design
--&gt;</description>
      <pubDate>Thu, 30 Jul 2026 20:50:02 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/skill-or-sub-agent-choosing-ai-capabilities-you-will-actually/ba-p/4542099</guid>
      <dc:creator>KishoreKumarPattabiraman</dc:creator>
      <dc:date>2026-07-30T20:50:02Z</dc:date>
    </item>
    <item>
      <title>Mastering GitHub Copilot Budgets: How to Prevent Surprise Overages Without Blocking Devs</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/mastering-github-copilot-budgets-how-to-prevent-surprise/ba-p/4542073</link>
      <description>&lt;H3 data-path-to-node="11"&gt;&lt;STRONG data-path-to-node="11" data-index-in-node="0"&gt;1. The Core Mental Model: The Water Park Analogy&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P data-path-to-node="12"&gt;To understand Copilot billing, think of your enterprise as a &lt;STRONG data-path-to-node="12" data-index-in-node="61"&gt;water park&lt;/STRONG&gt;:&lt;/P&gt;
&lt;UL data-path-to-node="13"&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="13,0,0" data-index-in-node="0"&gt;The Shared Pool (Included Usage):&lt;/STRONG&gt; Every Copilot Business ($19/mo) and Enterprise ($39/mo) seat license adds a set number of free AI credits into a shared corporate water tank (1,900 or 3,900 credits/seat, respectively).&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="13,1,0" data-index-in-node="0"&gt;Phase 1 (Pool Phase):&lt;/STRONG&gt; All licensed developers drink from this central pool for free until it runs empty.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="13,2,0" data-index-in-node="0"&gt;Phase 2 (Metered Overage):&lt;/STRONG&gt; Once the pool is dry, extra water costs &lt;STRONG data-path-to-node="13,2,0" data-index-in-node="67"&gt;$0.01 per credit&lt;/STRONG&gt;. This phase only activates if your enterprise explicitly enables the &lt;STRONG data-path-to-node="13,2,0" data-index-in-node="153"&gt;"AI Credit Paid Usage"&lt;/STRONG&gt; policy.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="13,3,0" data-index-in-node="0"&gt;User-Level Budgets (ULBs):&lt;/STRONG&gt; Personal wristbands limiting how much water &lt;EM data-path-to-node="13,3,0" data-index-in-node="71"&gt;one person&lt;/EM&gt; can consume across both Phase 1 and Phase 2 combined.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="13,4,0" data-index-in-node="0"&gt;Group / Enterprise Budgets:&lt;/STRONG&gt; Group bar tabs that kick in &lt;STRONG data-path-to-node="13,4,0" data-index-in-node="56"&gt;only during Phase 2&lt;/STRONG&gt; to cap overage spending.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3 data-path-to-node="15"&gt;&lt;STRONG data-path-to-node="15" data-index-in-node="0"&gt;2. How GitHub Evaluates Requests (The 3-Step Decision Flow)&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P data-path-to-node="16"&gt;Every time a developer uses an AI credit feature (like GitHub Copilot Chat or custom Agents), GitHub evaluates the request through a 3-step hierarchy:&lt;/P&gt;
&lt;img /&gt;
&lt;H4 data-path-to-node="18"&gt;&lt;STRONG data-path-to-node="18" data-index-in-node="0"&gt;Step 1: The Personal Guardrail (User-Level Budgets)&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P data-path-to-node="19"&gt;ULBs cap total consumption (Pool + Metered). These are &lt;STRONG data-path-to-node="19" data-index-in-node="55"&gt;always hard stops&lt;/STRONG&gt;. The system selects the most specific rule:&lt;/P&gt;
&lt;OL data-path-to-node="20"&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="20,0,0" data-index-in-node="0"&gt;Individual ULB:&lt;/STRONG&gt; Overrides everything (e.g., $50 for a Lead Data Scientist).&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="20,1,0" data-index-in-node="0"&gt;Cost Center ULB:&lt;/STRONG&gt; Overrides Universal (e.g., $20/dev for Engineering, $5/dev for Marketing).&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="20,2,0" data-index-in-node="0"&gt;Universal ULB:&lt;/STRONG&gt; The global baseline for all licensed users.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H4 data-path-to-node="21"&gt;&lt;STRONG data-path-to-node="21" data-index-in-node="0"&gt;Step 2: The Included Pool Check&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P data-path-to-node="22"&gt;If the user hasn't hit their ULB, the system checks if pooled credits remain. You can also configure &lt;STRONG data-path-to-node="22" data-index-in-node="101"&gt;Included Usage Controls&lt;/STRONG&gt; on Cost Centers to ring-fence a team's pool draw to match their contributed seats.&lt;/P&gt;
&lt;H4 data-path-to-node="23"&gt;&lt;STRONG data-path-to-node="23" data-index-in-node="0"&gt;Step 3: Group &amp;amp; Enterprise Overage Caps&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P data-path-to-node="24"&gt;Once the pool is empty, metered charges begin ($0.01/credit). The system checks overage budgets in order:&lt;/P&gt;
&lt;OL data-path-to-node="25"&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="25,0,0" data-index-in-node="0"&gt;Cost Center Budget&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="25,1,0" data-index-in-node="0"&gt;Organization Budget&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="25,2,0" data-index-in-node="0"&gt;Enterprise Budget&lt;/STRONG&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;P data-path-to-node="26,0"&gt;&lt;STRONG data-path-to-node="26,0" data-index-in-node="0"&gt;⚠️ Critical FinOps Note:&lt;/STRONG&gt; Group/Enterprise budgets &lt;STRONG data-path-to-node="26,0" data-index-in-node="50"&gt;do NOT stop usage by default&lt;/STRONG&gt;. You must explicitly enable the setting:&lt;/P&gt;
&lt;P data-path-to-node="26,1"&gt;Stop usage when budget limit is reached&lt;/P&gt;
&lt;P data-path-to-node="26,2"&gt;Without this setting toggled &lt;STRONG data-path-to-node="26,2" data-index-in-node="29"&gt;ON&lt;/STRONG&gt;, overage charges will accrue uncapped!&lt;/P&gt;
&lt;H3 data-path-to-node="26,2"&gt;&lt;STRONG&gt;3. Key Takeaways &amp;amp; Common Pitfalls&lt;/STRONG&gt;&lt;/H3&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 100%; height: 198px; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr style="height: 34.8px;"&gt;&lt;td style="height: 34.8px;"&gt;&lt;STRONG&gt;Budget Control&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 34.8px;"&gt;&lt;STRONG&gt;What It Limits&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 34.8px;"&gt;&lt;STRONG&gt;Active Phase&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 34.8px;"&gt;&lt;STRONG&gt;Hard Stop Built-In?&lt;/STRONG&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr style="height: 58.8px;"&gt;&lt;td style="height: 58.8px;"&gt;&lt;SPAN data-path-to-node="29,1,0,0"&gt;&lt;STRONG data-path-to-node="29,1,0,0" data-index-in-node="0"&gt;Individual / Cost Center / Universal ULB&lt;/STRONG&gt;&lt;/SPAN&gt;&lt;/td&gt;&lt;td style="height: 58.8px;"&gt;&lt;SPAN data-path-to-node="29,1,1,0"&gt;Single user total credits&lt;/SPAN&gt;&lt;/td&gt;&lt;td style="height: 58.8px;"&gt;&lt;STRONG&gt;&lt;SPAN data-path-to-node="29,1,2,0"&gt;Pool + Metered&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 58.8px;"&gt;&lt;SPAN data-path-to-node="29,1,3,0"&gt;&lt;STRONG data-path-to-node="29,1,3,0" data-index-in-node="0"&gt;Yes&lt;/STRONG&gt; (Always)&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 34.8px;"&gt;&lt;td style="height: 34.8px;"&gt;&lt;SPAN data-path-to-node="29,2,0,0"&gt;&lt;STRONG data-path-to-node="29,2,0,0" data-index-in-node="0"&gt;Cost Center Budget&lt;/STRONG&gt;&lt;/SPAN&gt;&lt;/td&gt;&lt;td style="height: 34.8px;"&gt;&lt;SPAN data-path-to-node="29,2,1,0"&gt;Team overage spend&lt;/SPAN&gt;&lt;/td&gt;&lt;td style="height: 34.8px;"&gt;&lt;SPAN data-path-to-node="29,2,2,0"&gt;&lt;STRONG data-path-to-node="29,2,2,0" data-index-in-node="0"&gt;Metered Only&lt;/STRONG&gt;&lt;/SPAN&gt;&lt;/td&gt;&lt;td style="height: 34.8px;"&gt;&lt;SPAN data-path-to-node="29,2,3,0"&gt;Only if Stop Usage is &lt;STRONG&gt;ON&lt;/STRONG&gt;&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 34.8px;"&gt;&lt;td style="height: 34.8px;"&gt;&lt;SPAN data-path-to-node="29,3,0,0"&gt;&lt;STRONG data-path-to-node="29,3,0,0" data-index-in-node="0"&gt;Org Budget&lt;/STRONG&gt;&lt;/SPAN&gt;&lt;/td&gt;&lt;td style="height: 34.8px;"&gt;&lt;SPAN data-path-to-node="29,3,1,0"&gt;Organization overage spend&lt;/SPAN&gt;&lt;/td&gt;&lt;td style="height: 34.8px;"&gt;&lt;SPAN data-path-to-node="29,3,2,0"&gt;&lt;STRONG data-path-to-node="29,3,2,0" data-index-in-node="0"&gt;Metered Only&lt;/STRONG&gt;&lt;/SPAN&gt;&lt;/td&gt;&lt;td style="height: 34.8px;"&gt;&lt;SPAN data-path-to-node="29,3,3,0"&gt;Only if Stop Usage is &lt;STRONG&gt;ON&lt;/STRONG&gt;&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 34.8px;"&gt;&lt;td style="height: 34.8px;"&gt;&lt;SPAN data-path-to-node="29,4,0,0"&gt;&lt;STRONG data-path-to-node="29,4,0,0" data-index-in-node="0"&gt;Enterprise Budget&lt;/STRONG&gt;&lt;/SPAN&gt;&lt;/td&gt;&lt;td style="height: 34.8px;"&gt;&lt;SPAN data-path-to-node="29,4,1,0"&gt;Enterprise overage spend&lt;/SPAN&gt;&lt;/td&gt;&lt;td style="height: 34.8px;"&gt;&lt;SPAN data-path-to-node="29,4,2,0"&gt;&lt;STRONG data-path-to-node="29,4,2,0" data-index-in-node="0"&gt;Metered Only&lt;/STRONG&gt;&lt;/SPAN&gt;&lt;/td&gt;&lt;td style="height: 34.8px;"&gt;&lt;SPAN data-path-to-node="29,4,3,0"&gt;Only if Stop Usage is &lt;STRONG&gt;ON&lt;/STRONG&gt;&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H4 data-path-to-node="30"&gt;&lt;STRONG data-path-to-node="30" data-index-in-node="0"&gt;The "Lowest Headroom Wins" Rule&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P data-path-to-node="31"&gt;If a developer has $10 remaining on their personal ULB, but the Enterprise Overage Budget only has $1 left before hitting its limit, &lt;STRONG data-path-to-node="31" data-index-in-node="133"&gt;the Enterprise limit will block the user&lt;/STRONG&gt;. Whichever budget hits its ceiling first wins.&lt;/P&gt;
&lt;H4 data-path-to-node="32"&gt;&lt;STRONG data-path-to-node="32" data-index-in-node="0"&gt;What Gets Blocked vs. What Keeps Working?&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P data-path-to-node="33"&gt;When a developer hits a budget limit:&lt;/P&gt;
&lt;UL data-path-to-node="34"&gt;
&lt;LI&gt;❌ &lt;STRONG data-path-to-node="34,0,0" data-index-in-node="2"&gt;Blocked:&lt;/STRONG&gt; Copilot Chat, CLI, Agents, and complex reasoning models (anything consuming AI credits).&lt;/LI&gt;
&lt;LI&gt;✅ &lt;STRONG data-path-to-node="34,1,0" data-index-in-node="2"&gt;Unblocked:&lt;/STRONG&gt; Standard inline code completions and next edit suggestions (included free with seat licenses).&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3 data-path-to-node="36"&gt;&lt;STRONG data-path-to-node="36" data-index-in-node="0"&gt;Conclusion: Best Practices for Admins&lt;/STRONG&gt;&lt;/H3&gt;
&lt;OL data-path-to-node="37"&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="37,0,0" data-index-in-node="0"&gt;Establish a Universal ULB First:&lt;/STRONG&gt; Set a reasonable baseline (e.g., $10–$15/user) to ensure fair pool access.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="37,1,0" data-index-in-node="0"&gt;Use Cost Center ULBs for Group Granularity:&lt;/STRONG&gt; Avoid managing thousands of individual budgets by setting per-user caps at the Cost Center level.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="37,2,0" data-index-in-node="0"&gt;Always Turn On "Stop Usage":&lt;/STRONG&gt; Ensure your enterprise and cost center overage budgets have the hard stop toggle enabled to prevent unexpected monthly bills.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="37,3,0" data-index-in-node="0"&gt;Isolate R&amp;amp;D with Cost Center Exclusion:&lt;/STRONG&gt; Use exclusion rules for teams that require independent spending authority outside the general enterprise cap.&lt;/LI&gt;
&lt;/OL&gt;</description>
      <pubDate>Wed, 29 Jul 2026 17:33:31 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/mastering-github-copilot-budgets-how-to-prevent-surprise/ba-p/4542073</guid>
      <dc:creator>gauravbhardwaj</dc:creator>
      <dc:date>2026-07-29T17:33:31Z</dc:date>
    </item>
    <item>
      <title>Token Economics in Practice</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/token-economics-in-practice/ba-p/4540472</link>
      <description>&lt;H2 data-line="6"&gt;Introduction: The cheap-token trap&lt;/H2&gt;
&lt;P&gt;Token prices alone are a poor economic model for agents. The price of reaching a fixed capability has fallen sharply — In a&amp;nbsp;&lt;A class="lia-external-url" href="https://hai.stanford.edu/assets/files/hai_ai_index_report_2025.pdf" target="_blank" rel="noopener"&gt;2025 Report Stanford's AI Index&lt;/A&gt; reported a roughly 280-fold drop in the cost of GPT-3.5-level inference between late 2022 and late 2024, and Epoch AI tracks steep (if uneven) per-benchmark price declines. The intuitive conclusion is that agents are getting cheaper to run. The operational reality is the opposite.&lt;/P&gt;
&lt;P data-line="10"&gt;Agents turn cheaper inference into &lt;STRONG&gt;longer, stochastic trajectories&lt;/STRONG&gt;: growing context windows, repeated tool schemas, retries, reflection loops, and sub-agent fan-out. In one study of agentic coding, repeated runs of the&amp;nbsp;&lt;EM&gt;same&lt;/EM&gt;&amp;nbsp;agent on the&amp;nbsp;&lt;EM&gt;same&lt;/EM&gt;&amp;nbsp;task varied in token cost by as much as&amp;nbsp;&lt;STRONG&gt;30×&amp;nbsp;&lt;/STRONG&gt;for coding agents. When a single logical task can cost you thirty times more depending on the path the agent takes, optimizing&amp;nbsp;&lt;EM&gt;average cost per token&lt;/EM&gt; will happily make the wrong system look efficient. Similar argument can be made for other agentic systems where we may need more than one tries, more than one MCP Calls, Reasoning or use of multiple skills, hooks or tool calls to arrive at a completed task.&lt;/P&gt;
&lt;P data-line="12"&gt;So, the leading question of token economics isn't "what's the token price?" It's &lt;STRONG&gt;"what does it cost to get one accepted unit of useful work — and how confident can we be in that number before the agent runs?"&lt;/STRONG&gt;&lt;/P&gt;
&lt;H2 data-line="14"&gt;The unit that actually matters:&amp;nbsp; Cost per accepted task&lt;/H2&gt;
&lt;P data-line="16"&gt;I use&amp;nbsp;&lt;STRONG&gt;token economics&lt;/STRONG&gt;&amp;nbsp;to mean managing the unit economics of useful AI work under uncertainty. The meaningful unit is&amp;nbsp;&lt;STRONG&gt;cost per accepted task&lt;/STRONG&gt;, not cost per token.&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P data-line="18"&gt;Let A = 1 mean a task passed its acceptance rubric. The long-run unit cost of a policy π is approximately:&lt;/P&gt;
&lt;img /&gt;&lt;/BLOCKQUOTE&gt;
&lt;P data-line="24"&gt;The numerator is expected task cost; the denominator is the probability the output is actually acceptable. This follows the FinOps distinction between successful and unsuccessful AI outputs and the recommendation to connect cost with workload value. It is a working definition for this project, not a quoted standard — but it reframes the engineering problem immediately. A "cheaper" policy that halves cost while dropping acceptance from 95% to 70% is &lt;STRONG&gt;more expensive per accepted task&lt;/STRONG&gt;, and only this ratio makes that visible.&lt;/P&gt;
&lt;P data-line="26"&gt;That reframing turns "pick the cheapest model" into a five-step discipline:&lt;/P&gt;
&lt;OL data-line="28"&gt;
&lt;LI data-line="28"&gt;&lt;STRONG&gt;Forecast&lt;/STRONG&gt;&amp;nbsp;a distribution, not a single token estimate.&lt;/LI&gt;
&lt;LI data-line="29"&gt;&lt;STRONG&gt;Select&lt;/STRONG&gt;&amp;nbsp;a cost policy that is plausible for the task and its risk.&lt;/LI&gt;
&lt;LI data-line="30"&gt;&lt;STRONG&gt;Enforce&lt;/STRONG&gt;&amp;nbsp;routing, context, cache, and budget controls during execution.&lt;/LI&gt;
&lt;LI data-line="31"&gt;&lt;STRONG&gt;Evaluate&lt;/STRONG&gt;&amp;nbsp;whether the output still clears a workload-specific quality floor.&lt;/LI&gt;
&lt;LI data-line="32"&gt;&lt;STRONG&gt;Revert&lt;/STRONG&gt; unsafe savings, reconcile predicted vs. actual usage, and calibrate the next forecast.&lt;/LI&gt;
&lt;/OL&gt;
&lt;img /&gt;
&lt;H2 data-line="40"&gt;From a metric to a controller&lt;/H2&gt;
&lt;P data-line="42"&gt;If cost is a random variable, the objective is a stochastic one. Minimize&amp;nbsp;&lt;EM&gt;expected&lt;/EM&gt; task cost subject to two constraints — a quality floor on every workload segment, and a bound on how often you blow the budget&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;subject to a per-segment quality floor:&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;and a chance constraint on budget breach:&lt;/P&gt;
&lt;img /&gt;
&lt;P data-line="42"&gt;&amp;nbsp;&lt;/P&gt;
&lt;P data-line="60"&gt;Here π is the policy; C_task is total task cost; Q_s is quality for a supported segment s with floor Q_min; B is the budget; and ε is the tolerated breach probability. The pieces are all borrowed — stochastic optimization for the expected-cost objective; FrugalGPT and Confident Adaptive Language Modeling for the LLM precedent of cutting cost while preserving performance; SRE service-level objectives for treating "acceptable service" as an action-driving threshold and Group DRO for the insight that &lt;EM&gt;averages hide group failures; and&lt;/EM&gt;&amp;nbsp;Charnes–Cooper chance-constrained programming for the probabilistic budget limit. &lt;STRONG&gt;The synthesis — wiring them into one agent controller — is the contribution&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P data-line="62"&gt;Two honest caveats travel with this controller:&amp;nbsp;Q_s&amp;nbsp;needs a confidence-adjusted lower bound (sparse segments shouldn't trigger changes on two samples), and the chance constraint is&amp;nbsp;&lt;STRONG&gt;not a guarantee&lt;/STRONG&gt; until your forecast's percentile coverage is calibrated against real traces. A modeled P95 is a planning estimate, not a promised 5% breach bound.&lt;/P&gt;
&lt;H2 data-line="64"&gt;Two halves of the loop: feed-forward and feedback&lt;/H2&gt;
&lt;P data-line="66"&gt;The current work is result of two self-prototypes — &lt;STRONG&gt;FutureTokenPredictor&lt;/STRONG&gt;&amp;nbsp;and&amp;nbsp;&lt;STRONG&gt;TokenGov&lt;/STRONG&gt; — built to make agent unit economics operable on Azure. These are reusable implementation patterns and experiments.&lt;/P&gt;
&lt;P data-line="66"&gt;The controller splits cleanly into a planning half and a runtime half.&lt;/P&gt;
&lt;UL data-line="68"&gt;
&lt;LI data-line="68"&gt;&lt;STRONG&gt;FutureTokenPredictor is the feed-forward side.&lt;/STRONG&gt;&amp;nbsp;It models workflow archetypes and uncertain iteration counts to produce P50/P95-style planning estimates&amp;nbsp;&lt;EM&gt;before execution and&lt;/EM&gt; recommends a policy. It stays&amp;nbsp;&lt;STRONG&gt;outside the request path&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI data-line="69"&gt;&lt;STRONG&gt;TokenGov is the feedback side.&lt;/STRONG&gt; Its request path applies the admitted cost policy; an out-of-band control plane evaluates outcomes and changes externalized policy when quality regresses. Runtime telemetry then flows back to the predictor as calibration data for the next forecast.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="71"&gt;Neither half is sufficient alone. Prediction without control is a spreadsheet. Control without quality feedback silently degrades your hardest segments. The value is the&amp;nbsp;&lt;STRONG&gt;wire between them&lt;/STRONG&gt;: a forecast that becomes an enforceable policy, an eval verdict that can reverse a cost action, and actuals that sharpen the next forecast.&lt;/P&gt;
&lt;H2 data-line="73"&gt;How the equation lands on Azure&lt;/H2&gt;
&lt;P data-line="75"&gt;This is where token economics stops being a metric and becomes architecture. Each term in the controller maps to a concrete Azure control:&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Controller term&lt;/th&gt;&lt;th&gt;Azure control in practice&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;π&amp;nbsp;(policy)&lt;/td&gt;&lt;td&gt;Externalized in&amp;nbsp;&lt;STRONG&gt;Azure App Configuration&lt;/STRONG&gt;; enforced by&amp;nbsp;&lt;STRONG&gt;API Management&lt;/STRONG&gt;&amp;nbsp;GenAI gateway (routing, context, cache, token policies)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;E[C_task | π]&amp;nbsp;&lt;/P&gt;
&lt;P&gt;(expected cost)&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;Reconstructed from APIM gateway, model, and&amp;nbsp;&lt;STRONG&gt;Application Insights&lt;/STRONG&gt;&amp;nbsp;telemetry&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Q_s&amp;nbsp;(segment quality)&lt;/td&gt;&lt;td&gt;&lt;STRONG&gt;Azure AI Foundry evaluation&lt;/STRONG&gt;&amp;nbsp;over golden sets and sampled production traces&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;B,&amp;nbsp;ε&amp;nbsp;(budget, breach tolerance)&lt;/td&gt;&lt;td&gt;Forecast-informed limits and&amp;nbsp;&lt;STRONG&gt;Azure Monitor&lt;/STRONG&gt;&amp;nbsp;alerts;&amp;nbsp;&lt;STRONG&gt;Cost Management&lt;/STRONG&gt;&amp;nbsp;for allocation&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Reversion&lt;/td&gt;&lt;td&gt;A&amp;nbsp;&lt;STRONG&gt;Monitor-triggered Azure Function&lt;/STRONG&gt;&amp;nbsp;tightens or reverts policy in App Configuration — closing the eval-to-enforcement loop&amp;nbsp;&lt;STRONG&gt;without a code deployment&lt;/STRONG&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P data-line="85"&gt;Most of these primitives already exist and are individually documented: APIM provides token quotas, semantic caching, and token metrics; Foundry Model Router offers cost/balanced/quality routing modes; Foundry cloud evaluation scores datasets and sampled traces. The interesting gap they&amp;nbsp;&lt;EM&gt;don't&lt;/EM&gt;&amp;nbsp;close on their own is the connected mechanism — an evaluation verdict that can constrain or reverse a cost-saving action, and actual usage that improves the next forecast.&lt;/P&gt;
&lt;P data-line="87"&gt;Here is the full two-plane view. FutureTokenPredictor forecasts and recommends&amp;nbsp;&lt;EM&gt;before&lt;/EM&gt; execution; TokenGov owns runtime enforcement and quality-triggered reversion; prediction IDs join forecasts to actual telemetry so calibration can improve the next estimate.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;From Concept to Implementation&amp;nbsp;&lt;/H2&gt;
&lt;P&gt;&lt;STRONG&gt;Version 1&lt;/STRONG&gt; release the &lt;STRONG&gt;Token Prediction and forecast ability&lt;/STRONG&gt; using a local mcp server called FutureTokenPredictor using a local MCP server modeled behind a simple UI, where you can create an assessment for your UI Workload. It lets you simple describe the AI / Agentic Solution you want to build and suggested a topology for it. From there , depending on your model selection, the studio, helps you predict the range of token usage and its estimated costs.&lt;/P&gt;
&lt;P&gt;In full version, this forecast is used to build a policy and govern your AI Spend accordingly.&lt;/P&gt;
&lt;P&gt;If you want to read more about the FutureTokenPredictor and how it works, check out my earlier blog&amp;nbsp;&lt;A class="lia-external-url" href="https://www.linkedin.com/pulse/agentic-currency-tokens-ai-infra-full-stack-cost-agents-dhiman-rqd8e/" target="_blank" rel="noopener"&gt;Agentic Currency – Tokens and AI Infra: Full-Stack Cost Prediction for Autonomous Agents&lt;/A&gt;&lt;/P&gt;
&lt;img /&gt;&lt;img /&gt;&lt;img /&gt;
&lt;P&gt;Version 2 with full governance and control will be released soon.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;TokenEconomics is available in the GitHub Repo &lt;A href="https://github.com/pd-illinois/TokenEconomics.git" target="_blank"&gt;TokenEconomics&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;Clone it, experiment and test it out. Please provide feedback via a pull request on the repo or directly here via comments&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Happy Reading!&lt;/P&gt;
&lt;H5&gt;References&amp;nbsp;&lt;/H5&gt;
&lt;P&gt;&lt;A href="https://hai.stanford.edu/ai-index/2025-ai-index-report" target="_blank" rel="noopener"&gt;The 2025 AI Index Report | Stanford HAI&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://www.jstor.org/stable/2627476" target="_blank" rel="noopener"&gt;Chance-Constrained Programming | JSTOR&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://digitaleconomy.stanford.edu/news/how-are-ai-agents-spending-your-tokens/" target="_blank" rel="noopener"&gt;How are AI agents spending your tokens? - Stanford Digital Economy Lab&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://www.finops.org/wg/finops-for-ai-overview/" target="_blank" rel="noopener"&gt;FinOps for AI Overview&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/api-management/genai-gateway-capabilities" target="_blank" rel="noopener"&gt;AI gateway capabilities in Azure API Management | Microsoft Learn&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-router" target="_blank" rel="noopener"&gt;Model router for Microsoft Foundry concepts - Microsoft Foundry | Microsoft Learn&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Also Read&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://techcommunity.microsoft.com/blog/AzureArchitectureBlog/optimizing-github-copilot-cost-in-the-usage-based-billing-era/4534171?previewMessage=true" target="_blank" rel="noopener"&gt;Optimizing GitHub Copilot Cost in the Usage-Based Billing Era | Microsoft Community Hub&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://techcommunity.microsoft.com/blog/azuredevcommunityblog/token-economics-the-new-finops-for-agentic-ai/4533743" target="_blank" rel="noopener"&gt;Token Economics: The New FinOps for Agentic AI | Microsoft Community Hub&lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 28 Jul 2026 17:57:41 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/token-economics-in-practice/ba-p/4540472</guid>
      <dc:creator>prateekwrites</dc:creator>
      <dc:date>2026-07-28T17:57:41Z</dc:date>
    </item>
    <item>
      <title>Preserve a Legacy IP During Azure Migration with Private Link Service Direct Connect</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/preserve-a-legacy-ip-during-azure-migration-with-private-link/ba-p/4540575</link>
      <description>&lt;H1&gt;Why this matters&lt;/H1&gt;
&lt;P&gt;Phased migrations to Azure frequently stall on a single, unglamorous constraint: an application with a hard-coded dependency address. The usual options are to re-address the application or change its configuration, and both add development effort, regression testing, and risk to a migration that is already time-boxed.&lt;/P&gt;
&lt;P&gt;This article demonstrates a pattern that removes that constraint from the critical path. Using Azure Private Link Service Direct Connect together with a Private Endpoint, the migrated application continues to reach its dependency on exactly the same private address it has always used, while that dependency remains on premises during the interim phase.&lt;/P&gt;
&lt;P&gt;The practical benefits are:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Reduced migration risk. No application re-addressing or code change, so no regression cycle and a significantly smaller blast radius.&lt;/LI&gt;
&lt;LI&gt;Faster migration. Web and application tiers can move now rather than waiting for the backend dependency to be modernised, so phase one is no longer blocked by phase two.&lt;/LI&gt;
&lt;LI&gt;Avoided remediation cost. No development effort to change hard-coded addressing, and none of the associated testing spend.&lt;/LI&gt;
&lt;LI&gt;Business continuity preserved. The application connects to the same address throughout, which removes addressing-related cutover risk mid-migration.&lt;/LI&gt;
&lt;LI&gt;Overlapping address space handled. One of the most common reasons enterprise datacentre exits stall is addressed directly, without VNet peering.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The pattern is deliberately a bridge rather than a destination. It buys time to modernise properly; it does not remove the need to do so.&lt;/P&gt;
&lt;H1&gt;Customer challenge&lt;/H1&gt;
&lt;P&gt;The customer was planning a phased migration of a legacy application to Azure. The web and application tiers could move first, but a dependent database service needed to remain on premises during the initial phase. The application used a hard-coded private address for this dependency, and changing the application configuration or addressing would have increased migration risk and required additional development and regression testing.&lt;/P&gt;
&lt;H1&gt;Customer requirements&lt;/H1&gt;
&lt;P&gt;The customer needed transitional architecture that would allow the web and application tiers to run in Azure while continuing to reach the on-premises database through its existing application-facing address. The solution is needed to:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Preserve the existing application-facing database address.&lt;/LI&gt;
&lt;LI&gt;Migrate the web and application tiers to Azure without immediate application code or configuration changes.&lt;/LI&gt;
&lt;LI&gt;Keep the database on-premises until a later migration phase.&lt;/LI&gt;
&lt;LI&gt;Provide private, controlled connectivity between the Azure-hosted application tiers and the on-premises database.&lt;/LI&gt;
&lt;LI&gt;Validate the complete traffic path before considering the pattern for production use.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H1&gt;Solution and lab overview&lt;/H1&gt;
&lt;P&gt;This lab evaluates Azure Private Link Service Direct Connect as a transitional connectivity pattern for a phased migration. The capability can connect a Private Link service directly to a privately routable destination address without requiring the destination to sit behind a traditional load balancer. That makes it relevant to direct IP-based scenarios such as database connections, custom applications, and privately reachable on-premises resources.&lt;/P&gt;
&lt;P&gt;In the lab, the migrated web and application virtual machines connect to a Private Endpoint in an extended prod VNet. The endpoint maps to a Private Link Service Direct Connect resource in a hub VNet. The hub then routes the traffic across a site-to-site VPN to the simulated on-premises database server. The detailed IP values used to reproduce the lab are introduced in the architecture and deployment sections rather than in the customer narrative.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 99.2593%; border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Preview notice&lt;/STRONG&gt;&lt;BR /&gt;Private Link Service Direct Connect is in public preview and available only in selected regions. Review the documented requirements, limitations, and regional availability before enabling it in a subscription or using it in a production design.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 100.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H1&gt;Why this pattern works&lt;/H1&gt;
&lt;UL&gt;
&lt;LI&gt;Legacy applications often depend on fixed private addresses that cannot be changed safely during the first migration wave.&lt;/LI&gt;
&lt;LI&gt;Reusing or overlapping address space complicates conventional routed hybrid connectivity.&lt;/LI&gt;
&lt;LI&gt;Private Link Service Direct Connect can route traffic to a specific privately routable destination address and bypass the traditional load-balancer requirement.&lt;/LI&gt;
&lt;LI&gt;The pattern provides a temporary abstraction that lets the Azure-hosted application tier keep using its expected destination while the actual dependency remains on-premises.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H1&gt;When this pattern fits&lt;/H1&gt;
&lt;P&gt;Consider this pattern when:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;An application has a hard-coded or otherwise fixed dependency address that cannot be changed in the current migration phase.&lt;/LI&gt;
&lt;LI&gt;The application tier is ready to migrate ahead of its backend dependency.&lt;/LI&gt;
&lt;LI&gt;On-premises and Azure address spaces overlap, which rules out straightforward VNet peering.&lt;/LI&gt;
&lt;LI&gt;You need a controlled, reversible transition step rather than a single large cutover.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Look at other options when:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;The application can be reconfigured or modernised within the migration window. Where that is possible, resolving the dependency properly is the better outcome.&lt;/LI&gt;
&lt;LI&gt;You need a permanent end-state architecture. This pattern is transitional by design.&lt;/LI&gt;
&lt;LI&gt;Your target region does not yet support Private Link Service Direct Connect, which is in public preview with limited regional availability.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H1&gt;Architecture&lt;/H1&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;Figure 1. Logical architecture and traffic path&lt;/SPAN&gt;&lt;/DIV&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;/DIV&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Component&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Address space / IP&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Location&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Purpose&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Hub VNet&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;10.230.0.0/21&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;West US 2&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Hosts the VPN Gateway, Azure Bastion, test subnet, and Private Link Service subnet.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;On-premises VNet&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;10.220.0.0/21&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;North Central US&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Simulates the on-premises network with production and development subnets.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;On-premises production subnet&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;10.220.1.0/24&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;North Central US&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Hosts vm-db at 10.220.1.6. IIS on this VM simulates the test service on TCP port 80.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Extended production VNet&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;10.220.1.0/24&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;West US 2&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Hosts the migrated web and application VMs that retain the original application-facing address space.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Private Link Service subnet&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;10.230.3.0/24&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Hub VNet&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Hosts the Private Link Service NAT IP configurations 10.230.3.10 and 10.230.3.11.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Private Endpoint&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;10.220.1.6&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Extended production subnet&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Provides the local endpoint used by the migrated application tiers to reach the on-premises dependency.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H2&gt;Traffic flow&lt;/H2&gt;
&lt;P&gt;&lt;STRONG&gt;Testing note:&lt;/STRONG&gt; In this lab, &lt;STRONG&gt;vm-db&lt;/STRONG&gt; represents the on-premises database server. To validate the end-to-end network path without deploying a database engine, we install IIS on vm-db and use its web server on TCP port 80 as a lightweight test service. A successful HTTP response confirms that traffic reaches vm-db through the intended private connectivity path; it does not validate database functionality.&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;The web or application VM in the extended production subnet connects to &lt;STRONG&gt;10.220.1.6&lt;/STRONG&gt; on TCP port 80. For this test, that port is served by IIS running on vm-db.&lt;/LI&gt;
&lt;LI&gt;Because 10.220.1.6 is assigned to the Private Endpoint in the same subnet, the connection is sent to the Private Endpoint locally.&lt;/LI&gt;
&lt;LI&gt;The Private Endpoint connection maps to the Private Link Service Direct Connect resource in the hub VNet.&lt;/LI&gt;
&lt;LI&gt;The Private Link Service Direct Connect resource uses the destination IP address 10.220.1.6.&lt;/LI&gt;
&lt;LI&gt;The hub VNet routes 10.220.0.0/21 through the VPN Gateway to the simulated on-premises network.&lt;/LI&gt;
&lt;LI&gt;The simulated on-premises vm-db receives the request through IIS, and the HTTP response returns through the same private connectivity path.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H1&gt;Prerequisites and assumptions&lt;/H1&gt;
&lt;UL&gt;
&lt;LI&gt;An Azure subscription where the Microsoft.Network/AllowPrivateLinkserviceUDR feature flag can be registered.&lt;/LI&gt;
&lt;LI&gt;Azure CLI and Az PowerShell modules available in the execution environment. The commands below assume PowerShell syntax for variables.&lt;/LI&gt;
&lt;LI&gt;Permission to create resource groups, VNets, subnets, NSGs, public IPs, VPN gateways, local network gateways, Private Link Service, Private Endpoint, NICs, and VMs.&lt;/LI&gt;
&lt;LI&gt;A region supported by Private Link Service Direct Connect. This lab uses West US 2 and North Central US.&lt;/LI&gt;
&lt;LI&gt;This lab intentionally uses overlapping application-facing IP space in the extended Azure VNet and the on-premises production subnet. Do not peer the hub VNet and the extended production VNet directly.&lt;/LI&gt;
&lt;/UL&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 97.963%; border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Credential hygiene&lt;/STRONG&gt;&lt;BR /&gt;The original lab used a sample shared key and local administrator password. For publication, replace all secrets with strong unique values or Key Vault-backed automation. This document uses placeholders only.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 100.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H1&gt;Deploy the lab&lt;/H1&gt;
&lt;P&gt;Run the following sequence from an authenticated PowerShell session with Azure CLI available. Review every variable before execution. The lab uses IIS on the simulated database VM only as a lightweight TCP and HTTP target; the objective is to validate the network path rather than database behavior.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;# -----------------------------&lt;BR /&gt;# Variables&lt;BR /&gt;# -----------------------------&lt;BR /&gt;$RG="rg-pls-demo"&lt;BR /&gt;$LOCATION="westus2"&lt;BR /&gt;$VNET="vnet-hub"&lt;BR /&gt;$VNET_PREFIX="10.230.0.0/21"&lt;BR /&gt;$GW_SUBNET="GatewaySubnet"&lt;BR /&gt;$GW_PREFIX="10.230.1.0/24"&lt;BR /&gt;$BASTION_SUBNET="AzureBastionSubnet"&lt;BR /&gt;$BASTION_PREFIX="10.230.2.0/24"&lt;BR /&gt;$PLS_SUBNET="snet-pls"&lt;BR /&gt;$PLS_PREFIX="10.230.3.0/24"&lt;BR /&gt;$TEST_SUBNET="snet-test"&lt;BR /&gt;$TEST_PREFIX="10.230.4.0/24"&lt;BR /&gt;$ONPREM_RG="onprem-group"&lt;BR /&gt;$ONPREM_LOCATION="northcentralus"&lt;BR /&gt;$ONPREM_VNET="vnet-onprem"&lt;BR /&gt;$ONPREM_VNET_PREFIX="10.220.0.0/21"&lt;BR /&gt;$ONPREM_GW_SUBNET="GatewaySubnet"&lt;BR /&gt;$ONPREM_GW_PREFIX="10.220.0.0/24"&lt;BR /&gt;$ONPREM_PROD_SUBNET="snet-prod"&lt;BR /&gt;$ONPREM_PROD_PREFIX="10.220.1.0/24"&lt;BR /&gt;$ONPREM_DEV_SUBNET="snet-dev"&lt;BR /&gt;$ONPREM_DEV_PREFIX="10.220.2.0/24"&lt;BR /&gt;$ONPREM_DB_IP="10.220.1.6"&lt;BR /&gt;$VNET_EXTENDED="Extended-VNET"&lt;BR /&gt;$VNET_PREFIX_EXTENDED="10.220.1.0/24"&lt;BR /&gt;$EXT_SUBNET="snet-extended-prod"&lt;BR /&gt;$EXT_PREFIX="10.220.1.0/24"&lt;BR /&gt;$PLS_NAME="pls-directconnect-db"&lt;BR /&gt;$PLS_IP1="10.230.3.10"&lt;BR /&gt;$PLS_IP2="10.230.3.11"&lt;BR /&gt;$DESTINATION_IP=$ONPREM_DB_IP&lt;BR /&gt;$PE_NAME="pe-db"&lt;BR /&gt;$PE_CONNECTION_NAME="pe-to-pls"&lt;BR /&gt;$PLS_JSON="pls-ipconfigs.json"&lt;BR /&gt;$PE_JSON="pe-ipconfig.json"&lt;BR /&gt;$LNG_ONPREM="lng-onprem"&lt;BR /&gt;$LNG_AZURE="lng-azure"&lt;BR /&gt;$PIP_ONPREM_VPNGW="pip-onprem-vpngw"&lt;BR /&gt;$ONPREM_VPNGW="onprem-vpngw"&lt;BR /&gt;$PIP_VPNGW="pip-vpngw"&lt;BR /&gt;$VPNGW_HUB="vpngw-hub"&lt;BR /&gt;$SHAREDKEY="&amp;lt;replace-with-strong-shared-key&amp;gt;"&lt;BR /&gt;$ADMINUSER="adminuser"&lt;BR /&gt;$ADMIN_PASSWORD="&amp;lt;replace-with-strong-password&amp;gt;"&lt;BR /&gt;&lt;BR /&gt;# -----------------------------&lt;BR /&gt;# Create hub resource group, VNet, NSGs, and subnets&lt;BR /&gt;# -----------------------------&lt;BR /&gt;az group create --name $RG --location $LOCATION&lt;BR /&gt;az network vnet create --resource-group $RG --name $VNET --location $LOCATION --address-prefixes $VNET_PREFIX&lt;BR /&gt;az network nsg create -g $RG -n nsg-gateway -l $LOCATION&lt;BR /&gt;az network nsg create -g $RG -n nsg-bastion -l $LOCATION&lt;BR /&gt;az network nsg create -g $RG -n nsg-pls -l $LOCATION&lt;BR /&gt;az network nsg create -g $RG -n nsg-test -l $LOCATION&lt;BR /&gt;az network vnet subnet create -g $RG --vnet-name $VNET -n $GW_SUBNET --address-prefixes $GW_PREFIX&lt;BR /&gt;az network vnet subnet create -g $RG --vnet-name $VNET -n $BASTION_SUBNET --address-prefixes $BASTION_PREFIX&lt;BR /&gt;az network vnet subnet create -g $RG --vnet-name $VNET -n $PLS_SUBNET --address-prefixes $PLS_PREFIX --network-security-group nsg-pls --disable-private-link-service-network-policies true&lt;BR /&gt;az network vnet subnet create -g $RG --vnet-name $VNET -n $TEST_SUBNET --address-prefixes $TEST_PREFIX --network-security-group nsg-test&lt;BR /&gt;&lt;BR /&gt;# -----------------------------&lt;BR /&gt;# Configure Azure Bastion NSG rules and associate with the Azure bastion Subnet&lt;BR /&gt;# -----------------------------&lt;BR /&gt;$NSG="nsg-bastion"&lt;BR /&gt;$nsg=Get-AzNetworkSecurityGroup -ResourceGroupName $RG -Name $NSG&lt;BR /&gt;$nsg.SecurityRules.Add((New-AzNetworkSecurityRuleConfig -Name "AllowHttpsInbound" -Priority 120 -Direction Inbound -Access Allow -Protocol Tcp -SourceAddressPrefix Internet -SourcePortRange * -DestinationAddressPrefix * -DestinationPortRange 443))&lt;BR /&gt;$nsg.SecurityRules.Add((New-AzNetworkSecurityRuleConfig -Name "AllowGatewayManagerInbound" -Priority 130 -Direction Inbound -Access Allow -Protocol Tcp -SourceAddressPrefix GatewayManager -SourcePortRange * -DestinationAddressPrefix * -DestinationPortRange 443))&lt;BR /&gt;$nsg.SecurityRules.Add((New-AzNetworkSecurityRuleConfig -Name "AllowAzureLoadBalancerInbound" -Priority 140 -Direction Inbound -Access Allow -Protocol Tcp -SourceAddressPrefix AzureLoadBalancer -SourcePortRange * -DestinationAddressPrefix * -DestinationPortRange 443))&lt;BR /&gt;$nsg.SecurityRules.Add((New-AzNetworkSecurityRuleConfig -Name "AllowBastionHostCommunication" -Priority 150 -Direction Inbound -Access Allow -Protocol Tcp -SourceAddressPrefix VirtualNetwork -SourcePortRange * -DestinationAddressPrefix VirtualNetwork -DestinationPortRange 8080,5701))&lt;BR /&gt;$nsg.SecurityRules.Add((New-AzNetworkSecurityRuleConfig -Name "AllowSshRdpOutbound" -Priority 100 -Direction Outbound -Access Allow -Protocol Tcp -SourceAddressPrefix * -SourcePortRange * -DestinationAddressPrefix VirtualNetwork -DestinationPortRange 22,3389))&lt;BR /&gt;$nsg.SecurityRules.Add((New-AzNetworkSecurityRuleConfig -Name "AllowAzureCloudOutbound" -Priority 110 -Direction Outbound -Access Allow -Protocol Tcp -SourceAddressPrefix * -SourcePortRange * -DestinationAddressPrefix AzureCloud -DestinationPortRange 443))&lt;BR /&gt;$nsg.SecurityRules.Add((New-AzNetworkSecurityRuleConfig -Name "AllowBastionCommunicationOutbound" -Priority 120 -Direction Outbound -Access Allow -Protocol Tcp -SourceAddressPrefix VirtualNetwork -SourcePortRange * -DestinationAddressPrefix VirtualNetwork -DestinationPortRange 8080,5701))&lt;BR /&gt;$nsg.SecurityRules.Add((New-AzNetworkSecurityRuleConfig -Name "AllowHttpOutbound" -Priority 130 -Direction Outbound -Access Allow -Protocol Tcp -SourceAddressPrefix * -SourcePortRange * -DestinationAddressPrefix Internet -DestinationPortRange 80))&lt;BR /&gt;Set-AzNetworkSecurityGroup -NetworkSecurityGroup $nsg&lt;BR /&gt;az network vnet subnet update -g $RG --vnet-name $VNET -n $BASTION_SUBNET --network-security-group nsg-bastion&lt;BR /&gt;&lt;BR /&gt;# -----------------------------&lt;BR /&gt;# Deploy Azure Bastion and hub VPN Gateway&lt;BR /&gt;# -----------------------------&lt;BR /&gt;az network public-ip create -g $RG -n pip-bastion -l $LOCATION --sku Standard --zone 1 2 3&lt;BR /&gt;az network bastion create -g $RG -n bastion-hub --public-ip-address pip-bastion --vnet-name $VNET -l $LOCATION --sku Standard&lt;BR /&gt;az network public-ip create -g $RG -n $PIP_VPNGW -l $LOCATION --sku Standard --zone 1 2 3&lt;BR /&gt;az network vnet-gateway create -g $RG -n $VPNGW_HUB --public-ip-addresses $PIP_VPNGW --vnet $VNET --gateway-type Vpn --vpn-type RouteBased --sku VpnGw2AZ&lt;BR /&gt;&lt;BR /&gt;# -----------------------------&lt;BR /&gt;# Create extended Azure prod VNet and migrated web/app VMs(+ vm-test in hub VNET)&lt;BR /&gt;# -----------------------------&lt;BR /&gt;az network vnet create --resource-group $RG --name $VNET_EXTENDED --location $LOCATION --address-prefixes $VNET_PREFIX_EXTENDED&lt;BR /&gt;az network nsg create -g $RG -n nsg-extended -l $LOCATION&lt;BR /&gt;az network vnet subnet create --resource-group $RG --vnet-name $VNET_EXTENDED --name $EXT_SUBNET --address-prefixes $EXT_PREFIX --network-security-group nsg-extended&lt;BR /&gt;az network nic create -g $RG -n nic-testvm --vnet-name $VNET --subnet $TEST_SUBNET&lt;BR /&gt;az vm create -g $RG -n vm-test --nics nic-testvm --image Win2022Datacenter --admin-username $ADMINUSER --admin-password $ADMIN_PASSWORD&lt;BR /&gt;az network nic create -g $RG -n nic-web --vnet-name $VNET_EXTENDED --subnet $EXT_SUBNET --private-ip-address 10.220.1.4&lt;BR /&gt;az vm create -g $RG -n vm-web --nics nic-web --image Win2022Datacenter --admin-username $ADMINUSER --admin-password $ADMIN_PASSWORD&lt;BR /&gt;az network nic create -g $RG -n nic-app --vnet-name $VNET_EXTENDED --subnet $EXT_SUBNET --private-ip-address 10.220.1.5&lt;BR /&gt;az vm create -g $RG -n vm-app --nics nic-app --image Win2022Datacenter --admin-username $ADMINUSER --admin-password $ADMIN_PASSWORD&lt;BR /&gt;&lt;BR /&gt;# -----------------------------&lt;BR /&gt;# Create simulated on-premises VNet and VPN Gateway&lt;BR /&gt;# -----------------------------&lt;BR /&gt;az group create --name $ONPREM_RG --location $ONPREM_LOCATION&lt;BR /&gt;az network vnet create --resource-group $ONPREM_RG --name $ONPREM_VNET --location $ONPREM_LOCATION --address-prefixes $ONPREM_VNET_PREFIX&lt;BR /&gt;az network vnet subnet create --resource-group $ONPREM_RG --vnet-name $ONPREM_VNET --name $ONPREM_GW_SUBNET --address-prefixes $ONPREM_GW_PREFIX&lt;BR /&gt;az network vnet subnet create --resource-group $ONPREM_RG --vnet-name $ONPREM_VNET --name $ONPREM_PROD_SUBNET --address-prefixes $ONPREM_PROD_PREFIX&lt;BR /&gt;az network vnet subnet create --resource-group $ONPREM_RG --vnet-name $ONPREM_VNET --name $ONPREM_DEV_SUBNET --address-prefixes $ONPREM_DEV_PREFIX&lt;BR /&gt;az network public-ip create --resource-group $ONPREM_RG --name $PIP_ONPREM_VPNGW --location $ONPREM_LOCATION --sku Standard&lt;BR /&gt;az network vnet-gateway create --resource-group $ONPREM_RG --name $ONPREM_VPNGW --public-ip-addresses $PIP_ONPREM_VPNGW --vnet $ONPREM_VNET --gateway-type Vpn --vpn-type RouteBased --sku VpnGw2AZ&lt;BR /&gt;&lt;BR /&gt;# -----------------------------&lt;BR /&gt;# Create local network gateways and VPN connections&lt;BR /&gt;# -----------------------------&lt;BR /&gt;$AZURE_VPN_PUBLICIP = az network public-ip show -g $RG -n $PIP_VPNGW --query ipAddress -o tsv&lt;BR /&gt;$ONPREM_VPN_PUBLICIP = az network public-ip show -g $ONPREM_RG -n $PIP_ONPREM_VPNGW --query ipAddress -o tsv&lt;BR /&gt;Write-Host "Azure VPN IP&amp;nbsp; : $AZURE_VPN_PUBLICIP"&lt;BR /&gt;Write-Host "OnPrem VPN IP : $ONPREM_VPN_PUBLICIP"&lt;BR /&gt;az network local-gateway create -g $RG -n $LNG_ONPREM --gateway-ip-address $ONPREM_VPN_PUBLICIP --local-address-prefixes 10.220.0.0/21&lt;BR /&gt;az network local-gateway create -g $ONPREM_RG -n $LNG_AZURE --gateway-ip-address $AZURE_VPN_PUBLICIP --local-address-prefixes 10.230.0.0/21&lt;BR /&gt;az network vpn-connection create -g $RG -n conn-to-onprem --vnet-gateway1 $VPNGW_HUB --local-gateway2 $LNG_ONPREM --shared-key $SHAREDKEY&lt;BR /&gt;az network vpn-connection create -g $ONPREM_RG -n conn-to-azure --vnet-gateway1 $ONPREM_VPNGW --local-gateway2 $LNG_AZURE --shared-key $SHAREDKEY&lt;BR /&gt;az network vpn-connection show -g $RG -n conn-to-onprem --query "{Name:name,Status:connectionStatus}" -o table&lt;BR /&gt;az network vpn-connection show -g $ONPREM_RG -n conn-to-azure --query "{Name:name,Status:connectionStatus}" -o table&lt;BR /&gt;&lt;BR /&gt;# -----------------------------&lt;BR /&gt;# Create on-premises DB VM&lt;BR /&gt;# -----------------------------&lt;BR /&gt;az network nic create -g $ONPREM_RG -n nic-db --vnet-name $ONPREM_VNET --subnet $ONPREM_PROD_SUBNET --private-ip-address $ONPREM_DB_IP&lt;BR /&gt;az vm create -g $ONPREM_RG -n vm-db --nics nic-db --image Win2022Datacenter --admin-username $ADMINUSER --admin-password $ADMIN_PASSWORD&lt;BR /&gt;&lt;BR /&gt;# -----------------------------&lt;BR /&gt;# Register PLS Direct Connect preview feature&lt;BR /&gt;# -----------------------------&lt;BR /&gt;az feature register --namespace Microsoft.Network --name AllowPrivateLinkserviceUDR&lt;BR /&gt;az feature show --namespace Microsoft.Network --name AllowPrivateLinkserviceUDR --query properties.state -o tsv&lt;BR /&gt;az provider register --namespace Microsoft.Network&lt;BR /&gt;&lt;BR /&gt;# -----------------------------&lt;BR /&gt;# Create PLS Direct Connect IP configuration JSON&lt;BR /&gt;# -----------------------------&lt;BR /&gt;$PLS_SUBNET_ID = az network vnet subnet show -g $RG --vnet-name $VNET -n $PLS_SUBNET --query id -o tsv&lt;BR /&gt;@"&lt;BR /&gt;[&lt;BR /&gt;&amp;nbsp; {&lt;BR /&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; "name": "ipconfig1",&lt;BR /&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; "primary": true,&lt;BR /&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; "private-ip-allocation-method": "Static",&lt;BR /&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; "private-ip-address": "$PLS_IP1",&lt;BR /&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; "subnet": { "id": "$PLS_SUBNET_ID" }&lt;BR /&gt;&amp;nbsp; },&lt;BR /&gt;&amp;nbsp; {&lt;BR /&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; "name": "ipconfig2",&lt;BR /&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; "primary": false,&lt;BR /&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; "private-ip-allocation-method": "Static",&lt;BR /&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; "private-ip-address": "$PLS_IP2",&lt;BR /&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; "subnet": { "id": "$PLS_SUBNET_ID" }&lt;BR /&gt;&amp;nbsp; }&lt;BR /&gt;]&lt;BR /&gt;"@ | Out-File -FilePath $PLS_JSON -Encoding ascii&lt;BR /&gt;&lt;BR /&gt;az network private-link-service create -g $RG -n $PLS_NAME -l $LOCATION --destination-ip-address $DESTINATION_IP --ip-configurations "@$PLS_JSON"&lt;BR /&gt;az network private-link-service show -g $RG -n $PLS_NAME --query "name:name,provisioningState:provisioningState,destinationIpAddress:destinationIPAddress}" -o table&lt;BR /&gt;&lt;BR /&gt;# -----------------------------&lt;BR /&gt;# Create Private Endpoint in the extended Azure production subnet&lt;BR /&gt;# -----------------------------&lt;BR /&gt;az network vnet subnet update -g $RG --vnet-name $VNET_EXTENDED -n $EXT_SUBNET --private-endpoint-network-policies Disabled&lt;BR /&gt;$PLS_ID = az network private-link-service show -g $RG -n $PLS_NAME --query id -o tsv&lt;BR /&gt;@"&lt;BR /&gt;[&lt;BR /&gt;&amp;nbsp; {&lt;BR /&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; "name": "pe-ipconfig1",&lt;BR /&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; "private-ip-address": "$ONPREM_DB_IP"&lt;BR /&gt;&amp;nbsp; }&lt;BR /&gt;]&lt;BR /&gt;"@ | Out-File -FilePath $PE_JSON -Encoding ascii&lt;BR /&gt;&lt;BR /&gt;az network private-endpoint create -g $RG -n $PE_NAME -l $LOCATION --vnet-name $VNET_EXTENDED --subnet $EXT_SUBNET --private-connection-resource-id $PLS_ID --connection-name $PE_CONNECTION_NAME --ip-configs "@$PE_JSON"&lt;BR /&gt;az network private-link-service show -g $RG -n $PLS_NAME --query privateEndpointConnections -o table&lt;BR /&gt;&lt;BR /&gt;# -----------------------------&lt;BR /&gt;# Install IIS on the on-premises DB VM and publish a test page&lt;BR /&gt;# -----------------------------&lt;BR /&gt;az vm run-command invoke -g $ONPREM_RG -n vm-db --command-id RunPowerShellScript --scripts "Install-WindowsFeature Web-Server -IncludeManagementTools"&lt;BR /&gt;az vm run-command invoke -g $ONPREM_RG -n vm-db --command-id RunPowerShellScript --scripts "Set-Content -Path 'C:\inetpub\wwwroot\index.html' -Value '&amp;lt;h1&amp;gt;OnPrem DB Server IIS Page&amp;lt;/h1&amp;gt;'"&lt;BR /&gt;az vm run-command invoke -g $ONPREM_RG -n vm-db --command-id RunPowerShellScript --scripts "Get-WindowsFeature Web-Server; netstat -ano | findstr :80"&lt;BR /&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 100.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H1&gt;Validate the solution&lt;/H1&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Checkpoint&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Command&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Expected result&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;VPN tunnel status&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;az network vpn-connection show -g $RG -n conn-to-onprem --query connectionStatus -o tsv&lt;BR /&gt;az network vpn-connection show -g $ONPREM_RG -n conn-to-azure --query connectionStatus -o tsv&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Both connections show Connected.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;DB web service&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;az vm run-command invoke -g $ONPREM_RG -n vm-db --command-id RunPowerShellScript --scripts "netstat -ano | findstr :80"&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;The DB VM is listening on TCP port 80.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Private Endpoint connection&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;az network private-link-service show -g $RG -n $PLS_NAME --query "privateEndpointConnections[].{Connection:name,Status:privateLinkServiceConnectionState.status}" -o table&lt;BR /&gt;az network private-endpoint show -g $RG -n $PE_NAME --query "{Name:name,ProvisioningState:provisioningState}" -o table&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;The connection status is Approved, and the Private Endpoint provisioning state is Succeeded.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Hard-coded IP test from web and app tiers&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;az vm run-command invoke -g $RG -n vm-web --command-id RunPowerShellScript --scripts "Test-NetConnection 10.220.1.6 -Port 80"&lt;BR /&gt;az vm run-command invoke -g $RG -n vm-app --command-id RunPowerShellScript --scripts "Test-NetConnection 10.220.1.6 -Port 80"&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;TcpTestSucceeded is True.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;HTTP test from web and app tiers&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;az vm run-command invoke -g $RG -n vm-web --command-id RunPowerShellScript --scripts "Invoke-WebRequest http://10.220.1.6 -UseBasicParsing"&lt;BR /&gt;az vm run-command invoke -g $RG -n vm-app --command-id RunPowerShellScript --scripts "Invoke-WebRequest http://10.220.1.6 -UseBasicParsing"&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;The response contains the sample IIS page.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H1&gt;Production considerations&lt;/H1&gt;
&lt;UL&gt;
&lt;LI&gt;Do not peer the extended prod VNet with the Azure hub VNet. The extended VNet uses 10.220.1.0/24, which overlaps with the 10.220.0.0/21 on-premises prefix that the hub routes through the VPN Gateway. The repeated database IP is intentionally exposed to the web and app tiers only through the Private Endpoint abstraction, not through direct VNet peering.&lt;/LI&gt;
&lt;LI&gt;The PLS subnet must have Private Link service network policies disabled. The Private Endpoint subnet must have Private Endpoint network policies disabled for this lab pattern.&lt;/LI&gt;
&lt;LI&gt;Private Link Service Direct Connect requires at least two IP configurations. The lab uses 10.230.3.10 and 10.230.3.11 for high availability alignment.&lt;/LI&gt;
&lt;LI&gt;The destination IP must be privately routable and reachable from the PLS path. In this lab, the hub routes 10.220.0.0/21 to the on-premises side through VPN Gateway.&lt;/LI&gt;
&lt;LI&gt;For production, validate return path, NSG rules, route tables, firewall policies, DNS behavior, logging, and operational ownership before adopting the pattern.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H1&gt;Clean up&lt;/H1&gt;
&lt;P&gt;When the lab is no longer required, remove the lab resource groups. Validate that no shared resources exist in these groups before deletion.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;az group delete --name rg-pls-demo --yes --no-wait&lt;BR /&gt;az group delete --name onprem-group --yes --no-wait&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 100.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H1&gt;Conclusion&lt;/H1&gt;
&lt;P&gt;This lab shows how a Private Endpoint and Private Link Service Direct Connect can provide a controlled transition path when an application tier moves to Azure before a fixed-address dependency. The pattern can reduce immediate application change, but it should remain an interim architecture: validate routing symmetry, network security, monitoring, regional availability, operational ownership, and a clear plan to remove the legacy address dependency before production adoption.&lt;/P&gt;
&lt;P&gt;Used well, this approach reduces migration risk, avoids application remediation cost, and keeps a datacentre exit moving when it would otherwise stall. In a recent engagement it unblocked a phased migration for an enterprise customer whose application carried a hard-coded dependency address, without a single change to application code.&lt;/P&gt;
&lt;P&gt;Private Link Service Direct Connect is in public preview and available in selected regions, so confirm regional availability, routing, DNS behaviour, and operational ownership before any production adoption.&lt;/P&gt;
&lt;P&gt;If you try the lab, I would welcome your feedback, and any variations you find useful in your own migrations.&lt;/P&gt;
&lt;H1&gt;References&lt;/H1&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/private-link/configure-private-link-service-direct-connect" target="_blank" rel="noopener"&gt;Configure Private Link Service Direct Connect&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/private-link/private-link-service-overview" target="_blank" rel="noopener"&gt;What is Azure Private Link Service?&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/private-link/" target="_blank" rel="noopener"&gt;Azure Private Link documentation&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/vpn-gateway/" target="_blank" rel="noopener"&gt;Azure VPN Gateway documentation&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/vpn-gateway/tutorial-site-to-site-portal" target="_blank" rel="noopener"&gt;Create a site-to-site VPN connection&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/bastion/bastion-nsg" target="_blank" rel="noopener"&gt;Configure NSG rules for Azure Bastion&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H1&gt;About the author&lt;/H1&gt;
&lt;P&gt;Kumar Kaushal is a Senior Digital Cloud Solutions Architect at Microsoft, focused on helping customers design and validate practical Azure architectures for cloud migration, hybrid connectivity, networking, resiliency, and modernization.&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Mon, 27 Jul 2026 12:31:30 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/preserve-a-legacy-ip-during-azure-migration-with-private-link/ba-p/4540575</guid>
      <dc:creator>kumarshashikaushal</dc:creator>
      <dc:date>2026-07-27T12:31:30Z</dc:date>
    </item>
    <item>
      <title>Reminder: Path to Production for Agents Webinar Series Starts Next Week</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/reminder-path-to-production-for-agents-webinar-series-starts/ba-p/4539877</link>
      <description>&lt;P&gt;Next week, join Microsoft for the&amp;nbsp;&lt;STRONG&gt;Path to Production for Agents&lt;/STRONG&gt; webinar series—a free, six-session technical training designed to help organizations move from AI experimentation to secure, scalable, production-ready agent solutions. The series takes place &lt;STRONG&gt;July 27–28&lt;/STRONG&gt; and is aimed at architects, technical leaders, engineers, and AI practitioners looking to operationalize AI at enterprise scale.&lt;/P&gt;
&lt;P&gt;Across six expert-led sessions, you'll learn how to:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Establish governance foundations for AI at scale&lt;/LI&gt;
&lt;LI&gt;Design production-ready AI platforms and landing zones&lt;/LI&gt;
&lt;LI&gt;Build reliable multi-agent architectures&lt;/LI&gt;
&lt;LI&gt;Implement AgentOps practices for deployment and observability&lt;/LI&gt;
&lt;LI&gt;Secure and govern AI systems&lt;/LI&gt;
&lt;LI&gt;Optimize performance, cost, and scalability&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Each session includes proven Microsoft architecture patterns, real-world engineering guidance, and practical techniques you can apply immediately to accelerate your path from prototype to production.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Register today:&lt;/STRONG&gt; &lt;A href="https://aka.ms/AccelerateThePathToProduction" target="_blank"&gt;Path to Production for Agents Registration&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;Don't miss this opportunity to gain the knowledge and frameworks needed to confidently deploy production-grade AI agents at scale. We look forward to seeing you next week!&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Thu, 23 Jul 2026 00:15:07 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/reminder-path-to-production-for-agents-webinar-series-starts/ba-p/4539877</guid>
      <dc:creator>brauerblogs</dc:creator>
      <dc:date>2026-07-23T00:15:07Z</dc:date>
    </item>
    <item>
      <title>Hypervelocity Engineering: Accelerating Enterprise AI with Azure AI Landing Zones</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/hypervelocity-engineering-accelerating-enterprise-ai-with-azure/ba-p/4536192</link>
      <description>&lt;P&gt;Artificial Intelligence is evolving at unprecedented speed. The challenge for enterprises is no longer building AI solutions—it is engineering AI platforms that can adapt, scale, and govern innovation continuously. Hypervelocity Engineering (HVE) provides the engineering operating model that enables this transformation.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Executive Summary&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Artificial Intelligence is transforming how enterprises design and operate digital platforms. Traditional Enterprise Architecture practices—built around static documentation, periodic governance reviews, and manual implementation—are struggling to keep pace with the rapid evolution of AI workloads.&lt;/P&gt;
&lt;P&gt;HyperVelocity Engineering (HVE) introduces a disciplined engineering operating model that combines AI-assisted decision making, platform engineering, automation, and continuous governance. When applied to Azure AI Landing Zone, HVE enables architects to evolve from producing architecture documents to continuously engineering secure, governed, and scalable AI platforms.&lt;/P&gt;
&lt;P&gt;This article demonstrates how HVE Core principles map naturally to Azure AI Landing Zone and how Enterprise Architects can use RPIR (Research, Plan, Implement, Review) as a continuous architecture lifecycle.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Why Enterprise AI Needs a New Engineering Model&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Organizations across every industry are rapidly adopting Generative AI, intelligent agents, and AI-assisted business processes. Yet many enterprise AI initiatives encounter the same obstacles:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;AI projects are developed independently across business units.&lt;/LI&gt;
&lt;LI&gt;Platform capabilities evolve slower than AI innovation.&lt;/LI&gt;
&lt;LI&gt;Governance and security become reactive rather than proactive.&lt;/LI&gt;
&lt;LI&gt;Infrastructure is treated as a one-time deployment instead of a continuously evolving product.&lt;/LI&gt;
&lt;LI&gt;Engineering teams spend excessive time provisioning environments instead of delivering business value.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Traditional cloud engineering practices were designed for application modernization. Enterprise AI introduces new demands—rapid experimentation, scalable model deployment, secure data access, and continuous compliance. Meeting these demands requires a fundamentally different engineering approach.&lt;/P&gt;
&lt;P&gt;Hypervelocity Engineering (HVE) addresses this challenge by enabling organizations to build &lt;STRONG&gt;platforms that evolve at the speed of AI while maintaining enterprise-grade governance, security, and operational excellence.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;From Traditional Engineering to Hypervelocity Engineering&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;Enterprise architecture is shifting from project-centric delivery to platform-centric engineering.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Traditional Engineering&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Hypervelocity Engineering&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Project-based delivery&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Product &amp;amp; platform engineering&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Manual provisioning&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Infrastructure as Code&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Governance after deployment&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Governance by Design&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Static architecture&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Continuous evolution&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Infrastructure owned by IT&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Self-service engineering platforms&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Periodic releases&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Continuous delivery&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Manual operations&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;AI-assisted engineering&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;This transformation allows engineering organizations to focus less on repetitive operational tasks and more on innovation.&lt;/P&gt;
&lt;H4&gt;&lt;STRONG&gt;What is HyperVelocity Engineering?&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P&gt;Hypervelocity Engineering (HVE) is Microsoft's engineering operating model designed to accelerate software and platform delivery through automation, reusable engineering patterns, platform engineering, AI-assisted development, and continuous feedback.&lt;/P&gt;
&lt;P&gt;Unlike traditional methodologies, &lt;STRONG&gt;HVE is not another architecture framework or software product&lt;/STRONG&gt;. Instead, it defines &lt;STRONG&gt;how engineering organizations operate&lt;/STRONG&gt; to deliver secure, governed, and continuously improving solutions.&lt;/P&gt;
&lt;P&gt;Its core philosophy is straightforward:&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Engineer platforms, automate everything practical, measure continuously, and improve through rapid feedback while keeping humans accountable for critical decisions.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;HVE combines several modern engineering disciplines into a unified operating model:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Platform Engineering&lt;/LI&gt;
&lt;LI&gt;Infrastructure as Code (IaC)&lt;/LI&gt;
&lt;LI&gt;Policy as Code&lt;/LI&gt;
&lt;LI&gt;Security by Design&lt;/LI&gt;
&lt;LI&gt;AI-assisted engineering&lt;/LI&gt;
&lt;LI&gt;Continuous observability&lt;/LI&gt;
&lt;LI&gt;DevSecOps&lt;/LI&gt;
&lt;LI&gt;Human-in-the-loop governance&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Together, these capabilities enable organizations to innovate faster without compromising security or compliance.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Azure AI Landing Zone as the Reference Implementation&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Azure AI Landing Zone provides the foundational platform required to deploy enterprise AI workloads securely and consistently.&lt;/P&gt;
&lt;P&gt;Typical capabilities include:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Microsoft Entra ID&lt;/LI&gt;
&lt;LI&gt;Management Groups&lt;/LI&gt;
&lt;LI&gt;Azure Policy&lt;/LI&gt;
&lt;LI&gt;Hub-and-Spoke Networking&lt;/LI&gt;
&lt;LI&gt;Private Endpoints&lt;/LI&gt;
&lt;LI&gt;Azure Firewall&lt;/LI&gt;
&lt;LI&gt;Azure API Management&lt;/LI&gt;
&lt;LI&gt;Azure AI Foundry&lt;/LI&gt;
&lt;LI&gt;Azure AI Search&lt;/LI&gt;
&lt;LI&gt;Azure Kubernetes Service (AKS)&lt;/LI&gt;
&lt;LI&gt;Azure Monitor&lt;/LI&gt;
&lt;LI&gt;Defender for Cloud&lt;/LI&gt;
&lt;LI&gt;Infrastructure as Code (Bicep/Terraform)&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;HVE provides the engineering operating model that continuously evolves this platform.&lt;/P&gt;
&lt;H4&gt;The Core Principles of Hypervelocity Engineering&lt;/H4&gt;
&lt;P&gt;Hypervelocity Engineering is built around six complementary principles.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Outcome-Driven Engineering&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Every engineering decision should align with measurable business outcomes rather than technology adoption alone. Success is measured by customer value, not infrastructure deployment.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Platform Engineering&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Rather than building infrastructure repeatedly for every project, organizations create reusable, self-service platforms that accelerate application delivery while maintaining consistency.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Automation Everywhere&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Automation extends beyond deployments. Infrastructure provisioning, governance, security validation, policy enforcement, testing, and operations should all be automated wherever possible.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Security by Design&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Security is integrated into every engineering stage—from identity and networking to deployment pipelines and operational monitoring—reducing risk while improving delivery speed.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;AI-Assisted Engineering&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;AI enhances developer productivity by generating code, documentation, infrastructure templates, architecture recommendations, and testing artifacts. Humans remain responsible for validation and governance.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Continuous Observability and Feedback&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Operational insights, telemetry, cost analysis, and user feedback continuously improve future engineering decisions.&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;Applying RPIR to Enterprise Architecture&lt;/STRONG&gt;&lt;/H5&gt;
&lt;P&gt;Rather than treating architecture as a one-time activity, HVE applies the RPIR cycle continuously.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Phase&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Enterprise Architecture Activities&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Azure Capabilities&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Expected Outcome&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Research&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Gather business requirements, architecture standards, security baselines, and reference guidance&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Azure AI Search, CAF, Well-Architected Framework&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Evidence-based architecture decisions&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Plan&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Define target architecture, evaluate options, create ADRs, prioritize roadmap&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Azure Landing Zones, Azure Policy, Management Groups&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Governed architecture blueprint&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Implement&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Deploy infrastructure, policies, networking, AI services, and automation&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Bicep, Terraform, Azure DevOps, GitHub Actions&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Repeatable, automated platform deployment&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Review&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Validate security, compliance, reliability, cost, and operational readiness&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Defender for Cloud, Azure Monitor, Azure Advisor&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Continuous improvement and governance&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;Instead of ending after deployment, the Review phase feeds directly back into Research, enabling continuous architecture evolution.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Mapping HVE Principles to Azure AI Landing Zone&lt;/STRONG&gt;&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;HVE Principle&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Azure AI Landing Zone Implementation&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Outcome Driven&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Align landing zone design with business outcomes and AI strategy&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Platform Engineering&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Shared AI Hub, reusable landing zone modules, self-service provisioning&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Automation Everywhere&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Infrastructure as Code, CI/CD pipelines, Policy as Code&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Security by Design&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Zero Trust, Microsoft Entra ID, Key Vault, Defender for Cloud&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Observability &amp;amp; Feedback&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Azure Monitor, Log Analytics, cost insights, operational metrics&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;AI-Assisted Engineering&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Azure AI Foundry, GitHub Copilot, AI Search, architecture assistants&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;These principles ensure that governance, security, and operational excellence are embedded throughout the platform lifecycle rather than applied only during reviews.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Reference Architecture Walkthrough&lt;/STRONG&gt;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;The architecture consists of four logical layers:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;AI Hub&lt;/STRONG&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;The AI Hub provides centralized ingress and shared platform services using Azure Application Gateway, Azure Firewall, and Azure API Management. These services enforce security, routing, and governance before requests reach AI workloads.&lt;/P&gt;
&lt;OL start="2"&gt;
&lt;LI&gt;&lt;STRONG&gt;Shared Platform Services&lt;/STRONG&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Shared Azure Kubernetes Service (AKS), Azure AI Foundry, and Azure AI Search provide reusable AI capabilities for multiple business units. Centralizing these services improves scalability, operational consistency, and cost efficiency.&lt;/P&gt;
&lt;OL start="3"&gt;
&lt;LI&gt;&lt;STRONG&gt;AI Spoke Environments&lt;/STRONG&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Each business unit deploys isolated AI workloads into dedicated spoke environments. Private networking, isolated data stores, and workload-specific resources maintain tenant separation while consuming shared platform capabilities.&lt;/P&gt;
&lt;OL start="4"&gt;
&lt;LI&gt;&lt;STRONG&gt;Platform Governance&lt;/STRONG&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Infrastructure as Code, Azure Policy, Microsoft Entra ID, Defender for Cloud, Azure Monitor, and FinOps capabilities operate across every layer to provide continuous governance, security, and operational visibility.&lt;/P&gt;
&lt;P&gt;This architecture reflects HVE's platform engineering philosophy by separating shared capabilities from workload-specific implementations while maintaining centralized governance.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Benefits for Enterprise Architects&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Adopting HVE alongside Azure AI Landing Zone provides measurable advantages:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Accelerates architecture delivery through AI-assisted design.&lt;/LI&gt;
&lt;LI&gt;Standardizes landing zone deployments using reusable platform modules.&lt;/LI&gt;
&lt;LI&gt;Embeds governance through Policy as Code and Infrastructure as Code.&lt;/LI&gt;
&lt;LI&gt;Improves security with Zero Trust and continuous compliance validation.&lt;/LI&gt;
&lt;LI&gt;Enables continuous architecture evolution using the RPIR lifecycle.&lt;/LI&gt;
&lt;LI&gt;Reduces operational overhead through automation and observability.&lt;/LI&gt;
&lt;LI&gt;Supports FinOps practices with integrated cost monitoring and optimization.&lt;/LI&gt;
&lt;LI&gt;Creates auditable architecture decisions through Architecture Decision Records (ADRs).&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Rather than replacing architects, HVE allows architects to focus on strategic decisions while AI assists with repetitive engineering activities.&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;Common Anti-Patterns to Avoid&lt;/STRONG&gt;&lt;/H5&gt;
&lt;P&gt;Organizations should avoid several common pitfalls when adopting enterprise AI:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Building separate AI platforms for every project&lt;/LI&gt;
&lt;LI&gt;Treating governance as a post-deployment activity&lt;/LI&gt;
&lt;LI&gt;Relying on manual infrastructure provisioning&lt;/LI&gt;
&lt;LI&gt;Operating without continuous telemetry and feedback&lt;/LI&gt;
&lt;LI&gt;Deploying AI services without standardized platform engineering&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Hypervelocity Engineering addresses these anti-patterns by embedding automation, governance, and continuous improvement into the engineering lifecycle.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;Design Recommendations&lt;/STRONG&gt;&lt;/H5&gt;
&lt;P&gt;When adopting Hypervelocity Engineering for Azure AI Landing Zones:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Start with measurable business outcomes rather than technology selection.&lt;/LI&gt;
&lt;LI&gt;Treat your AI Landing Zone as a reusable platform product.&lt;/LI&gt;
&lt;LI&gt;Automate infrastructure, security, and governance using Infrastructure as Code and Policy as Code.&lt;/LI&gt;
&lt;LI&gt;Integrate AI-assisted engineering while maintaining human oversight for architectural decisions.&lt;/LI&gt;
&lt;LI&gt;Use the RPIR loop to continuously evolve your platform based on operational insights and business feedback.&lt;/LI&gt;
&lt;LI&gt;Embed security, observability, and FinOps from the outset to support sustainable growth.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Conclusion&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Enterprise AI success depends on more than deploying advanced models—it requires an engineering operating model capable of delivering secure, scalable, and continuously evolving platforms.&lt;/P&gt;
&lt;P&gt;Hypervelocity Engineering provides that model by combining platform engineering, automation, AI-assisted development, and continuous feedback into a unified approach. When paired with Azure AI Landing Zones, it transforms static infrastructure into a living platform that accelerates innovation while preserving governance, security, and operational excellence.&lt;/P&gt;
&lt;P&gt;For enterprise architects, the value is clear: &lt;STRONG&gt;Azure AI Landing Zones establish the foundation, and Hypervelocity Engineering ensures that foundation continuously adapts to changing business needs and the rapid pace of AI innovation.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;EM&gt;In the era of enterprise AI, success belongs to organizations that engineer for continuous evolution—not just initial deployment. Hypervelocity Engineering offers the operating model to achieve exactly that.&lt;/EM&gt;&lt;/P&gt;</description>
      <pubDate>Fri, 17 Jul 2026 14:57:27 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/hypervelocity-engineering-accelerating-enterprise-ai-with-azure/ba-p/4536192</guid>
      <dc:creator>VimalVerma</dc:creator>
      <dc:date>2026-07-17T14:57:27Z</dc:date>
    </item>
    <item>
      <title>From Policy to Proof: Governing AI to Scale Human Ambition and Machine Intelligence</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/from-policy-to-proof-governing-ai-to-scale-human-ambition-and/ba-p/4535137</link>
      <description>&lt;P&gt;&lt;SPAN data-contrast="none"&gt;Enterprises are moving from AI experiments to AI in production. Copilots are in the flow of daily work, custom applications are built on foundation models, and autonomous agents are beginning to take actions on behalf of the business. As organizations combine&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;human ambition with machine intelligence&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN data-contrast="none"&gt;&lt;STRONG&gt;, trust becomes the defining requirement&lt;/STRONG&gt; for scaling AI responsibly and realizing AI's full potential.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:160,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;That shift changes the governance question from&amp;nbsp;&lt;/SPAN&gt;&lt;EM&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;"can we use AI responsibly?"&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/EM&gt;&lt;SPAN data-contrast="none"&gt;&lt;EM&gt;&lt;STRONG&gt;&amp;nbsp;&lt;/STRONG&gt;&lt;/EM&gt;to&amp;nbsp;a harder one:&lt;EM&gt;&lt;STRONG&gt;&amp;nbsp;&lt;/STRONG&gt;&lt;/EM&gt;&lt;/SPAN&gt;&lt;EM&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;"can we prove, continuously, that every AI system and agent in our estate is safe, compliant, observable, and accountable?"&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:160,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;The organizations getting this right have stopped treating governance as a static document that lives in a policy binder. They treat it as an operating system: a connected set of policies, controls, telemetry, and evidence that runs alongside AI everywhere it&amp;nbsp;operates.&amp;nbsp;Governance defines what should happen.&amp;nbsp;Observability and evaluations verify what is actually happening.&amp;nbsp;Audit and response prove it, and feed what they learn back into policy. The loop, not any single control, is what makes governance real.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:160,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;Every control in this map ladders up to a principled foundation:&amp;nbsp;&lt;/SPAN&gt;&lt;A href="https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/microsoft/final/en-us/microsoft-brand/documents/Microsoft-Responsible-AI-Standard-General-Requirements.pdf?culture=en-us&amp;amp;country=us" target="_blank" rel="noopener"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;Microsoft's Responsible AI Standard&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="none"&gt;, whose six principles (fairness, reliability and safety, privacy and security, inclusiveness, transparency, and accountability) set the bar these controls exist to meet.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;This post lays out a practical reference map for that operating system: the&amp;nbsp;domains&amp;nbsp;AI governance must cover, the cross-cutting runtime enforcement layer that operationalizes them, and the Microsoft services that support each one.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;U&gt;&lt;SPAN class="lia-text-color-21"&gt;&lt;STRONG&gt;The four pillars of AI governance&lt;/STRONG&gt;&lt;/SPAN&gt;&lt;/U&gt;&lt;/H2&gt;
&lt;P&gt;At its core, AI governance does four things. Skip any one of them and a gap opens up -policy without control is aspirational, control without visibility is blind, and visibility without proof cannot withstand an audit.&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;Policy - Define the rules&lt;/STRONG&gt;, and who owns them. What's acceptable use, which use cases are approved, how each is classified by risk, and -critically -who is accountable for outcomes at each stage.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Control - Enforce the boundaries&lt;/STRONG&gt;, proportionate to risk. Turn policy into guardrails, access decisions, and lifecycle gates that constrain systems at runtime -with tighter controls on higher-risk use cases.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Visibility - Observe behavior&lt;/STRONG&gt;, including fairness and quality. Capture logs, metrics, traces, token consumption, tool calls, and runtime signals that reveal what AI is truly doing -not just latency and cost, but grounding, drift, and fairness signals.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Proof - Audit, prove, and improve&lt;/STRONG&gt;. Produce the evidence to demonstrate compliance and transparency (explainability and traceability you can show regulators and customers), investigate incidents, and feed what you learn back into policy.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H1&gt;&lt;U&gt;&lt;SPAN class="lia-text-color-21"&gt;&lt;STRONG&gt;What AI governance covers: the domain map&lt;/STRONG&gt;&lt;/SPAN&gt;&lt;/U&gt;&lt;/H1&gt;
&lt;P&gt;A complete AI governance program spans nine core domains, with runtime enforcement operating as a cross-cutting layer across several of them. Use this taxonomy to map your own requirements to concrete controls:&amp;nbsp;&amp;nbsp;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Policy:&lt;/STRONG&gt; Responsible AI policies, approval workflows, and guardrails.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Data governance:&lt;/STRONG&gt; Classification, sensitivity labels, DLP, retention, and lineage.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Model governance:&lt;/STRONG&gt; Validation, versioning, documentation, and change control.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Observability:&lt;/STRONG&gt; Logs, metrics, traces, token usage, and runtime signals.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Evaluations:&lt;/STRONG&gt; Quality, safety, grounding, drift, and agent task success.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Security:&lt;/STRONG&gt; Threat detection, posture management, proactive AI red teaming, and prompt-injection resistance.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Identity &amp;amp; access:&lt;/STRONG&gt; RBAC, least privilege, conditional access, and agent identities.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Audit &amp;amp; compliance:&lt;/STRONG&gt; Evidence, eDiscovery, legal hold, and reporting.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Agent governance:&lt;/STRONG&gt; Registry, lifecycle, policy enforcement, and fleet visibility.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Cross-cutting runtime enforcement:&lt;/STRONG&gt; Runtime enforcement is not a separate governance program. It is the control layer that operationalizes policy, security, identity, data protection, observability, and agent governance while AI systems are running. It governs live interactions between users, agents, models, tools, APIs, MCP servers, A2A agent APIs, and enterprise systems through authentication, authorization, traffic controls, token governance, quotas, policy checks, and operational guardrails.&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;The Microsoft governance stack&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;Microsoft's approach is a composed platform rather than a single product — AI lifecycle, data governance, identity, security, observability, and audit working together:&lt;/P&gt;
&lt;img /&gt;
&lt;H3&gt;1. Policy and the control plane: define and enforce&lt;/H3&gt;
&lt;P&gt;Governance starts with a place to define the rules and decide who owns them. This is where responsible AI and acceptable use policies are authored; where new use cases come in through intake, get classified by risk, and are approved or rejected; where you decide up front which decisions and agent actions require human sign-off versus running autonomously within guardrails; and where accountability is assigned across the deploy, update, and retire lifecycle. Enforcing those rules through runtime guardrails, security posture, agent-level controls, and API-level controls is the job of the sections that follow; the work here is defining what “acceptable” means before anything is switched on.&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Services that support it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Microsoft Foundry&lt;/STRONG&gt; -Compliance workspace for authoring and housing responsible AI policies and use case approvals.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure Policy&lt;/STRONG&gt; -Resource level governance to codify those rules as enforceable baselines.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure API Management (AI Gateway):&amp;nbsp;&lt;/STRONG&gt;Runtime enforcement for AI applications, models, agents, tools, and APIs, including authentication, authorization, token governance, quota management, and policy enforcement.&amp;nbsp;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;(Agent specific policy and lifecycle controls via Microsoft Agent 365 are covered in Section 6.)&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Practical takeaway: &lt;/STRONG&gt;establish a governance baseline before AI and agent adoption scales -including deciding up front where a human stays in the loop. Retrofitting controls onto a sprawling estate is far harder than starting from one.&lt;/P&gt;
&lt;H3&gt;2. Data governance and compliance: Purview at the center&lt;/H3&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;Most AI risk is&amp;nbsp;ultimately data&amp;nbsp;risk: the&amp;nbsp;wrong information&amp;nbsp;reaching a model, leaving through a response, or being&amp;nbsp;retained&amp;nbsp;without a policy.&amp;nbsp;&lt;/SPAN&gt;&lt;A href="https://www.microsoft.com/en-us/security/business/microsoft-purview" target="_blank" rel="noopener"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;Microsoft Purview&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="none"&gt;&amp;nbsp;is the primary governance engine for AI interactions and enterprise data controls, and it works in three moves:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Discover - &lt;/STRONG&gt;Identify AI usage, sensitive prompts and responses, and overall risk posture with DSPM for AI.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Protect - &lt;/STRONG&gt;Apply sensitivity labels, DLP, access boundaries, and retention policies.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Prove - &lt;/STRONG&gt;Generate audit trails, eDiscovery evidence, and compliance reporting.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;If the question is &lt;EM&gt;“how do we govern the data flowing through AI?”&lt;/EM&gt; - Purview is the anchor.&lt;/P&gt;
&lt;H3&gt;3. Model governance, safety, and evaluations: quality at build time and runtime&lt;/H3&gt;
&lt;P&gt;Evaluations are what make governance measurable, and they belong at two points in the lifecycle:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Pre-deployment - &lt;/STRONG&gt;Release gating on safety, grounding, quality, fairness, and task success before anything ships.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Post-deployment - &lt;/STRONG&gt;Runtime scoring for drift, quality thresholds, and harmful-output alerts once a system is live.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Services that support it:&amp;nbsp;&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="7" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559683&amp;quot;:0,&amp;quot;335559684&amp;quot;:-2,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/foundry/concepts/built-in-evaluators" target="_blank" rel="noopener"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;Microsoft Foundry evaluators&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="none"&gt;:&lt;/SPAN&gt;&lt;SPAN data-contrast="none"&gt;&amp;nbsp;quality, safety, RAG grounding, and agent task metrics&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="7" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559683&amp;quot;:0,&amp;quot;335559684&amp;quot;:-2,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/ai-services/content-safety/overview" target="_blank" rel="noopener"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;Azure AI Content Safety&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="none"&gt;:&lt;/SPAN&gt;&lt;SPAN data-contrast="none"&gt;&amp;nbsp;detects unsafe or harmful content in prompts and responses. This is also integrated in Microsoft Foundry as guardrails that protect against unsafe content and direct and indirect&amp;nbsp;jailbreak&amp;nbsp;prompt injection attacks.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="7" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559683&amp;quot;:0,&amp;quot;335559684&amp;quot;:-2,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;A href="https://github.com/responsibleai/ASSERT" target="_blank" rel="noopener"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;ASSERT&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="none"&gt;&amp;nbsp;(open source):&lt;/SPAN&gt;&lt;SPAN data-contrast="none"&gt;&amp;nbsp;Microsoft's open-source, policy-driven evaluation framework, short for Adaptive Spec-driven Scoring for Evaluation and Regression Testing. ASSERT turns your&amp;nbsp;specific requirements and&amp;nbsp;organizational policies into targeted, safety-focused test cases rather than generic benchmarks, and runs across frameworks such as&amp;nbsp;LangChain,&amp;nbsp;CrewAI,&amp;nbsp;LiteLLM, and OpenAI, so evaluation is not locked to any single stack. ASSERT also pairs with runtime controls to close the loop: measure a failure rate, apply a control (see&amp;nbsp;&lt;/SPAN&gt;&lt;A href="https://commandline.microsoft.com/agent-control-specification-runtime-governance/" target="_blank" rel="noopener"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;Agent Control Specification&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="none"&gt;&amp;nbsp;in section 6), then re-run the same evaluation to prove the rate dropped. That turns "we mitigated it" into a measured before-and-after.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Evaluations are not separate from governance - they are what turn governance from a promise into measurable proof.&amp;nbsp;&lt;/P&gt;
&lt;H3&gt;4. Observability and runtime monitoring: know what is actually happening&lt;/H3&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;Once AI is in production, you need to see its behavior in real time. Capture four kinds of signal:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="6" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559683&amp;quot;:0,&amp;quot;335559684&amp;quot;:-2,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;Logs:&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN data-contrast="none"&gt;&amp;nbsp;prompts, responses, and decisions&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="6" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559683&amp;quot;:0,&amp;quot;335559684&amp;quot;:-2,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;Metrics:&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN data-contrast="none"&gt;&amp;nbsp;latency, token usage, quality scores, request volume, and quota consumption&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="6" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559683&amp;quot;:0,&amp;quot;335559684&amp;quot;:-2,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;Traces:&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN data-contrast="none"&gt;&lt;STRONG&gt;&amp;nbsp;&lt;/STRONG&gt;agent reasoning paths, tool calls, model calls, and inter-service dependencies&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="6" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559683&amp;quot;:0,&amp;quot;335559684&amp;quot;:-2,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;Alerts:&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN data-contrast="none"&gt;&lt;STRONG&gt;&amp;nbsp;&lt;/STRONG&gt;policy violations, drift, harmful outputs, and abnormal traffic patterns&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;Microsoft Foundry provides purpose-built observability for AI applications and agents across three connected capabilities. Evaluation measures quality, safety, and reliability throughout development. Monitoring tracks deployed systems in real-world conditions. Tracing, built on OpenTelemetry standards, captures the execution flow of LLM calls, tool invocations, and agent decisions, so you can debug multi-step reasoning rather than guess at it, across frameworks including LangChain, LangGraph, the OpenAI Agents SDK, and the Microsoft Agent Framework. The Agent Monitoring Dashboard brings these signals together for production traffic: token usage, latency, run success rates, evaluation scores, and red-teaming results in one view, backed by Azure Monitor, Application Insights. From the same surface you can turn on continuous evaluation to score sampled live responses, scheduled evaluations to detect drift against benchmarks, scheduled red team scans to probe deployed agents for emerging risks, and alerts that fire when outputs fail quality thresholds or produce harmful content.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;Critically, this coverage is not limited to agents built on Foundry. Agents running elsewhere can be onboarded through the AI Gateway and instrumented with&amp;nbsp;OpenTelemetry&amp;nbsp;semantic conventions to send telemetry to the same Application Insights instance, so one dashboard, one evaluation pipeline, and one alerting surface cover the whole estate.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;The broader observability fabric&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;Foundry observability composes with the rest of the platform: Azure Monitor, Application Insights, Log Analytics, Purview Audit, and Microsoft Agent 365 dashboards. Azure API Management (AI Gateway) extends visibility further by capturing AI traffic patterns, token consumption, API usage, policy violations, and agent-to-tool interactions,&amp;nbsp;providing&amp;nbsp;runtime insight into AI workloads at the network boundary.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;Practical takeaway:&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN data-contrast="none"&gt;&lt;STRONG&gt;&amp;nbsp;&lt;/STRONG&gt;wire&amp;nbsp;observability&amp;nbsp;before launch, not after the first incident. Continuous evaluation on sampled production traffic is the cheapest early-warning system you can buy, and the trace data it rides on is the same evidence an auditor will ask for later.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H3&gt;5. Security and identity governance: protect the AI estate&lt;/H3&gt;
&lt;P&gt;AI expands the attack surface, so governance has to include security and identity across the following:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Threat protection -&lt;/STRONG&gt;Prompt injection, jailbreak, data exfiltration, and unsafe actions -with Microsoft Defender for Cloud and Microsoft Defender XDR.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Proactive adversarial testing -&lt;/STRONG&gt;Find failures before attackers do, with the AI Red Teaming Agent in Microsoft Foundry for automated agent probing and PyRIT, Microsoft’s open-source red-teaming framework.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Security posture -&lt;/STRONG&gt;Misconfigurations, exposed endpoints, and risky resources -with Defender for Cloud and Azure Policy.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Identity &amp;amp; access -&lt;/STRONG&gt;Who can use AI, and what data, apps, and agents it can reach -with Microsoft Entra ID and RBAC.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Runtime access enforcement:&amp;nbsp;&lt;/STRONG&gt;Who can call which model, tool, API, or MCP server, under which conditions, and with what level of throttling, inspection, and monitoring.&amp;nbsp;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Secrets &amp;amp; network -&lt;/STRONG&gt;Protecting keys and enforcing private access and network boundaries -with Azure Key Vault and Private Link.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Device management&lt;/STRONG&gt;-Discover, monitor, and govern unmanaged AI agents and the endpoints they run on -with Microsoft Intune and Microsoft Defender for Endpoint (MDE).&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Cross-cutting runtime enforcement&amp;nbsp;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;As AI systems increasingly rely on APIs, tools, MCP servers, enterprise applications, and agent-to-agent interactions, governance must extend beyond model build-time controls to the live interactions that&amp;nbsp;connect&amp;nbsp;those systems. Runtime enforcement helps operationalize the broader AI governance program by enforcing who can call what, under which conditions, at what scale, and with what level of monitoring.&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Services that support it&amp;nbsp;:&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure API Management (AI Gateway):&amp;nbsp;&lt;/STRONG&gt;Centralized authentication, authorization, token controls, traffic shaping, quota enforcement, content safety policies, and policy governance for AI models, agents, tools, MCP servers, A2A agent APIs, and enterprise APIs.&amp;nbsp;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Microsoft Entra ID: &lt;/STRONG&gt;Identity and access control for applications, users, and agents.&amp;nbsp;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure RBAC:&amp;nbsp;&lt;/STRONG&gt;Fine-grained authorization for AI resources and workloads.&amp;nbsp;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;The mindset shift: &lt;/STRONG&gt;treat agents as non-human identities with their own permissions, lifecycle, runtime access and monitoring.&lt;/P&gt;
&lt;H3&gt;6. Agent governance: the emerging control plane&lt;/H3&gt;
&lt;P&gt;Agents are a new kind of enterprise actor, and governing them takes two complementary pieces: a control plane to manage the whole fleet, and a portable control standard to enforce safety inside each agent's workflow.&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;Microsoft Agent 365: govern the agent fleet&lt;/STRONG&gt;&lt;/H5&gt;
&lt;P&gt;Agents introduce a new problem - sprawl. As teams build agents on different platforms, organizations need one way to see and govern all of them. Microsoft Agent 365, now generally available, is the control plane for AI agents: it gives each agent its own identity and manages agents with the admin tools you already use - Microsoft Entra, Defender, Purview, Intune, and the Microsoft 365 admin center. It works across agents built on Microsoft, open-source, and third-party platforms, through five core capabilities:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Registry - &lt;/STRONG&gt;Discover and catalog every agent - including shadow agents - and quarantine the unsanctioned ones.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Access control - &lt;/STRONG&gt;Give each agent a unique Microsoft Entra Agent ID, and enforce least-privilege, risk-based access with policy templates.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Visualization - &lt;/STRONG&gt;Unified dashboards, telemetry, and alerts across the entire agent fleet.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Interoperability - &lt;/STRONG&gt;Works with agents from Microsoft, open-source, and third-party frameworks, alongside your Microsoft 365 apps and data.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Security - &lt;/STRONG&gt;Threat detection, data protection, and compliance for agents at scale, through Defender, Purview, and Entra.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H5&gt;&lt;STRONG&gt;Agent Control Specification (ACS): portable runtime controls&lt;/STRONG&gt;&lt;/H5&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;Governing agents also means enforcing controls consistently, no matter which framework an agent is built on. The&amp;nbsp;&lt;/SPAN&gt;&lt;A href="https://commandline.microsoft.com/agent-control-specification-runtime-governance/" target="_blank" rel="noopener"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;Agent Control Specification (ACS)&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="none"&gt;&amp;nbsp;is an open industry standard, part of Microsoft's Agent Governance Toolkit, for placing deterministic safety and security controls at defined checkpoints in an agent's workflow. Just as the Model Context Protocol (MCP) standardized how agents connect to tools, and A2A standardized how agents communicate with each other, ACS aims to standardize how agents are governed: one portable control layer that any framework can adopt, with Microsoft providing reference implementations.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="2" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559683&amp;quot;:0,&amp;quot;335559684&amp;quot;:-2,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;Control checkpoints:&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN data-contrast="none"&gt;&amp;nbsp;validation at defined intervention points across the agent's lifecycle, spanning input, model invocation, state, tool selection and execution, and output&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="2" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559683&amp;quot;:0,&amp;quot;335559684&amp;quot;:-2,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;Composable controls:&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN data-contrast="none"&gt;&amp;nbsp;deterministic checks (allow/deny rules, custom filters) alongside model-based checks (classifier endpoints, LLM judges), each placed exactly where it is needed&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="2" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559683&amp;quot;:0,&amp;quot;335559684&amp;quot;:-2,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;Portable and auditable:&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN data-contrast="none"&gt;&amp;nbsp;expressed as a declarative manifest with policy-as-code evaluation, so controls are&amp;nbsp;versionable, auditable, and travel with the agent across any stack&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Together they cover both altitudes of agent governance: Agent 365 governs the fleet from the outside, while ACS enforces safety inside each agent's workflow. Those checkpoints can also&amp;nbsp;&lt;STRONG&gt;route to a human for approval before a high-impact action&lt;/STRONG&gt; -the runtime side of human-in-the-loop -while Agent 365 defines which agents and actions require sign-off.&lt;/P&gt;
&lt;H1&gt;&lt;U&gt;&lt;STRONG&gt;The governance service map&lt;/STRONG&gt;&lt;/U&gt;&lt;/H1&gt;
&lt;P&gt;Use this as a quick reference for matching each governance area to the Microsoft services that support it. Runtime enforcement is shown as a cross-cutting capability because it operationalizes several domains rather than replacing them:&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;H1&gt;&lt;U&gt;&lt;STRONG&gt;A practical path forward&lt;/STRONG&gt;&lt;/U&gt;&lt;/H1&gt;
&lt;P&gt;You do not have to govern everything at once. A pragmatic sequence gets you to a defensible position quickly:&amp;nbsp;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Start with scope:&amp;nbsp;&lt;/STRONG&gt;Inventory the AI apps, Copilots, agents,&amp;nbsp;APIs, tools, MCP servers,&amp;nbsp;and custom Microsoft Foundry workloads that exist today.&amp;nbsp;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Classify risk: &lt;/STRONG&gt;Determine&amp;nbsp;what data is used, what actions AI can take,&amp;nbsp;what systems it can reach,&amp;nbsp;and which regulations apply.&amp;nbsp;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Apply controls: &lt;/STRONG&gt;Bring in Purview, Entra, Defender, Microsoft Foundry guardrails,&amp;nbsp;Azure API Management AI Gateway policies,&amp;nbsp;Agent 365 policies, and portable ACS checkpoints for agents.&amp;nbsp;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Measure and prove: &lt;/STRONG&gt;Add evaluations, including policy-driven ASSERT runs and adversarial testing with the AI Red Teaming Agent, plus observability,&amp;nbsp;APIM runtime telemetry,&amp;nbsp;audit evidence, and operational dashboards.&amp;nbsp;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Governance works best when it follows the business rather than leading it. This journey view shows how the controls layer in progressively starting from a high-value business use case and adding only what each next step requires, so governance&amp;nbsp;&lt;EM&gt;accelerates&lt;/EM&gt;&amp;nbsp;safe delivery instead of gating it. The path runs through four stages:&amp;nbsp;&lt;STRONG&gt;align&lt;/STRONG&gt;&amp;nbsp;on the use case and its value, stand up a&amp;nbsp;&lt;STRONG&gt;safe foundation&lt;/STRONG&gt;&amp;nbsp;of minimum viable controls,&amp;nbsp;&lt;STRONG&gt;build and validate&lt;/STRONG&gt;&amp;nbsp;quality and safety, then&amp;nbsp;&lt;STRONG&gt;operate and scale&lt;/STRONG&gt;&amp;nbsp;-monitoring, auditing, enforcing runtime controls and governing agents in production.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;&lt;STRONG&gt;Microsoft gives you the building blocks to govern AI as a live operating system - not a static policy document.&lt;/STRONG&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-contrast="none"&gt;To go deeper, explore&amp;nbsp;Microsoft Foundry, Microsoft Purview, Microsoft Defender for Cloud, Microsoft&amp;nbsp;Entra, Azure&amp;nbsp;API Management&amp;nbsp;and Microsoft Agent 365&amp;nbsp;-plus the open-source&amp;nbsp;&lt;/SPAN&gt;&lt;A href="https://github.com/responsibleai/ASSERT" target="_blank" rel="noopener"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;ASSERT&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="none"&gt;&amp;nbsp;evaluation framework and the&amp;nbsp;&lt;/SPAN&gt;&lt;A href="https://commandline.microsoft.com/agent-control-specification-runtime-governance/" target="_blank" rel="noopener"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;Agent Control Specification (ACS)&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="none"&gt;&amp;nbsp;for portable, framework-agnostic agent controls.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:160,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H1&gt;&lt;U&gt;&lt;STRONG&gt;References&lt;/STRONG&gt;&lt;/U&gt;&lt;/H1&gt;
&lt;P&gt;&lt;STRONG&gt;Policy &amp;amp; lifecycle&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Microsoft Foundry&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/azure/foundry/what-is-foundry" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/azure/foundry/what-is-foundry&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Microsoft Foundry Compliance and Security&lt;/STRONG&gt;- &lt;A href="https://learn.microsoft.com/en-us/azure/foundry/control-plane/how-to-manage-compliance-security" target="_blank" rel="noopener"&gt;Manage compliance and security in Microsoft Foundry - Microsoft Foundry | Microsoft Learn&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure Policy&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/azure/governance/policy/overview" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/azure/governance/policy/overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure API Management AI Gateway&lt;/STRONG&gt; - https://learn.microsoft.com/en-us/azure/api-management/genai-gateway-capabilities&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Microsoft Agent 365&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/microsoft-agent-365/overview" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/microsoft-agent-365/overview&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Data governance&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Microsoft Purview&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/purview/purview" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/purview/purview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;DSPM for AI&lt;/STRONG&gt;&amp;nbsp;(Data Security Posture Management) —&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/purview/data-security-posture-management-learn-about" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/purview/data-security-posture-management-learn-about&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;DLP&lt;/STRONG&gt;&amp;nbsp;(Data Loss Prevention) —&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/purview/dlp-learn-about-dlp" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/purview/dlp-learn-about-dlp&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Sensitivity labels&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/purview/sensitivity-labels" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/purview/sensitivity-labels&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Model safety&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure AI Content Safety&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/azure/ai-services/content-safety/overview" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/azure/ai-services/content-safety/overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Prompt Shields&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/jailbreak-detection" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/jailbreak-detection&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Evaluations&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure AI Foundry evaluators&lt;/STRONG&gt;&amp;nbsp;(observability &amp;amp; evaluation) —&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/azure/foundry/concepts/observability" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/azure/foundry/concepts/observability&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;ASSERT&lt;/STRONG&gt;&amp;nbsp;— official announcement (open-source agent evals; no standalone Learn page yet):&amp;nbsp;&lt;A href="https://devblogs.microsoft.com/foundry/build-2026-open-trust-stack-ai-agents/" target="_blank" rel="noopener"&gt;https://devblogs.microsoft.com/foundry/build-2026-open-trust-stack-ai-agents/&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Observability&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure Monitor&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/azure/azure-monitor/fundamentals/overview" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/azure/azure-monitor/fundamentals/overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Application Insights&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/azure/azure-monitor/app/app-insights-overview" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/azure/azure-monitor/app/app-insights-overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Log Analytics&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/azure/azure-monitor/logs/log-analytics-workspace-overview" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/azure/azure-monitor/logs/log-analytics-workspace-overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Purview Audit&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/purview/audit-solutions-overview" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/purview/audit-solutions-overview&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Security&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Microsoft Defender for Cloud&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/azure/defender-for-cloud/defender-for-cloud-introduction" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/azure/defender-for-cloud/defender-for-cloud-introduction&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Microsoft Defender XDR&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/defender-xdr/microsoft-365-defender" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/defender-xdr/microsoft-365-defender&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Microsoft Sentinel&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/azure/sentinel/sentinel-overview" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/azure/sentinel/sentinel-overview&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Identity&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Microsoft Entra ID&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/entra/identity/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/entra/identity/&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure RBAC&lt;/STRONG&gt;&amp;nbsp;(role-based access control) —&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/azure/role-based-access-control/overview" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/azure/role-based-access-control/overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Conditional Access&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/entra/identity/conditional-access/overview" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/entra/identity/conditional-access/overview&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Audit &amp;amp; compliance&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Purview Audit&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/purview/audit-solutions-overview" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/purview/audit-solutions-overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;eDiscovery&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/purview/edisc" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/purview/edisc&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Compliance Manager&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/purview/compliance-manager" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/purview/compliance-manager&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Agent governance&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Microsoft Agent 365&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/microsoft-agent-365/overview" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/microsoft-agent-365/overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Agent Control Specification (ACS)&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://microsoft.github.io/agent-governance-toolkit/packages/agent-control-specification/" target="_blank" rel="noopener"&gt;https://microsoft.github.io/agent-governance-toolkit/packages/agent-control-specification/&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Agent Governance Toolkit&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://microsoft.github.io/agent-governance-toolkit/" target="_blank" rel="noopener"&gt;https://microsoft.github.io/agent-governance-toolkit/&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Purview for agents&lt;/STRONG&gt;&amp;nbsp;—&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/purview/purview" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/purview/purview&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;&lt;STRONG&gt;Contributors:&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;This article is maintained by Microsoft. It was originally written by the following contributors.&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;A href="https://www.linkedin.com/in/trmanasa" target="_blank" rel="noopener"&gt;Manasa Ramalinga&lt;/A&gt;&amp;nbsp;| Senior Principal Cloud Solution Architect – US Customer Success &lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://www.linkedin.com/in/mehrnoosh-sameki/" target="_blank" rel="noopener"&gt;Mehrnoosh Sameki&lt;/A&gt; | Principal PM Manager -AI Governance Product Team&lt;/LI&gt;
&lt;/OL&gt;</description>
      <pubDate>Wed, 15 Jul 2026 19:51:51 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/from-policy-to-proof-governing-ai-to-scale-human-ambition-and/ba-p/4535137</guid>
      <dc:creator>manasa_ramalinga</dc:creator>
      <dc:date>2026-07-15T19:51:51Z</dc:date>
    </item>
    <item>
      <title>The AI Agent Lifecycle: A Simple Guide</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/the-ai-agent-lifecycle-a-simple-guide/ba-p/4535729</link>
      <description>&lt;H2&gt;The Bigger Picture&lt;/H2&gt;
&lt;P&gt;Building an AI agent is fundamentally different from building traditional software.&lt;/P&gt;
&lt;P&gt;With a website or application, teams typically design, develop, test, and release. Once deployed, the focus shifts primarily to maintenance and feature enhancements. AI agents operate differently. They don't just execute predefined instructions they interpret information, reason, and make decisions.&lt;/P&gt;
&lt;P&gt;In banking, those decisions can influence customer experiences, operational efficiency, compliance outcomes, and risk management. As a result, deploying an AI agent is not the finish line; it's the beginning of an ongoing process of learning, monitoring, and improvement.&lt;/P&gt;
&lt;P&gt;The lifecycle of an enterprise AI agent reflects this reality.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;Stage 1 —&amp;gt; Design: Decide What It Can and Cannot Do&lt;/H2&gt;
&lt;P&gt;Before writing a single line of code, answer three questions:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;What is this agent allowed to do?&lt;/LI&gt;
&lt;LI&gt;What must it never do?&lt;/LI&gt;
&lt;LI&gt;Who is accountable when something goes wrong?&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;STRONG&gt;Simple example:&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;A loan agent is allowed to check credit scores, apply lending policy, and recommend a decision. It is never allowed to approve a loan above $50,000 without a human sign-off. The Head of Credit Risk is accountable.&lt;/P&gt;
&lt;P&gt;Design produces one critical output: the&amp;nbsp;&lt;STRONG&gt;risk classification&lt;/STRONG&gt;. A low-risk FAQ bot and a high-risk loan decisioning agent need completely different levels of testing, guardrails, and oversight. Getting this wrong at design time is expensive to fix later.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Microsoft helps here with:&lt;/STRONG&gt;&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Tool&lt;/th&gt;&lt;th&gt;How It Helps&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Microsoft Foundry Model Catalog&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Browse and compare models and select the right one for the risk level before any code is written&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Azure AI Content Safety&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Review the built-in risk categories to understand what the platform can enforce, informing the guardrail boundary decisions&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Microsoft Responsible AI Impact Assessment&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Structured tooling to assess and document harms, likelihood, severity, and mitigations, producing a risk classification artefact.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H2&gt;Stage 2 —&amp;gt; Build: Put the Safety Controls In, Not On&lt;/H2&gt;
&lt;P&gt;Build the agent but more importantly, build the safety controls at the same time. Not afterwards.&lt;/P&gt;
&lt;P&gt;The core stack:&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;STRONG&gt;Simple example:&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The loan agent is built with: a PII redaction step (strip account numbers before they reach the model), a credit bureau tool, a policy lookup tool, an output checker (does the response cite a real policy?), and a Human in the loop gate (flag any decision over $25k for human review).&lt;/P&gt;
&lt;P&gt;Establish the&amp;nbsp;&lt;STRONG&gt;golden dataset&lt;/STRONG&gt; during this phase: a representative set of real-world loan scenarios with SME validated expected outcomes. This serves as the ground truth for evaluating accuracy, consistency, and regression performance throughout the agent lifecycle.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Microsoft helps here with:&lt;/STRONG&gt;&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 100%; height: 448px; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr style="height: 35px;"&gt;&lt;th style="height: 35px;"&gt;Tool&lt;/th&gt;&lt;th style="height: 35px;"&gt;How It Helps&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr style="height: 59px;"&gt;&lt;td style="height: 59px;"&gt;&lt;STRONG&gt;Microsoft Foundry Agent Service&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 59px;"&gt;The primary platform for building and hosting the agent, tool registration, memory, and orchestration in one place&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 59px;"&gt;&lt;td style="height: 59px;"&gt;&lt;STRONG&gt;Azure OpenAI Service&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 59px;"&gt;The LLM backbone with configurable built-in content filters on every inference call&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 59px;"&gt;&lt;td style="height: 59px;"&gt;&lt;STRONG&gt;Azure AI Content Safety&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 59px;"&gt;Input and output guardrails, content moderation and Prompt Shield for injection and jailbreak detection&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 59px;"&gt;&lt;td style="height: 59px;"&gt;&lt;STRONG&gt;Azure AI Language&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 59px;"&gt;PII detection and redaction across 100+ entity types before data reaches the model&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 59px;"&gt;&lt;td style="height: 59px;"&gt;&lt;STRONG&gt;Azure AI Search&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 59px;"&gt;The RAG pipeline retrieves verified policy documents to ground every agent response&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 59px;"&gt;&lt;td style="height: 59px;"&gt;&lt;STRONG&gt;Azure Functions (Premium)&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 59px;"&gt;Hosts custom guardrail logic (e.g. policy compliance checks) inside the bank's private network&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 59px;"&gt;&lt;td style="height: 59px;"&gt;&lt;STRONG&gt;Microsoft Foundry Tracing&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 59px;"&gt;Instruments every tool call and reasoning step, essential for evaluation and audit&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H2&gt;Stage 3 —&amp;gt; Test: Find the Failures Before Customers Do&lt;/H2&gt;
&lt;P&gt;Testing occurs in three waves:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Automated testing&lt;/STRONG&gt; evaluates the agent against the golden dataset, measuring accuracy, groundedness, and safety.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Human review&lt;/STRONG&gt; brings in domain experts to assess decision quality, reasoning, and compliance.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Red teaming&lt;/STRONG&gt; stress-tests the agent with adversarial prompts to uncover vulnerabilities and safety gaps.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The stage concludes with a quality gate, a formal sign-off that the agent meets the required standards. No sign-off, no deployment.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Simple example:&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The loan agent achieves&amp;nbsp;98% accuracy against the golden dataset. A compliance officer reviews a sample of 50 decisions and confirms that the reasoning meets requirements. During red-team testing, a vulnerability is discovered: the agent can be manipulated through instructions embedded within a PDF. The issue is addressed and remediated before deployment.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Microsoft helps here with:&lt;/STRONG&gt;&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Tool&lt;/th&gt;&lt;th&gt;How It Helps&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Microsoft Foundry Evaluation SDK&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Runs the full golden dataset evaluation in parallel structured scores per row, side-by-side comparison between runs&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Built-in Safety Evaluators&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Out-of-the-box scoring for violence, hate, self-harm, sexual content, and indirect prompt injection&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Built-in Quality Evaluators&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Groundedness, relevance, coherence, and fluency no configuration needed&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Agent Evaluators&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;TaskAdherence and ToolCallAccuracy&amp;nbsp; checks the agent followed the right process, not just gave the right answer&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Microsoft Foundry Versioned Datasets&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Locks the golden dataset by version the same benchmark is used for every regression test&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H2&gt;Stage 4 —&amp;gt; Deploy: Start Small, Expand Carefully&lt;/H2&gt;
&lt;P&gt;Do not flip a switch and send all traffic to the new agent. Start in shadow mode.&lt;/P&gt;
&lt;P&gt;Shadow mode: Agent processes requests → responses NOT shown to customers Purpose: does it behave in production like it did in test? Pilot (5%): A small slice of real customers get agent responses Watch error rates for 2 weeks Full rollout: Expand only when quality metrics stay within thresholds&lt;/P&gt;
&lt;P&gt;Everything must be live before the first customer interaction: monitoring dashboards, alerting, human review queues, and a fallback plan if the agent needs to be pulled.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Simple example:&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The loan agent goes live in shadow mode for one week. No unexpected failures. Expands to 5% of applications. Error rate stays below 0.1% for two weeks. Full rollout approved.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Microsoft helps here with:&lt;/STRONG&gt;&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Tool&lt;/th&gt;&lt;th&gt;How It Helps&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Microsoft Foundry Agent Services, Azure Kubernetes Services, Azure Container Apps&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Hosts and auto-scales the agent runtime canary deployments enable the staged rollout without a full infrastructure team&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Azure API Management&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;The API gateway enforces rate limits, authentication, and routing before any request reaches the agent&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Azure Application Insights&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Latency, volume, and error rate dashboards live from the first interaction&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Azure Private Endpoints + Managed Identity&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;All traffic stays inside the bank's network no public endpoints, no passwords in code&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Microsoft Foundry Deployment Management&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Version-pins the model deployment enables instant rollback if the new version degrades&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H2&gt;Stage 5 —&amp;gt; Operate: Watch Everything, Always&lt;/H2&gt;
&lt;P&gt;A deployed agent is not a finished product. It is a living system. Watch five things continuously:&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;What to Watch&lt;/th&gt;&lt;th&gt;Why&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;What's coming in&lt;/td&gt;&lt;td&gt;Are users trying to manipulate the agent?&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;How fast it responds&lt;/td&gt;&lt;td&gt;Is it meeting SLA?&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Quality of outputs&lt;/td&gt;&lt;td&gt;Is it still giving correct answers?&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Guardrail trigger rates&lt;/td&gt;&lt;td&gt;Are more things being blocked or slipping through?&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Business outcomes&lt;/td&gt;&lt;td&gt;Are loan decisions still aligned with policy?&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;Sample a portion of live interactions and route them to human reviewers. When a reviewer corrects the agent, that correction is a training signal — collected, annotated, and fed back into the next iteration.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Simple example:&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Three weeks after launch, monitoring shows the agent's policy compliance score has dropped from 98% to 94%. Human reviewers identify that a recent policy update was not reflected in the agent's RAG knowledge base. The team is alerted before any customers are affected.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Microsoft helps here with:&lt;/STRONG&gt;&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Tool&lt;/th&gt;&lt;th&gt;How It Helps&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Microsoft Foundry Online Evaluation&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Asynchronously samples live traffic and evaluate it, providing continuous quality monitoring without impacting response latency.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Azure AI Content Safety (runtime)&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Prompt Shield enforces guardrails on every production interaction in real time&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Azure Monitor + KQL&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Provides dashboards and alerts across all monitoring signals, including latency, quality, guardrail compliance, and business outcomes.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Microsoft Foundry Tracing&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Captures every production trace, including tool interactions and execution history, providing a complete audit trail for review and compliance&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Traces to Dataset&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Automatically converts production traces into versioned evaluation datasets, feeding seamlessly into the next optimization cycle.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H2&gt;Stage 6 —&amp;gt; Iterate: The Agent Is Never Finished&lt;/H2&gt;
&lt;P&gt;Every signal from production triggers a loop back into the lifecycle.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Signal&lt;/th&gt;&lt;th&gt;What Happens&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Quality score drops&lt;/td&gt;&lt;td&gt;Loop back to Build and update the RAG index&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;New attack pattern detected&lt;/td&gt;&lt;td&gt;Loop back to Build and patch the guardrail, re-test&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Human overrides spiking&lt;/td&gt;&lt;td&gt;Loop back to Design and rethink the HITL threshold&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;New regulation published&lt;/td&gt;&lt;td&gt;Loop back to Design and full cycle restarts&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;&lt;STRONG&gt;Simple example:&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;APRA publishes new guidance on AI in credit decisions. The loan agent must be updated to include a new mandatory disclosure in every decision output. The team loops back to Design, specifies the new requirement, updates the agent in Build, re-tests against an updated golden dataset, and redeploys within three weeks.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Microsoft helps here with:&lt;/STRONG&gt;&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 100%; height: 295px; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr style="height: 35px;"&gt;&lt;th style="height: 35px;"&gt;Tool&lt;/th&gt;&lt;th style="height: 35px;"&gt;How It Helps&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr style="height: 67px;"&gt;&lt;td style="height: 67px;"&gt;&lt;STRONG&gt;Microsoft Foundry Fine-tuning&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 67px;"&gt;
&lt;P&gt;Fine-tunes the model using human-reviewed annotations, enabling the agent to improve with every feedback cycle.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 67px;"&gt;&lt;td style="height: 67px;"&gt;&lt;STRONG&gt;Microsoft Foundry Dataset Versioning&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 67px;"&gt;
&lt;P&gt;Promotes newly annotated traces into the next version of the golden dataset, ensuring regression testing remains current.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 67px;"&gt;&lt;td style="height: 67px;"&gt;&lt;STRONG&gt;Microsoft Foundry Experiment Tracking&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 67px;"&gt;
&lt;P&gt;Maps evaluation outcomes to specific prompt revisions, making it easy to identify the exact change that introduced a regression.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 59px;"&gt;&lt;td style="height: 59px;"&gt;&lt;STRONG&gt;Azure Monitor Alerts&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 59px;"&gt;
&lt;P&gt;Automatically triggers when quality thresholds are breached, initiating the optimization cycle without requiring manual intervention.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H2&gt;Why the Loop Is the Most Important Part&lt;/H2&gt;
&lt;P&gt;Most teams focus on stages 1–4. The loop through stage 6 is what separates agents that stay safe from agents that drift into risk over time.&lt;/P&gt;
&lt;P&gt;Regulators do not just ask "was it safe when you launched it?" They ask, "is it safe now and can you prove it has been improving?"&lt;/P&gt;
&lt;P&gt;The iterate loop, supported by Microsoft Foundry's continuous evaluation and monitoring capabilities, is how you answer yes.&lt;/P&gt;
&lt;H2&gt;Summary&lt;/H2&gt;
&lt;P&gt;Design what the agent can and cannot do. Build the safety controls into the system at the same time as the agent itself. Microsoft Foundry Agent Service, Azure AI Content Safety, and Azure AI Search provide the core infrastructure. Test it in three waves before any customer sees it, using Microsoft Foundry's evaluation SDK and built-in evaluators. Deploy it carefully in stages, watching every metric through Azure Monitor and Application Insights. Once live, monitor it continuously using Microsoft Foundry's online evaluation. And when something changes a policy, a regulation, a performance drift&amp;nbsp; loop back to the right stage and run the cycle again. The agent is never finished. The loop is the product.&lt;/P&gt;</description>
      <pubDate>Mon, 13 Jul 2026 00:52:49 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/the-ai-agent-lifecycle-a-simple-guide/ba-p/4535729</guid>
      <dc:creator>supriyas</dc:creator>
      <dc:date>2026-07-13T00:52:49Z</dc:date>
    </item>
    <item>
      <title>Beyond the Canvas: The Azure Architecture Diagram Builder Becomes Agent-Ready</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/beyond-the-canvas-the-azure-architecture-diagram-builder-becomes/ba-p/4534590</link>
      <description>&lt;P&gt;&lt;STRONG&gt;AZURE ARCHITECTURE BLOG&lt;/STRONG&gt; · 8 MIN READ&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Author:&lt;/STRONG&gt; Arturo Quiroga, Senior Partner Solutions Architect — Microsoft&lt;/P&gt;
&lt;HR /&gt;
&lt;P&gt;Two months ago I published &lt;A href="https://techcommunity.microsoft.com/blog/azurearchitectureblog/from-prompt-to-production-building-azure-architecture-diagrams-with-ai/4520336" target="_blank" rel="noopener"&gt;&lt;EM&gt;From Prompt to Production: Building Azure Architecture Diagrams with AI&lt;/EM&gt;&lt;/A&gt;, introducing the open-source &lt;A href="https://aka.ms/diagram-builder" target="_blank" rel="noopener"&gt;Azure Architecture Diagram Builder&lt;/A&gt;. The response was humbling — thousands of you read it, tried the tool, and filed issues and feature requests. A follow-up on &lt;A href="https://techcommunity.microsoft.com/post-2-waf-validation/blog-draft-waf-validation.md" target="_blank" rel="noopener"&gt;how the Well-Architected Framework scoring works&lt;/A&gt; went deep on validation.&lt;/P&gt;
&lt;P&gt;You asked, and the tool grew. This post is about what’s new since May — and one change big enough to reframe the whole project: &lt;STRONG&gt;the Azure Architecture Diagram Builder is no longer just an app you click. It’s a partner you chat with, and a tool other agents can call.&lt;/STRONG&gt;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;TL;DR.&lt;/STRONG&gt; Three arcs of new capability: (1) &lt;STRONG&gt;Architecture Chat&lt;/STRONG&gt; turns diagram design into a multi-turn conversation over the live canvas; (2) &lt;STRONG&gt;Blueprint Diagrams&lt;/STRONG&gt; produce hand-drawn, whiteboard-style deliverables alongside the formal topology; and (3) the app now exposes its capabilities as a &lt;STRONG&gt;Model Context Protocol (MCP) server&lt;/STRONG&gt;, so AI agents can generate, validate, cost, and render Azure architectures programmatically. Plus a &lt;STRONG&gt;13-model fleet&lt;/STRONG&gt;, deployment guides grounded in &lt;STRONG&gt;Microsoft Learn&lt;/STRONG&gt;, and July output enhancements.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;HR /&gt;
&lt;H2 id="whats-new-at-a-glance"&gt;What’s new at a glance&lt;/H2&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 100%; height: 510.903px; border-width: 1px;"&gt;&lt;colgroup&gt;&lt;col style="width: 49.9588%" /&gt;&lt;col style="width: 49.9588%" /&gt;&lt;/colgroup&gt;&lt;thead&gt;&lt;tr style="height: 35.1215px;"&gt;&lt;th style="height: 35.1215px;"&gt;Capability&lt;/th&gt;&lt;th&gt;
&lt;P&gt;&lt;STRONG&gt;What it does&lt;/STRONG&gt;&lt;/P&gt;
&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr style="height: 83.1424px;"&gt;&lt;td style="height: 83.1424px;"&gt;&lt;STRONG&gt;Architecture Chat&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Refine a diagram by conversation — “add Front Door with WAF,” then“now make it zone-redundant.” Each turn reads the live canvas and auto-saves to history.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 83.1424px;"&gt;&lt;td style="height: 83.1424px;"&gt;&lt;STRONG&gt;Blueprint Diagrams (BETA)&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Hand-drawn, whiteboard-style renders with nested zones and numbered&amp;nbsp;flow arrows. Topology, Blueprint, or Both.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 59.1319px;"&gt;&lt;td style="height: 59.1319px;"&gt;&lt;STRONG&gt;A fleet of 14 models&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Multi-provider roster — GPT-5.x, DeepSeek, Grok, Mistral, and Kimi — with side-by-side comparison to pick the right brain per task.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 84.0799px;"&gt;&lt;td style="height: 84.0799px;"&gt;&lt;STRONG&gt;MCP server&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;
&lt;P&gt;The app is now a remote MCP server. Agents can list_services, validate_architecture, estimate_costs, generate_bicep and render_diagram with typed, structured outputs.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 83.1424px;"&gt;&lt;td style="height: 83.1424px;"&gt;&lt;STRONG&gt;Microsoft Learn grounding&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Deployment guides now cite live Microsoft Learn documentation.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 83.1424px;"&gt;&lt;td style="height: 83.1424px;"&gt;&lt;STRONG&gt;Output enhancements (July 2026)&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Cost badges, light/dark render themes, and metadata panels in rendered diagrams.&lt;/P&gt;
&amp;nbsp;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;HR /&gt;
&lt;H2 id="from-clicking-to-conversing-architecture-chat"&gt;From clicking to conversing: Architecture Chat&lt;/H2&gt;
&lt;P&gt;The single most common request after the launch post was some version of &lt;EM&gt;“I love the first diagram, but I want to iterate without re-writing the whole prompt.”&lt;/EM&gt; Regenerating from scratch every time you tweak a requirement is slow and loses context.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Architecture Chat&lt;/STRONG&gt; solves this. It’s a conversational panel that sits alongside the canvas and treats your diagram as a living document. Each message is a turn in an ongoing design session:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;EM&gt;“Add an Azure Front Door with WAF in front of the app tier.”&lt;/EM&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;EM&gt;“Now make the data layer zone-redundant.”&lt;/EM&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;EM&gt;“Swap the SQL Database for Cosmos DB and update the connections.”&lt;/EM&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Every turn reads the &lt;STRONG&gt;current state of the canvas&lt;/STRONG&gt; — not the original prompt — so refinements compound naturally the way they would with a human architect at a whiteboard. The conversation auto-saves to history, so you can step back through the evolution of a design or branch from an earlier point.&lt;/P&gt;
&lt;FIGURE&gt;&lt;BR /&gt;
&lt;FIGCAPTION aria-hidden="true"&gt;Architecture Chat panel beside the canvas, showing a multi-turn conversation that incrementally adds and modifies services on the diagram.&lt;/FIGCAPTION&gt;
&lt;/FIGURE&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 1.&lt;/STRONG&gt; Architecture Chat treats the diagram as a living document. Each message refines the current canvas — adding services, changing SKUs, or reorganizing groups — and the full exchange is saved to history.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;The shift is subtle but important: architecture design stops being a one-shot prompt and becomes an &lt;STRONG&gt;iterative dialogue&lt;/STRONG&gt;.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="the-whiteboard-deliverable-blueprint-diagrams-beta"&gt;The whiteboard deliverable: Blueprint Diagrams (BETA)&lt;/H2&gt;
&lt;P&gt;Formal topology diagrams with official Azure icons are perfect for documentation and stakeholder decks. But early-stage design conversations often want something looser — the hand-drawn feel of a whiteboard sketch that communicates intent without implying finality.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Blueprint Diagrams&lt;/STRONG&gt; generate exactly that: a whiteboard-style render with nested zones (subscription → VNet → subnet), numbered flow arrows, and a deliberately sketchy aesthetic. You choose the output mode:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Topology&lt;/STRONG&gt; — the formal, icon-based diagram from the launch post&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Blueprint&lt;/STRONG&gt; — the hand-drawn whiteboard style&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Both&lt;/STRONG&gt; — generate the two side by side&lt;/LI&gt;
&lt;/UL&gt;
&lt;FIGURE&gt;&lt;BR /&gt;
&lt;FIGCAPTION aria-hidden="true"&gt;The formal topology diagram of an architecture shown next to a Blueprint-style hand-drawn version of the same design with nested zones and numbered flow arrows.&lt;/FIGCAPTION&gt;
&lt;/FIGURE&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 2.&lt;/STRONG&gt; The same architecture in two visual languages. &lt;STRONG&gt;Left:&lt;/STRONG&gt; the formal, icon-based topology. &lt;STRONG&gt;Right:&lt;/STRONG&gt; Blueprint mode — a whiteboard-style render with nested zones and numbered flow steps, plus a numbered legend explaining each hop. Use Blueprint for early design conversations and Topology for final documentation.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;It’s the same underlying architecture — two visual languages for two different moments in the design lifecycle.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="a-fleet-of-13-models-pick-the-right-brain-per-task"&gt;A fleet of 14 models: pick the right brain per task&lt;/H2&gt;
&lt;P&gt;The launch post shipped with multi-model support. That fleet has grown to &lt;STRONG&gt;13 models across five providers&lt;/STRONG&gt;, so you can match the model to the job — fast models for iteration, reasoning models for complex designs, code-optimized models for Bicep generation:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;OpenAI GPT-5.x&lt;/STRONG&gt; — GPT-5.1, GPT-5.2, GPT-5.6 Sol, Terra and Luna, GPT-5.4, GPT-5.4 Mini&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;DeepSeek&lt;/STRONG&gt; — V3.2 Speciale, V4 Pro&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;xAI Grok&lt;/STRONG&gt; — 4.1 Fast, 4.3&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Mistral&lt;/STRONG&gt; — Large 3&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;MoonshotAI Kimi&lt;/STRONG&gt; — K2.5, K2.7 Code&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The &lt;STRONG&gt;Compare Models&lt;/STRONG&gt; feature runs the same prompt through any subset of these in parallel and ranks them on service count, token usage, latency, and cost — with &lt;EM&gt;Fastest / Cheapest / Most Thorough&lt;/EM&gt; badges — so you can make an evidence-based choice rather than a guess.&lt;/P&gt;
&lt;FIGURE&gt;&lt;BR /&gt;
&lt;FIGCAPTION aria-hidden="true"&gt;Compare Models results grid showing side-by-side metrics across all 13 models with Fastest, Cheapest, and Most Thorough badges.&lt;/FIGCAPTION&gt;
&lt;/FIGURE&gt;
&lt;img /&gt;
&lt;FIGURE&gt;&lt;BR /&gt;
&lt;FIGCAPTION aria-hidden="true"&gt;AI Critique panel with an overall ranking and per-model analysis generated by a critic model.&lt;/FIGCAPTION&gt;
&lt;/FIGURE&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 3.&lt;/STRONG&gt; Multi-model comparison across the full 13-model fleet. &lt;STRONG&gt;Top:&lt;/STRONG&gt; the results grid ranks every model on service count, connections, token usage, latency, and cost, with &lt;EM&gt;Fastest / Cheapest / Most Thorough&lt;/EM&gt; badges. &lt;STRONG&gt;Bottom:&lt;/STRONG&gt; an optional AI Critique uses a critic model to rank the outputs and explain each model’s strengths and gaps.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;Adding a model is now a small, well-understood change — a testament to how the multi-provider abstraction has matured since May.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="the-headline-the-diagram-builder-is-now-an-mcp-server"&gt;The headline: the Diagram Builder is now an MCP server&lt;/H2&gt;
&lt;P&gt;Here’s the change that reframes the project. Everything above is about a &lt;EM&gt;person&lt;/EM&gt; using a &lt;EM&gt;web app&lt;/EM&gt;. But the same capabilities — generating a diagram, validating it against WAF, estimating its cost, producing Bicep — are exactly the things an &lt;STRONG&gt;AI agent&lt;/STRONG&gt; needs when it reasons about Azure architecture.&lt;/P&gt;
&lt;P&gt;So we exposed them. The Azure Architecture Diagram Builder now runs as a &lt;STRONG&gt;Model Context Protocol (MCP) server&lt;/STRONG&gt;. Any MCP-capable agent can call its tools with typed inputs and structured outputs:&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 98.7963%; height: 259.653px; border-width: 1px;"&gt;&lt;colgroup&gt;&lt;col style="width: 50.0054%" /&gt;&lt;col style="width: 50.0054%" /&gt;&lt;/colgroup&gt;&lt;thead&gt;&lt;tr style="height: 35.1215px;"&gt;&lt;th style="height: 35.1215px;"&gt;Tool&lt;/th&gt;&lt;th&gt;
&lt;P&gt;&lt;STRONG&gt;What the agent gets&lt;/STRONG&gt;&lt;/P&gt;
&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr style="height: 35.434px;"&gt;&lt;td style="height: 35.434px;"&gt;&lt;CODE&gt;list_services&lt;/CODE&gt;&lt;/td&gt;&lt;td&gt;
&lt;P&gt;The catalog of supported Azure services and categories&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 35.434px;"&gt;&lt;td style="height: 35.434px;"&gt;&lt;CODE&gt;validate_architecture&lt;/CODE&gt;&lt;/td&gt;&lt;td&gt;
&lt;P&gt;A WAF assessment with pillar scores and findings&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 59.1146px;"&gt;&lt;td style="height: 59.1146px;"&gt;&lt;CODE&gt;estimate_costs&lt;/CODE&gt;&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Multi-region cost estimates from the Azure Retail Prices API&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 35.434px;"&gt;&lt;td style="height: 35.434px;"&gt;&lt;CODE&gt;generate_bicep&lt;/CODE&gt;&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Infrastructure-as-Code templates for the design&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 59.1146px;"&gt;&lt;td style="height: 59.1146px;"&gt;&lt;CODE&gt;render_diagram&lt;/CODE&gt;&lt;/td&gt;&lt;td&gt;
&lt;P&gt;A rendered diagram (topology or blueprint) of the architecture&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;&lt;STRONG&gt;This means an agent can hold a conversation like &lt;EM&gt;“design a HIPAA-compliant platform, check it against the Well-Architected Framework, tell me the monthly cost in West Europe, and give me the Bicep”&lt;/EM&gt; — and the Diagram Builder answers each part programmatically, returning structured data the agent can reason over and chain.&lt;/STRONG&gt;&lt;/P&gt;
&lt;FIGURE&gt;&lt;BR /&gt;
&lt;FIGCAPTION aria-hidden="true"&gt;Microsoft Scout invoking the Diagram Builder’s render_diagram MCP tool, showing the tool-call parameters and saving the generated SVG to the workspace.&lt;/FIGCAPTION&gt;
&lt;/FIGURE&gt;
&lt;img /&gt;
&lt;FIGURE&gt;&lt;BR /&gt;
&lt;FIGCAPTION aria-hidden="true"&gt;The Azure architecture diagram rendered by the MCP tool and displayed inline in the Microsoft Scout conversation.&lt;/FIGCAPTION&gt;
&lt;/FIGURE&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 4.&lt;/STRONG&gt; The Diagram Builder as an MCP server inside &lt;STRONG&gt;Microsoft Scout&lt;/STRONG&gt;. &lt;STRONG&gt;Top:&lt;/STRONG&gt; from a natural-language request, the agent calls the &lt;CODE&gt;render_diagram&lt;/CODE&gt; tool with structured parameters (title, format, direction, theme, region) and saves the returned SVG to its workspace. &lt;STRONG&gt;Bottom:&lt;/STRONG&gt; the rendered architecture — grouped zones, labeled flows, and cost badges — appears inline in the conversation, generated entirely through agent tool calls.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;The tool that started as a canvas for humans is now also a &lt;STRONG&gt;building block for agents&lt;/STRONG&gt;. That’s the arc: &lt;EM&gt;from an app you click, to a partner you chat with, to a tool other agents call.&lt;/EM&gt;&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="grounded-in-microsoft-learn-and-sharper-output"&gt;Grounded in Microsoft Learn, and sharper output&lt;/H2&gt;
&lt;P&gt;Two smaller-but-meaningful improvements round out the release:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Microsoft Learn grounding.&lt;/STRONG&gt; Deployment guides now search official &lt;A href="https://learn.microsoft.com/" target="_blank" rel="noopener"&gt;Microsoft Learn&lt;/A&gt; documentation at generation time and cite it, so the guidance reflects current, authoritative practice rather than a model’s training snapshot.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Output enhancements (July 2026).&lt;/STRONG&gt; Rendered diagrams now carry per-service &lt;STRONG&gt;cost badges&lt;/STRONG&gt;, support &lt;STRONG&gt;light and dark render themes&lt;/STRONG&gt;, and include &lt;STRONG&gt;metadata panels&lt;/STRONG&gt; that summarize the architecture — service counts, regions, and estimated cost — directly on the image.&lt;/LI&gt;
&lt;/UL&gt;
&lt;HR /&gt;
&lt;H2 id="highlights"&gt;Highlights&lt;/H2&gt;
&lt;P&gt;Since the May launch, the Azure Architecture Diagram Builder has grown from a design tool into an agent-ready platform:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Conversational design&lt;/STRONG&gt;: iterate on a diagram by chatting over the live canvas, with full history&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Two visual languages&lt;/STRONG&gt;: formal topology and hand-drawn Blueprint, from the same architecture&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;13 models, five providers&lt;/STRONG&gt;: choose the right brain per task, with evidence-based comparison&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Agent-ready&lt;/STRONG&gt;: an MCP server exposing generation, validation, costing, and IaC as callable tools&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Grounded guidance&lt;/STRONG&gt;: deployment guides cite live Microsoft Learn documentation&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Still open source&lt;/STRONG&gt;: every capability above is available to inspect, extend, and contribute to&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2 id="try-it-today"&gt;Try It Today&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Live demo&lt;/STRONG&gt;: &lt;A href="https://aka.ms/diagram-builder" target="_blank" rel="noopener"&gt;https://aka.ms/diagram-builder&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Source code&lt;/STRONG&gt;: &lt;A href="https://github.com/Arturo-Quiroga-MSFT/azure-architecture-diagram-builder" target="_blank" rel="noopener"&gt;GitHub repository&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Documentation&lt;/STRONG&gt;: See the &lt;A href="https://github.com/Arturo-Quiroga-MSFT/azure-architecture-diagram-builder/blob/main/DOCS/getting-started-guide.md" target="_blank" rel="noopener"&gt;Getting Started Guide&lt;/A&gt; for setup, and the repository’s MCP server directory for agent integration.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;If you read the first post and tried the tool — thank you. The features above exist because you told me what you needed. Keep the feedback coming via GitHub Issues.&lt;/P&gt;
&lt;HR /&gt;
&lt;P&gt;&lt;STRONG&gt;Tags:&lt;/STRONG&gt; &lt;CODE&gt;artificial intelligence&lt;/CODE&gt; · &lt;CODE&gt;application&lt;/CODE&gt; · &lt;CODE&gt;apps &amp;amp; devops&lt;/CODE&gt; · &lt;CODE&gt;well architected&lt;/CODE&gt; · &lt;CODE&gt;infrastructure&lt;/CODE&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 14 Jul 2026 14:02:30 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/beyond-the-canvas-the-azure-architecture-diagram-builder-becomes/ba-p/4534590</guid>
      <dc:creator>arturoqu</dc:creator>
      <dc:date>2026-07-14T14:02:30Z</dc:date>
    </item>
    <item>
      <title>Golden Paths Are a Product. Treat Them Like One.</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/golden-paths-are-a-product-treat-them-like-one/ba-p/4533707</link>
      <description>&lt;P&gt;&lt;STRONG&gt;Audience:&lt;/STRONG&gt; Cloud architects, platform engineers, engineering leaders&lt;/P&gt;
&lt;HR /&gt;
&lt;P&gt;In my earlier Cloud Native Platforms articles, I focused on how teams &lt;A href="https://techcommunity.microsoft.com/blog/azurearchitectureblog/cloud-native-platforms-build/4519605" target="_blank" rel="noopener"&gt;build&lt;/A&gt;, &lt;A href="https://techcommunity.microsoft.com/blog/azurearchitectureblog/cloud-native-platforms-run/4520188" target="_blank" rel="noopener"&gt;run&lt;/A&gt;, and &lt;A href="https://techcommunity.microsoft.com/blog/azurearchitectureblog/cloud-native-platforms-evolve/4520195" target="_blank" rel="noopener"&gt;evolve&lt;/A&gt; modern engineering foundations. This article focuses on what keeps those capabilities useful after launch.&lt;/P&gt;
&lt;P&gt;Most teams do not lose adoption because the first release was poor. They lose adoption because the path was launched like a project and left to survive like a product.&lt;/P&gt;
&lt;P&gt;The pattern is familiar. A capability ships with good intent, reasonable documentation, and early enthusiasm. A few months later, teams start bypassing it, exceptions multiply, support load rises, and the platform team quietly becomes a ticket queue.&lt;/P&gt;
&lt;P&gt;If a capability is expected to be adopted, trusted, improved, and measured over time, it needs product discipline from the start.&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;How this connects to forward-deployed engineering&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Forward-deployed engineering (FDE) is a delivery model where an engineer embeds directly with a customer or team, learns the real operating context, and builds a working solution in the field. The model was &lt;A href="https://blog.palantir.com/a-day-in-the-life-of-a-palantir-forward-deployed-software-engineer-45ef2de257b1" target="_blank" rel="noopener"&gt;popularized by Palantir&lt;/A&gt;, whose engineers embed with customers. The same pattern now appears inside large engineering organizations, where platform engineers embed with internal product teams. Golden paths are a logical next step. FDE is the front of the funnel, where bespoke solutions get built one team at a time. Between the two sits a graduation decision. When the same pattern shows up across several independent teams with the same friction, it is ready to become a safe default. A golden path is the back end. The part that repeats across teams becomes that default, so you stop solving the same problem by hand every time. FDE does not stop once a path exists. New teams, new domains, and new edges keep producing forward-deployed work in parallel. The loop closes through what the path cannot absorb. Where teams bypass the path or hit a case it does not cover, those gaps become the next forward-deployed engagement, and the cycle repeats.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;EM&gt;Figure 1. Forward-deployed work is the continuous front of the funnel. Golden paths are how the repeatable parts scale.&lt;/EM&gt;&lt;/P&gt;
&lt;H2&gt;Quick context before we begin&lt;/H2&gt;
&lt;P&gt;&lt;a id="community--1-term-golden-path" class="lia-anchor"&gt;&lt;/a&gt;&lt;STRONG&gt;Golden path&lt;/STRONG&gt; is the easiest safe default way to deliver a repeated engineering outcome.&lt;/P&gt;
&lt;P&gt;&lt;a id="community--1-term-nps" class="lia-anchor"&gt;&lt;/a&gt;&lt;STRONG&gt;Net Promoter Score (NPS)&lt;/STRONG&gt; measures how likely someone is to recommend something to others. In this context, &lt;STRONG&gt;internal NPS&lt;/STRONG&gt; means how likely one internal engineering team is to recommend a capability to another internal team. It is useful as a directional trust signal, not as a standalone operating metric.&lt;/P&gt;
&lt;P&gt;&lt;a id="community--1-term-slo" class="lia-anchor"&gt;&lt;/a&gt;&lt;STRONG&gt;Service Level Objective (SLO)&lt;/STRONG&gt; is a quantitative target for a service characteristic such as availability, latency, or error rate.&lt;/P&gt;
&lt;P&gt;&lt;a id="community--1-term-sla" class="lia-anchor"&gt;&lt;/a&gt;&lt;STRONG&gt;Service Level Agreement (SLA)&lt;/STRONG&gt; is a formal service commitment, often external or contract-backed. Many internal engineering capabilities will use &lt;A title="A quantitative target for availability, latency, or error rate" rel="noopener" target="_blank"&gt;SLO&lt;/A&gt;s without using formal SLAs.&lt;/P&gt;
&lt;P&gt;&lt;a id="community--1-term-deprecation-lane" class="lia-anchor"&gt;&lt;/a&gt;&lt;STRONG&gt;Deprecation lane&lt;/STRONG&gt; is the planned and time-bound way to retire an older pattern without causing unnecessary disruption.&lt;/P&gt;
&lt;H2&gt;How to identify a true golden path&lt;/H2&gt;
&lt;P&gt;Not every good practice is a golden path. A path qualifies when most of the following are true:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Repeated use case across teams&lt;/LI&gt;
&lt;LI&gt;High cost of inconsistency&lt;/LI&gt;
&lt;LI&gt;High cognitive load from scratch&lt;/LI&gt;
&lt;LI&gt;A safe default can be defined&lt;/LI&gt;
&lt;LI&gt;Measurable outcome exists&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;If those signals are absent, teams are usually looking at a local implementation pattern, not a &lt;A title="The easiest safe default way to deliver a repeated engineering outcome" rel="noopener" target="_blank"&gt;golden path&lt;/A&gt;.&lt;/P&gt;
&lt;P&gt;That distinction matters. Internal capability teams need to invest their product energy where standardization meaningfully reduces risk, cost, delay, or complexity.&lt;/P&gt;
&lt;H2&gt;Two transformation threads that make this real&lt;/H2&gt;
&lt;P&gt;To keep this practical, this article refers back to two transformation threads throughout the discussion.&lt;/P&gt;
&lt;P&gt;The first was an AKS migration effort. The existing hosting model worked until scale, operational control, and predictable performance started pulling in opposite directions. The move to AKS was not treated as a one-time infrastructure shift. Workloads were phased over time, traffic was controlled through a single API entry point, and teams were trained to operate in a container-first model before the migration accelerated. The outcome was not only better control. It created a stronger foundation for observability, scaling, policy-driven operations, and future platform standardization.&lt;/P&gt;
&lt;P&gt;The second was a security modernization effort around a more unified authentication and authorization model. Identity logic had grown across services over time. Permissions were fragmented. Operational inconsistency was becoming risk. The move toward a more centralized approach depended as much on rollout safety as on design quality. Feature flags, parallel writes during transition, read-repair, and rollback paths made that shift survivable on a live platform while preserving performance and user continuity.&lt;/P&gt;
&lt;P&gt;These two threads matter because they show the same principle in different settings. A capability becomes durable when teams combine the right architecture with the right product mindset, operating model, and feedback loop.&lt;/P&gt;
&lt;H2&gt;The golden path operating canvas&lt;/H2&gt;
&lt;P&gt;A useful way to think about a &lt;A title="The easiest safe default way to deliver a repeated engineering outcome" rel="noopener" target="_blank"&gt;golden path&lt;/A&gt; is through five decision areas:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Product framing&lt;/LI&gt;
&lt;LI&gt;Operating model&lt;/LI&gt;
&lt;LI&gt;Engineering guardrails&lt;/LI&gt;
&lt;LI&gt;Adoption and change strategy&lt;/LI&gt;
&lt;LI&gt;Measurement and feedback loop&lt;/LI&gt;
&lt;/OL&gt;
&lt;img /&gt;
&lt;P&gt;&lt;EM&gt;Figure 2. The measurement and feedback area is what keeps the other four honest over time.&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;Individually, each area matters. Together, they determine whether a capability becomes the default, stays trusted, and improves over time.&lt;/P&gt;
&lt;H2&gt;1. Product framing&lt;/H2&gt;
&lt;P&gt;A common first mistake is treating a capability as an output instead of a product.&lt;/P&gt;
&lt;P&gt;Product framing starts with four simple questions:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Who is the consumer&lt;/LI&gt;
&lt;LI&gt;What repeated problem are we solving&lt;/LI&gt;
&lt;LI&gt;What safe default are we standardizing&lt;/LI&gt;
&lt;LI&gt;What measurable outcome proves value&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;This matters equally for a platform capability and for a shared service. An internal deployment template, a security validation path, a standardized API onboarding flow, or the opinionated consumption path around a reusable identity service can all become &lt;A title="The easiest safe default way to deliver a repeated engineering outcome" rel="noopener" target="_blank"&gt;golden paths&lt;/A&gt; if they remove repeated friction and create better defaults.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The path gets clearer when teams name the repeated decision they want to remove. In the AKS migration, the real question was not "should we modernize hosting." It was "what is the safest default for workloads that now need predictable performance and stronger operational control." That framing forces sharper scoping. It also prevents vague launch language from hiding the actual default teams are expected to adopt.&lt;/P&gt;
&lt;P&gt;If teams are defining a new &lt;A title="The easiest safe default way to deliver a repeated engineering outcome" rel="noopener" target="_blank"&gt;golden path&lt;/A&gt;, product framing is where they start.&lt;/P&gt;
&lt;P&gt;If teams already have one and see drift, they should return here first and ask whether the original consumer problem is still sharp, still current, and still visible.&lt;/P&gt;
&lt;H2&gt;2. Operating model&lt;/H2&gt;
&lt;P&gt;A &lt;A title="The easiest safe default way to deliver a repeated engineering outcome" rel="noopener" target="_blank"&gt;golden path&lt;/A&gt; does not stay healthy because documentation exists. It stays healthy because ownership exists.&lt;/P&gt;
&lt;P&gt;The operating model answers questions like these:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Who owns the path&lt;/LI&gt;
&lt;LI&gt;Who contributes to it&lt;/LI&gt;
&lt;LI&gt;Who makes change decisions&lt;/LI&gt;
&lt;LI&gt;What review cadence exists&lt;/LI&gt;
&lt;LI&gt;How support expectations are communicated&lt;/LI&gt;
&lt;LI&gt;How exceptions are granted and retired&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Without this, teams end up with a queue, not a product.&lt;/P&gt;
&lt;P&gt;This is where many internal capabilities weaken over time. Platform teams become catch-all owners. Security teams review too late. Consumer teams give feedback, but no one is clearly accountable for converting that feedback into decisions.&lt;/P&gt;
&lt;P&gt;Team Topologies describes a similar problem through cognitive load and enabling-team boundaries. When ownership and interaction modes are unclear, friction rises and good patterns fail to scale (&lt;A href="https://teamtopologies.com/" target="_blank" rel="noopener"&gt;Team Topologies&lt;/A&gt;).&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works is simple: one accountable owner, one recurring decision forum, and one visible exception model. The exception model matters more than teams expect. A path becomes brittle when consumers can go off-path with no review, and it becomes unusable when every exception requires escalation with no expiry. Temporary exceptions need an owner, a review date, and a retirement path.&lt;/P&gt;
&lt;P&gt;If teams are starting new, they should define one accountable owner and one recurring review rhythm.&lt;/P&gt;
&lt;P&gt;If teams already see drift, they should look for ambiguity in decision rights, ownership, or support expectations.&lt;/P&gt;
&lt;H2&gt;3. Engineering guardrails&lt;/H2&gt;
&lt;P&gt;A &lt;A title="The easiest safe default way to deliver a repeated engineering outcome" rel="noopener" target="_blank"&gt;golden path&lt;/A&gt; is only valuable if it makes the safe path the easy path.&lt;/P&gt;
&lt;P&gt;Engineering guardrails are the non-negotiables built into the path itself. These typically include:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Security posture&lt;/LI&gt;
&lt;LI&gt;Reliability targets&lt;/LI&gt;
&lt;LI&gt;Cost boundaries&lt;/LI&gt;
&lt;LI&gt;Operational readiness&lt;/LI&gt;
&lt;LI&gt;Observability and telemetry standards&lt;/LI&gt;
&lt;LI&gt;Supportability expectations&lt;/LI&gt;
&lt;LI&gt;Versioning and compatibility expectations&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;This is one place where product, service, and platform thinking often get separated unnecessarily. They should not be. A shared identity service needs guardrails just as much as a platform deployment path does. A platform path without operational clarity is as incomplete as a service without security consistency.&lt;/P&gt;
&lt;P&gt;The unified auth work shows why. Centralizing identity enforcement was not just an architectural choice. It was a way to make security control more consistent, measurable, and operable. Feature-flag strategy, rollback readiness, and performance targets were all part of the guardrail model, not postscript concerns.&lt;/P&gt;
&lt;P&gt;Microsoft's Well-Architected guidance reinforces the same principle. Security, reliability, and cost are not side dimensions. They are core decision lenses that shape how capabilities are designed and operated (&lt;A href="https://learn.microsoft.com/azure/well-architected/" target="_blank" rel="noopener"&gt;Microsoft Well-Architected Framework&lt;/A&gt;).&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The unified auth work makes this concrete. The technical design mattered, but the real strength came from turning security rules into operating defaults: central enforcement, predictable rollout controls, and a clear rollback posture. The same principle applies elsewhere. If teams do not know the compatibility expectations, cost envelope, or minimum operational standard up front, they invent their own.&lt;/P&gt;
&lt;P&gt;If teams are defining a new path, they should decide the minimum guardrails before expansion.&lt;/P&gt;
&lt;P&gt;If an existing path is drifting, teams should look for where consumers are stepping outside the guardrails or where the guardrails were never made explicit.&lt;/P&gt;
&lt;H2&gt;4. Adoption and change strategy&lt;/H2&gt;
&lt;P&gt;A &lt;A title="The easiest safe default way to deliver a repeated engineering outcome" rel="noopener" target="_blank"&gt;golden path&lt;/A&gt; does not become real because it is published. It becomes real because teams can adopt it safely and progressively.&lt;/P&gt;
&lt;P&gt;That means rollout strategy must be designed as carefully as the solution itself. That includes:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Migration sequencing&lt;/LI&gt;
&lt;LI&gt;Consumer onboarding&lt;/LI&gt;
&lt;LI&gt;Enablement and training&lt;/LI&gt;
&lt;LI&gt;Shadow testing where needed&lt;/LI&gt;
&lt;LI&gt;Rollback readiness&lt;/LI&gt;
&lt;LI&gt;Transition governance&lt;/LI&gt;
&lt;LI&gt;Minimum viable self-service&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;This was one of the clearest lessons from the AKS migration. The technical work mattered, but the people ramp-up mattered just as much. Moving a large engineering group from one operational model to another required structured enablement, practical migration guidance, and a way to build confidence in steps rather than all at once.&lt;/P&gt;
&lt;P&gt;This mirrors the intent behind Microsoft's Cloud Adoption Framework. Large-scale change lands better when teams combine technical design with lifecycle planning, capability enablement, and operating model clarity (&lt;A href="https://learn.microsoft.com/azure/cloud-adoption-framework/" target="_blank" rel="noopener"&gt;Microsoft Cloud Adoption Framework&lt;/A&gt;).&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The AKS migration worked because adoption was designed, not assumed. Teams had a phased route in, a safe path back, and enough enablement to operate the new model with confidence. Just as important, the path moved toward self-service instead of permanent dependency on the central team. If a so-called path still requires repeated manual intervention for normal use, it is closer to a managed service queue than a true default.&lt;/P&gt;
&lt;P&gt;If teams are building a new path, they should design adoption as a first-class workstream.&lt;/P&gt;
&lt;P&gt;If teams already have a path that is drifting, they should treat re-adoption as a change program, not a documentation refresh.&lt;/P&gt;
&lt;H2&gt;5. Measurement and the feedback loop&lt;/H2&gt;
&lt;P&gt;If feedback does not change decisions, feedback collection becomes theater.&lt;/P&gt;
&lt;P&gt;A &lt;A title="The easiest safe default way to deliver a repeated engineering outcome" rel="noopener" target="_blank"&gt;golden path&lt;/A&gt; needs a small, meaningful set of signals that tells teams whether it is healthy. That usually includes some mix of:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Adoption depth&lt;/LI&gt;
&lt;LI&gt;Time to first success&lt;/LI&gt;
&lt;LI&gt;On-path versus off-path incidents&lt;/LI&gt;
&lt;LI&gt;Cost efficiency trend&lt;/LI&gt;
&lt;LI&gt;Reliability trend against &lt;A title="Service Level Objective" rel="noopener" target="_blank"&gt;SLO&lt;/A&gt;s&lt;/LI&gt;
&lt;LI&gt;Internal &lt;A title="Net Promoter Score" rel="noopener" target="_blank"&gt;NPS&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Deprecation progress&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The metric set does not need to be large. It needs to be actionable.&lt;/P&gt;
&lt;P&gt;Internal &lt;A title="Net Promoter Score" rel="noopener" target="_blank"&gt;NPS&lt;/A&gt; is especially useful when paired with qualitative input. A team may still be using a capability while quietly warning other teams away from it. That is an early drift signal. It is still only one signal. Mandatory adoption, small sample sizes, and local politics can all distort it.&lt;/P&gt;
&lt;P&gt;DORA research reinforces the broader point that improvement depends on disciplined measurement and learning loops, not on collecting more indicators than teams can act on (&lt;A href="https://cloud.google.com/devops" target="_blank" rel="noopener"&gt;DORA research&lt;/A&gt;).&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The most useful feedback loop is visible. Teams should be able to see what was heard, what changed, and what did not make the cut. That is what turns trust into an operating behavior instead of a sentiment score. Without that trace, even good metrics become noise.&lt;/P&gt;
&lt;P&gt;If teams are launching a new path, they should define baseline metrics before expansion.&lt;/P&gt;
&lt;P&gt;If teams already have a path and see drift, they should run a focused diagnostic on adoption, incidents, bypass behavior, and team sentiment before redesigning anything.&lt;/P&gt;
&lt;H2&gt;The mindset shift that makes this work&lt;/H2&gt;
&lt;P&gt;Project thinking asks whether the capability shipped.&lt;/P&gt;
&lt;P&gt;Product thinking asks whether it is being adopted, trusted, improved, and measured over time.&lt;/P&gt;
&lt;P&gt;That shift changes how teams define success, stage rollout, involve consumers, and respond to drift. It applies to internal platforms, shared services, reusable components, and enablement models. The operating details are not identical across all of them, but the product discipline is still useful wherever teams expect repeat consumers and recurring outcomes.&lt;/P&gt;
&lt;P&gt;The CNCF Platforms white paper captures a related idea through paved paths and shared internal capabilities. The technical foundation matters, but the sustained value comes from how teams shape adoption, ownership, and evolution around it (&lt;A href="https://tag-app-delivery.cncf.io/whitepapers/platforms/" target="_blank" rel="noopener"&gt;CNCF Platforms White Paper&lt;/A&gt;).&lt;/P&gt;
&lt;P&gt;That is the broader point of this article. Golden path thinking is not just for platform teams. It is a repeatability and trust model for internal engineering capability design.&lt;/P&gt;
&lt;H2&gt;Signals your golden path is drifting&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;Adoption stalls or declines&lt;/LI&gt;
&lt;LI&gt;Side paths become normal&lt;/LI&gt;
&lt;LI&gt;Support load rises while roadmap clarity drops&lt;/LI&gt;
&lt;LI&gt;Consumers cannot name the owner&lt;/LI&gt;
&lt;LI&gt;Breaking changes surprise teams&lt;/LI&gt;
&lt;LI&gt;Feedback is collected with no decision trail&lt;/LI&gt;
&lt;LI&gt;Teams still use the path, but would not recommend it&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;These are usually not random symptoms. They point back to weak framing, weak ownership, weak guardrails, weak adoption design, or a weak feedback loop.&lt;/P&gt;
&lt;H2&gt;Immediate next steps&lt;/H2&gt;
&lt;H3&gt;If you are defining a new golden path&lt;/H3&gt;
&lt;P&gt;Start where you already embed. Your forward-deployed and high-touch engagements are the best source of golden path candidates.&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Pick one repeated, high-friction workflow with high risk if implemented inconsistently.&lt;/LI&gt;
&lt;LI&gt;Define the consumer, safe default, and one measurable success outcome.&lt;/LI&gt;
&lt;LI&gt;Set minimum guardrails for security, reliability, observability, cost, compatibility, and rollback.&lt;/LI&gt;
&lt;LI&gt;Create a 30-day pilot with a named owner, a review cadence, and a visible exception path.&lt;/LI&gt;
&lt;LI&gt;Capture internal &lt;A title="Net Promoter Score" rel="noopener" target="_blank"&gt;NPS&lt;/A&gt; plus qualitative feedback, then publish the first improvement decision.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H3&gt;If you already have a golden path and see drift&lt;/H3&gt;
&lt;OL&gt;
&lt;LI&gt;Identify the top drift signals, such as bypass behavior, adoption drop, or trust erosion.&lt;/LI&gt;
&lt;LI&gt;Diagnose which area is weak: framing, ownership, guardrails, adoption, or feedback loop.&lt;/LI&gt;
&lt;LI&gt;Run a focused 30-day correction cycle rather than a broad redesign.&lt;/LI&gt;
&lt;LI&gt;Reset one decision forum, one owner, and one exception model before changing anything larger.&lt;/LI&gt;
&lt;LI&gt;Publish what changed, then re-measure adoption, sentiment, and operational metrics after one cycle.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2&gt;Key takeaways&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;Golden paths reduce repeated risk and effort&lt;/LI&gt;
&lt;LI&gt;Product discipline keeps them useful&lt;/LI&gt;
&lt;LI&gt;Product discipline applies beyond platforms&lt;/LI&gt;
&lt;LI&gt;Safe defaults need explicit guardrails&lt;/LI&gt;
&lt;LI&gt;Trust grows when feedback changes decisions&lt;/LI&gt;
&lt;LI&gt;Golden paths are how the repeatable parts of forward-deployed work scale&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;What this means&lt;/H2&gt;
&lt;P&gt;Good engineering capabilities do not fail only because the technology is weak. They often fail because the operating model around them is too thin.&lt;/P&gt;
&lt;P&gt;That is why treating a &lt;A title="The easiest safe default way to deliver a repeated engineering outcome" rel="noopener" target="_blank"&gt;golden path&lt;/A&gt; like a product matters. It turns a useful pattern into a dependable default.&lt;/P&gt;
&lt;P&gt;It also creates a stronger bridge back to the earlier Cloud Native Platforms &lt;A href="https://techcommunity.microsoft.com/blog/azurearchitectureblog/cloud-native-platforms-run/4520188" target="_blank" rel="noopener"&gt;Run article&lt;/A&gt; and &lt;A href="https://techcommunity.microsoft.com/blog/azurearchitectureblog/cloud-native-platforms-evolve/4520195" target="_blank" rel="noopener"&gt;Evolve article&lt;/A&gt;. Building, running, and evolving technical foundations is one part of the problem. Sustaining trust and adoption around those foundations is the next one.&lt;/P&gt;
&lt;P&gt;The real shift is this: a golden path is not complete when it launches. It is complete when teams still trust it after the first wave of exceptions, change pressure, and scale.&lt;/P&gt;
&lt;HR /&gt;
&lt;P&gt;&lt;STRONG&gt;Want to discuss?&lt;/STRONG&gt; Drop a comment with patterns you have seen in your environment. I read every reply.&lt;/P&gt;</description>
      <pubDate>Tue, 07 Jul 2026 18:26:04 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/golden-paths-are-a-product-treat-them-like-one/ba-p/4533707</guid>
      <dc:creator>KishoreKumarPattabiraman</dc:creator>
      <dc:date>2026-07-07T18:26:04Z</dc:date>
    </item>
    <item>
      <title>Optimizing GitHub Copilot Cost in the Usage-Based Billing Era</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/optimizing-github-copilot-cost-in-the-usage-based-billing-era/ba-p/4534171</link>
      <description>&lt;P&gt;GitHub Copilot has become a core part of the modern developer workflow. We use it to complete code, explain unfamiliar repositories, write tests, refactor legacy applications, review pull requests, generate documentation, and automate repetitive engineering work.&lt;/P&gt;
&lt;P&gt;But as GitHub Copilot moves further into usage-based billing, teams and users are asking a very practical question:&lt;/P&gt;
&lt;P&gt;How do we keep getting value from GitHub Copilot without letting costs become unpredictable?&lt;/P&gt;
&lt;P&gt;The answer is not to blindly slash GitHub Copilot usage—that would defeat the purpose of adopting AI-assisted development in the first place. The better answer is to use GitHub Copilot more intentionally.&lt;/P&gt;
&lt;P&gt;Usage-based billing shifts our mindset from:&lt;/P&gt;
&lt;P&gt;“Can I use GitHub Copilot?” to: “Am I using the right GitHub Copilot capability, with the right model, the right context, and the right level of automation for this specific task?”&lt;/P&gt;
&lt;P&gt;This guide outlines the major ways developers and organizations can optimize GitHub Copilot costs while continuing to accelerate productivity.&lt;/P&gt;
&lt;P&gt;For a strategic look at how organizations are scaling these financial practices across modern infrastructure, read the playbook &lt;A href="https://techcommunity.microsoft.com/blog/azuredevcommunityblog/token-economics-the-new-finops-for-agentic-ai/4533743" target="_blank"&gt;Token Economics: The New FinOps for Agentic AI | Microsoft Community Hub&lt;/A&gt;&lt;/P&gt;
&lt;H2 data-path-to-node="11"&gt;Understanding the New Cost Model&lt;/H2&gt;
&lt;P data-path-to-node="12"&gt;Under usage-based billing, GitHub Copilot costs are driven primarily by two factors:&lt;/P&gt;
&lt;OL data-path-to-node="13"&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="13,0,0" data-index-in-node="0"&gt;The model used&lt;/STRONG&gt; (e.g., frontier vs. lightweight models)&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="13,1,0" data-index-in-node="0"&gt;The number of tokens consumed&lt;/STRONG&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;P data-path-to-node="14"&gt;A token is a small unit of text processed by the AI. This includes what you send to the model (input), what the model generates back (output), and, in some cases, cached context that the model reuses.&lt;/P&gt;
&lt;P data-path-to-node="15"&gt;This means a short question using a lightweight model consumes very little. Conversely, a long, multi-file agentic session using a frontier model can consume significantly more. GitHub Copilot cost is no longer just about how many people have seats—&lt;STRONG data-path-to-node="15" data-index-in-node="249"&gt;it is about how those seats are being used.&lt;/STRONG&gt;&lt;/P&gt;
&lt;H3 data-path-to-node="16"&gt;The Biggest Cost Drivers:&lt;/H3&gt;
&lt;UL data-path-to-node="17"&gt;
&lt;LI&gt;Using expensive models for simple tasks&lt;/LI&gt;
&lt;LI&gt;Overly large context windows&lt;/LI&gt;
&lt;LI&gt;Long chat sessions with repeated context&lt;/LI&gt;
&lt;LI&gt;Agentic workflows that inspect too many files&lt;/LI&gt;
&lt;LI&gt;Enabling tools when they aren't needed&lt;/LI&gt;
&lt;LI&gt;Dumping large terminal logs or full repository scans into the chat&lt;/LI&gt;
&lt;LI&gt;A lack of active budgets and usage monitoring&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-path-to-node="18"&gt;The good news? &lt;STRONG data-path-to-node="18" data-index-in-node="15"&gt;Most of these are completely controllable.&lt;/STRONG&gt;&lt;/P&gt;
&lt;H2 data-path-to-node="20"&gt;1. Set Spending Guardrails First&lt;/H2&gt;
&lt;P data-path-to-node="21"&gt;The first cost-reduction step isn't technical—it’s financial governance. Before teams start heavy usage, set spending caps and budget controls. This is critical for organizations where many developers share a pool of AI credits.&lt;/P&gt;
&lt;H3 data-path-to-node="22"&gt;Recommended Actions:&lt;/H3&gt;
&lt;UL data-path-to-node="23"&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="23,0,0" data-index-in-node="0"&gt;Set a monthly budget&lt;/STRONG&gt; and enable hard-stop controls where available.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="23,1,0" data-index-in-node="0"&gt;Segment limits:&lt;/STRONG&gt; Create separate limits for standard users and power users.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="23,2,0" data-index-in-node="0"&gt;Track granularly:&lt;/STRONG&gt; Monitor consumption by user, team, organization, or cost center.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="23,3,0" data-index-in-node="0"&gt;Review early:&lt;/STRONG&gt; Closely analyze the first few weeks of usage and adjust limits based on real consumption patterns.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-path-to-node="24"&gt;This prevents a single heavy user, an runaway automated agent session, or an experimental workflow from consuming a massive chunk of shared capacity early in the billing cycle.&lt;/P&gt;
&lt;P data-path-to-node="25"&gt;For enterprises, &lt;STRONG data-path-to-node="25" data-index-in-node="17"&gt;avoid a one-size-fits-all budget.&lt;/STRONG&gt; A junior developer asking occasional chat questions does not need the same budget as a platform engineer running massive migration tasks across multiple repositories.&lt;/P&gt;
&lt;H3 data-path-to-node="26"&gt;A Smarter Budgeting Model:&lt;/H3&gt;
&lt;UL data-path-to-node="27"&gt;
&lt;LI&gt;Standard developer budget&lt;/LI&gt;
&lt;LI&gt;Power user budget&lt;/LI&gt;
&lt;LI&gt;Pilot team budget&lt;/LI&gt;
&lt;LI&gt;Agentic workflow budget&lt;/LI&gt;
&lt;LI&gt;Innovation/experimentation budget&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-path-to-node="28"&gt;This gives leaders control without blocking productivity, allowing organizations to analyze and adjust budgets frequently.&lt;/P&gt;
&lt;H2 data-path-to-node="30"&gt;2. Use the Right Model for the Right Task&lt;/H2&gt;
&lt;P data-path-to-node="31"&gt;&lt;STRONG data-path-to-node="31" data-index-in-node="0"&gt;Not every task requires the most powerful model.&lt;/STRONG&gt; This is one of the biggest mindset shifts in usage-based billing. Developers often default to the strongest model because it feels "safer," but a smaller, cheaper model is often more than enough for daily tasks.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Use Lightweight / Lower-Cost Models For:&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;&lt;STRONG&gt;Use Stronger / Premium Models For:&lt;/STRONG&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;SPAN data-path-to-node="32,1,0,0"&gt;Simple code explanations&lt;/SPAN&gt;&lt;/td&gt;&lt;td&gt;&lt;SPAN data-path-to-node="32,1,1,0"&gt;Architecture &amp;amp; system design decisions&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;SPAN data-path-to-node="32,2,0,0"&gt;Documentation &amp;amp; comment updates&lt;/SPAN&gt;&lt;/td&gt;&lt;td&gt;&lt;SPAN data-path-to-node="32,2,1,0"&gt;Complex debugging &amp;amp; multi-file reasoning&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;SPAN data-path-to-node="32,3,0,0"&gt;Small unit test generation&lt;/SPAN&gt;&lt;/td&gt;&lt;td&gt;&lt;SPAN data-path-to-node="32,3,1,0"&gt;Security-sensitive reviews&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;SPAN data-path-to-node="32,4,0,0"&gt;Basic debugging &amp;amp; syntax help&lt;/SPAN&gt;&lt;/td&gt;&lt;td&gt;&lt;SPAN data-path-to-node="32,4,1,0"&gt;Legacy modernization &amp;amp; code translation&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;SPAN data-path-to-node="32,5,0,0"&gt;Regex help &amp;amp; simple refactoring&lt;/SPAN&gt;&lt;/td&gt;&lt;td&gt;&lt;SPAN data-path-to-node="32,5,1,0"&gt;Performance tuning &amp;amp; optimization&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;SPAN data-path-to-node="32,6,0,0"&gt;Translating code from one style to another&lt;/SPAN&gt;&lt;/td&gt;&lt;td&gt;&lt;SPAN data-path-to-node="32,6,1,0"&gt;Ambiguous, cross-service production issues&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;SPAN data-path-to-node="32,7,0,0"&gt;Summarizing small, single files&lt;/SPAN&gt;&lt;/td&gt;&lt;td&gt;&lt;SPAN data-path-to-node="32,7,1,0"&gt;Large pull request reviews&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H3 data-path-to-node="33"&gt;Team Golden Rule:&lt;/H3&gt;
&lt;P data-path-to-node="34,0"&gt;Start with the cheapest model that can do the task well. Move to a stronger model &lt;EM data-path-to-node="34,0" data-index-in-node="82"&gt;only&lt;/EM&gt; when the task requires deeper reasoning. This keeps high-cost models reserved for where they provide the highest value.&lt;/P&gt;
&lt;H2 data-path-to-node="36"&gt;3. Lean More on Code Completions&lt;/H2&gt;
&lt;P data-path-to-node="37"&gt;Inline code completions and next edit suggestions are still some of the most cost-effective GitHub Copilot experiences available. For paid GitHub Copilot plans, these inline experiences &lt;STRONG data-path-to-node="37" data-index-in-node="186"&gt;are not billed against your AI credits.&lt;/STRONG&gt; Developers should aggressively lean on completions when they already know &lt;EM data-path-to-node="37" data-index-in-node="300"&gt;what&lt;/EM&gt; they want to build.&lt;/P&gt;
&lt;H3 data-path-to-node="38"&gt;Ideal Completion Scenarios:&lt;/H3&gt;
&lt;UL data-path-to-node="39"&gt;
&lt;LI&gt;Writing a function after creating the signature&lt;/LI&gt;
&lt;LI&gt;Completing repetitive boilerplate code&lt;/LI&gt;
&lt;LI&gt;Filling in obvious implementation logic or adding similar methods&lt;/LI&gt;
&lt;LI&gt;Writing simple, predictable unit tests&lt;/LI&gt;
&lt;LI&gt;Completing configuration files&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-path-to-node="40,0"&gt;&lt;STRONG data-path-to-node="40,0" data-index-in-node="0"&gt;The Rule of Thumb:&lt;/STRONG&gt; Use chat when you need &lt;EM data-path-to-node="40,0" data-index-in-node="42"&gt;reasoning&lt;/EM&gt;. Use completions when you need &lt;EM data-path-to-node="40,0" data-index-in-node="83"&gt;acceleration&lt;/EM&gt;.&lt;/P&gt;
&lt;UL data-path-to-node="41"&gt;
&lt;LI&gt;❌ &lt;STRONG data-path-to-node="41,0,0" data-index-in-node="2"&gt;Bad Habit:&lt;/STRONG&gt; Opening a new chat window for every small function.&lt;/LI&gt;
&lt;LI&gt;👉 &lt;STRONG data-path-to-node="41,1,0" data-index-in-node="3"&gt;Better Habit:&lt;/STRONG&gt; Write the function name, a descriptive comment, or the first few lines, and let GitHub Copilot fill in the implementation inline. Open chat only if the completion misses the mark or the logic requires deeper discussion.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2 data-path-to-node="43"&gt;4. Use Tools Only When Needed&lt;/H2&gt;
&lt;P data-path-to-node="44"&gt;Agents are incredibly powerful because they can interact with tools: reading files, searching repositories, editing code, running commands, inspecting test failures, or connecting to external systems via Model Context Protocol (MCP) servers.&lt;/P&gt;
&lt;P data-path-to-node="45"&gt;However, &lt;STRONG data-path-to-node="45" data-index-in-node="9"&gt;tools heavily amplify costs.&lt;/STRONG&gt; Every tool call adds more context, more model invocations, and more output back into the token loop. If an agent reads too many files or dumps raw logs into chat, token counts skyrocket.&lt;/P&gt;
&lt;P data-path-to-node="46"&gt;The goal is not to ban tools, but to &lt;STRONG data-path-to-node="46" data-index-in-node="37"&gt;scope them correctly&lt;/STRONG&gt;:&lt;/P&gt;
&lt;UL data-path-to-node="47"&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="47,0,0" data-index-in-node="0"&gt;Don't enable everything:&lt;/STRONG&gt; Avoid activating every tool for every agent. Use only what is required.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="47,1,0" data-index-in-node="0"&gt;Specialized agents:&lt;/STRONG&gt; Create custom agents with specialized tool access (e.g., a documentation agent needs read/search tools, but likely doesn't need terminal or cloud database access).&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="47,2,0" data-index-in-node="0"&gt;Avoid full scans:&lt;/STRONG&gt; Ask Copilot to inspect specific files or folders rather than scanning the entire repository.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="47,3,0" data-index-in-node="0"&gt;Summarize logs:&lt;/STRONG&gt; Ask Copilot to summarize terminal outputs or run the narrowest relevant unit test instead of dumping pages of raw test logs.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3 data-path-to-node="48"&gt;A Cost-Aware Prompting Example:&lt;/H3&gt;
&lt;P data-path-to-node="49,0"&gt;&lt;EM data-path-to-node="49,0" data-index-in-node="0"&gt;"Use tools only if needed. Start by reading the files I mention. Do not scan the entire repository. If tests are needed, run only the relevant unit tests and summarize the failure instead of pasting full logs."&lt;/EM&gt;&lt;/P&gt;
&lt;H2 data-path-to-node="51"&gt;5. Narrow the Context Window&lt;/H2&gt;
&lt;P data-path-to-node="52"&gt;Context is useful only when it is relevant. A common cost driver is issuing broad prompts that pull massive chunks of irrelevant code into the session.&lt;/P&gt;
&lt;UL data-path-to-node="53"&gt;
&lt;LI&gt;❌ &lt;STRONG data-path-to-node="53,0,0" data-index-in-node="2"&gt;Broad Prompt:&lt;/STRONG&gt; &lt;EM data-path-to-node="53,0,0" data-index-in-node="16"&gt;"Analyze this repo and fix the bug."&lt;/EM&gt; (Forces the agent to scan, search, and load unnecessary files).&lt;/LI&gt;
&lt;LI&gt;👉 &lt;STRONG data-path-to-node="53,1,0" data-index-in-node="3"&gt;Targeted Prompt:&lt;/STRONG&gt; &lt;EM data-path-to-node="53,1,0" data-index-in-node="20"&gt;"The bug appears to be in /src/auth/tokenValidator.ts and /src/auth/sessionStore.ts. Review only these files and their related tests. Propose the smallest safe fix."&lt;/EM&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3 data-path-to-node="54"&gt;Practical Ways to Reduce Context:&lt;/H3&gt;
&lt;UL data-path-to-node="55"&gt;
&lt;LI&gt;Explicitly mention exact files, folders, error messages, or failing tests.&lt;/LI&gt;
&lt;LI&gt;Instruct the model on what &lt;EM data-path-to-node="55,1,0" data-index-in-node="27"&gt;not&lt;/EM&gt; to change.&lt;/LI&gt;
&lt;LI&gt;Explicitly exclude generated files, vendor folders, build artifacts, and lock files.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="55,3,0" data-index-in-node="0"&gt;Clear your sessions:&lt;/STRONG&gt; Staying in one massive chat session all day causes historical context to stack up. Start a fresh session when your topic/intent changes, when frequent compactions occur, or when different tools are required.&lt;/LI&gt;
&lt;LI&gt;&lt;EM data-path-to-node="55,4,0" data-index-in-node="0"&gt;Pro tip:&lt;/EM&gt; If you are ending a long session but need to carry forward the conclusion, ask Copilot to generate a quick Markdown summary of the session to paste as the starting context of your new, clean session.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2 data-path-to-node="57"&gt;6. Use Plan-First Prompts for Complex Work&lt;/H2&gt;
&lt;P data-path-to-node="58"&gt;For large-scale tasks, do not let GitHub Copilot immediately start editing files in agent mode. &lt;STRONG data-path-to-node="58" data-index-in-node="96"&gt;Always ask it to plan first.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P data-path-to-node="59"&gt;This is crucial for large refactoring, multi-file changes, legacy modernization, dependency upgrades, security fixes, and performance tuning.&lt;/P&gt;
&lt;H3 data-path-to-node="60"&gt;A Plan-First Prompting Pattern:&lt;/H3&gt;
&lt;P data-path-to-node="61,0"&gt;&lt;EM data-path-to-node="61,0" data-index-in-node="0"&gt;"Before editing files, create a short plan. Identify the files you need to inspect, the likely root cause, the risk level, and the smallest validation test. Do not make changes until the plan is clear."&lt;/EM&gt;&lt;/P&gt;
&lt;H3 data-path-to-node="62"&gt;Why this works:&lt;/H3&gt;
&lt;OL data-path-to-node="63"&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="63,0,0" data-index-in-node="0"&gt;Prevents wasted work:&lt;/STRONG&gt; It ensures the agent doesn't modify the wrong files or run unnecessary commands.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="63,1,0" data-index-in-node="0"&gt;Creates a cost checkpoint:&lt;/STRONG&gt; If the plan looks too broad or incorrect, you can narrow the scope &lt;EM data-path-to-node="63,1,0" data-index-in-node="94"&gt;before&lt;/EM&gt; GitHub Copilot kicks off an expensive, multi-step agentic loop.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2 data-path-to-node="65"&gt;7. Batch Related Questions in One Session&lt;/H2&gt;
&lt;P data-path-to-node="66"&gt;While giant, day-long sessions are bad for context bloat, opening a brand-new chat for every tiny, consecutive question introduces unnecessary context reload overhead. Finding the right balance is key.&lt;/P&gt;
&lt;UL data-path-to-node="67"&gt;
&lt;LI&gt;❌ &lt;STRONG data-path-to-node="67,0,0" data-index-in-node="2"&gt;Less Efficient (Fragmented):&lt;/STRONG&gt;
&lt;UL data-path-to-node="67,0,1"&gt;
&lt;LI&gt;&lt;EM data-path-to-node="67,0,1,0,0" data-index-in-node="0"&gt;Chat 1:&lt;/EM&gt; "What does this function do?"&lt;/LI&gt;
&lt;LI&gt;&lt;EM data-path-to-node="67,0,1,1,0" data-index-in-node="0"&gt;Chat 2:&lt;/EM&gt; "Can you write tests for it?"&lt;/LI&gt;
&lt;LI&gt;&lt;EM data-path-to-node="67,0,1,2,0" data-index-in-node="0"&gt;Chat 3:&lt;/EM&gt; "Can you check it for security issues?"&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;LI&gt;👉 &lt;STRONG data-path-to-node="67,1,0" data-index-in-node="3"&gt;More Efficient (Batched):&lt;/STRONG&gt;
&lt;UL data-path-to-node="67,1,1"&gt;
&lt;LI&gt;&lt;EM data-path-to-node="67,1,1,0,0" data-index-in-node="0"&gt;Single Chat:&lt;/EM&gt; "Review this function for correctness, security, performance, and test coverage. Give me the top issues first, then suggest the smallest safe improvement."&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-path-to-node="68"&gt;&lt;STRONG data-path-to-node="68" data-index-in-node="0"&gt;The Balance:&lt;/STRONG&gt; Batch highly related work in a single conversation, but hit the "New Chat" button the moment you pivot to an entirely new topic.&lt;/P&gt;
&lt;H2 data-path-to-node="70"&gt;8. Keep Custom Instructions Short&lt;/H2&gt;
&lt;P data-path-to-node="71"&gt;Custom instructions (.github/copilot-instructions.md) and AGENT.md files are powerful tools for enforcing team standards, but they are appended as base tokens to &lt;STRONG data-path-to-node="71" data-index-in-node="162"&gt;every single chat interaction&lt;/STRONG&gt;. If they are bloated, they act as a hidden tax on every prompt.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;✅ Keep it Concise &amp;amp; Rules-Based:&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;&lt;STRONG&gt;❌ Avoid Attaching to Every Prompt:&lt;/STRONG&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;SPAN data-path-to-node="72,1,0,0"&gt;&lt;EM data-path-to-node="72,1,0,0" data-index-in-node="0"&gt;"Prefer small, testable changes."&lt;/EM&gt;&lt;/SPAN&gt;&lt;/td&gt;&lt;td&gt;&lt;SPAN data-path-to-node="72,1,1,0"&gt;Massive architecture blueprints&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;SPAN data-path-to-node="72,2,0,0"&gt;&lt;EM data-path-to-node="72,2,0,0" data-index-in-node="0"&gt;"Do not modify public APIs without calling it out."&lt;/EM&gt;&lt;/SPAN&gt;&lt;/td&gt;&lt;td&gt;&lt;SPAN data-path-to-node="72,2,1,0"&gt;Full enterprise coding standards documents&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;SPAN data-path-to-node="72,3,0,0"&gt;&lt;EM data-path-to-node="72,3,0,0" data-index-in-node="0"&gt;"Use our standard logging pattern."&lt;/EM&gt;&lt;/SPAN&gt;&lt;/td&gt;&lt;td&gt;&lt;SPAN data-path-to-node="72,3,1,0"&gt;Entire infrastructure runbooks&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;SPAN data-path-to-node="72,4,0,0"&gt;&lt;EM data-path-to-node="72,4,0,0" data-index-in-node="0"&gt;"When writing tests, follow the existing test style."&lt;/EM&gt;&lt;/SPAN&gt;&lt;/td&gt;&lt;td&gt;&lt;SPAN data-path-to-node="72,4,1,0"&gt;Pages of multi-shot coding examples&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P data-path-to-node="73"&gt;&lt;STRONG data-path-to-node="73" data-index-in-node="0"&gt;The Strategy:&lt;/STRONG&gt; Use global/repository instructions &lt;EM data-path-to-node="73" data-index-in-node="49"&gt;only&lt;/EM&gt; for non-negotiable, high-level rules. For specific workflows, use localized prompt files, specialized skills, or custom agents.&lt;/P&gt;
&lt;H2 data-path-to-node="75"&gt;9. Use Skills and Specialized Agents for Repeatable Work&lt;/H2&gt;
&lt;P data-path-to-node="76"&gt;If your team frequently asks Copilot to perform identical workflows (e.g., API reviews, PR summarization, Accessibility audits, Terraform checks), turn them into reusable &lt;STRONG data-path-to-node="76" data-index-in-node="171"&gt;skills, prompt files, or custom agents&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P data-path-to-node="77"&gt;This ensures prompt formatting remains minimal, highly consistent, and optimized for low token usage. Furthermore, if you discover an excellent workflow during an active chat session, &lt;STRONG data-path-to-node="77" data-index-in-node="184"&gt;ask Copilot to turn that knowledge into a SKILL file.&lt;/STRONG&gt; This prevents future sessions from having to waste tokens "re-learning" the process.&lt;/P&gt;
&lt;H2 data-path-to-node="79"&gt;10. Avoid Unnecessary Agentic Mode&lt;/H2&gt;
&lt;P data-path-to-node="80"&gt;Agent mode is incredibly capable, but it shouldn't be the default default mode for standard queries.&lt;/P&gt;
&lt;UL data-path-to-node="81"&gt;
&lt;LI&gt;🛠️ &lt;STRONG data-path-to-node="81,0,0" data-index-in-node="4"&gt;Use Agent Mode For:&lt;/STRONG&gt; Multi-file bug fixes, generating and validating complex test suites, cross-file refactoring, reproducing test failures, or implementing full PR tasks.&lt;/LI&gt;
&lt;LI&gt;💬 &lt;STRONG data-path-to-node="81,1,0" data-index-in-node="3"&gt;Stick to Normal Chat/Completions For:&lt;/STRONG&gt; Simple syntax lookups, single-file explanations, writing minor snippets, basic documentation rewrites, or simple formatting tweaks.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-path-to-node="82,0"&gt;&lt;STRONG data-path-to-node="82,0" data-index-in-node="0"&gt;The Rule:&lt;/STRONG&gt; If a task doesn't explicitly require autonomous file editing, tool execution, or terminal commands, stick to standard chat or inline completions.&lt;/P&gt;
&lt;H2 data-path-to-node="84"&gt;11. Monitor Usage Weekly&lt;/H2&gt;
&lt;P data-path-to-node="85"&gt;Optimization without measurement is just guesswork. During your organization's transition into usage-based billing, establish a &lt;STRONG data-path-to-node="85" data-index-in-node="128"&gt;weekly review cadence&lt;/STRONG&gt; to look at your telemetry data.&lt;/P&gt;
&lt;H3 data-path-to-node="86"&gt;What to Look For:&lt;/H3&gt;
&lt;UL data-path-to-node="87"&gt;
&lt;LI&gt;Which models are drawing the most volume?&lt;/LI&gt;
&lt;LI&gt;Are premium frontier models being used for simple tasks?&lt;/LI&gt;
&lt;LI&gt;Which teams or workflows are spiking above their allocated credit budgets?&lt;/LI&gt;
&lt;LI&gt;Is tool usage or long context history driving up the average cost per session?&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-path-to-node="88"&gt;Use these insights to refine default model guidance, create user-level budget overrides for true power users, update prompt templates, and host short, continuous training sessions. &lt;STRONG data-path-to-node="88" data-index-in-node="181"&gt;The goal is never to shame high usage—high usage is fantastic if it yields high-value code. The goal is to eliminate systemic token waste.&lt;/STRONG&gt;&lt;/P&gt;
&lt;H2 data-path-to-node="90"&gt;12. Create a Simple Team Policy&lt;/H2&gt;
&lt;P data-path-to-node="91"&gt;To make this stick, give your developers a lightweight checklist they can actually memorize. Here is a great blueprint for an internal team policy:&lt;/P&gt;
&lt;UL data-path-to-node="92"&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="92,0,0" data-index-in-node="0"&gt;Completions First:&lt;/STRONG&gt; Use inline completions for normal code acceleration.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="92,1,0" data-index-in-node="0"&gt;Cheaper Models First:&lt;/STRONG&gt; Default to lightweight models; escalate to frontier models only for deep reasoning.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="92,2,0" data-index-in-node="0"&gt;Scope Wisely:&lt;/STRONG&gt; Mention specific files/folders; avoid blind repository scans.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="92,3,0" data-index-in-node="0"&gt;Plan First:&lt;/STRONG&gt; Ask for an architectural plan before allowing agents to execute large edits.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="92,4,0" data-index-in-node="0"&gt;Tool Hygiene:&lt;/STRONG&gt; Keep tools and custom instructions tightly scoped and concise.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="92,5,0" data-index-in-node="0"&gt;Review Cadence:&lt;/STRONG&gt; Check usage metrics weekly and adjust budgets monthly.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2 data-path-to-node="94"&gt;Reusable, Cost-Aware Prompt Templates&lt;/H2&gt;
&lt;P data-path-to-node="95"&gt;Copy and paste these templates into your daily workflows to keep token counts down:&lt;/P&gt;
&lt;UL data-path-to-node="96"&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="96,0,0" data-index-in-node="0"&gt;For Debugging:&lt;/STRONG&gt; &lt;EM data-path-to-node="96,0,0" data-index-in-node="15"&gt;"Analyze this error using only the files I mention. Do not scan the whole repository. First explain the likely root cause, then suggest the smallest fix."&lt;/EM&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="96,1,0" data-index-in-node="0"&gt;For Agent Mode:&lt;/STRONG&gt; &lt;EM data-path-to-node="96,1,0" data-index-in-node="16"&gt;"Use tools only if needed. Before editing, create a short plan and list the files you need to inspect. Run only the narrowest relevant test."&lt;/EM&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="96,2,0" data-index-in-node="0"&gt;For Code Review:&lt;/STRONG&gt; &lt;EM data-path-to-node="96,2,0" data-index-in-node="17"&gt;"Review this pull request for correctness, security, and maintainability. Focus on high-impact issues only. Do not rewrite the code unless necessary."&lt;/EM&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="96,3,0" data-index-in-node="0"&gt;For Testing:&lt;/STRONG&gt; &lt;EM data-path-to-node="96,3,0" data-index-in-node="13"&gt;"Generate unit tests for this function using the existing test style. Do not modify production code. Keep the test scope narrow."&lt;/EM&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="96,4,0" data-index-in-node="0"&gt;For Refactoring:&lt;/STRONG&gt; &lt;EM data-path-to-node="96,4,0" data-index-in-node="17"&gt;"Refactor this file only. Preserve behavior. Do not change public APIs. Explain the risk before making edits."&lt;/EM&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="96,5,0" data-index-in-node="0"&gt;For Large Repositories:&lt;/STRONG&gt; &lt;EM data-path-to-node="96,5,0" data-index-in-node="24"&gt;"Do not analyze the entire repository. Start with /src/payment and /tests/payment. Ask before expanding scope."&lt;/EM&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2 data-path-to-node="98"&gt;Summary: What to Avoid&lt;/H2&gt;
&lt;P data-path-to-node="99"&gt;To keep budgets predictable, coach your teams away from these common anti-patterns:&lt;/P&gt;
&lt;OL data-path-to-node="100"&gt;
&lt;LI&gt;Setting an expensive frontier model as the default for every single interaction.&lt;/LI&gt;
&lt;LI&gt;Running full agent loops for questions that a simple chat could answer.&lt;/LI&gt;
&lt;LI&gt;Allowing agents to freely scan entire codebases without folder constraints.&lt;/LI&gt;
&lt;LI&gt;Enabling every single available MCP tool by default.&lt;/LI&gt;
&lt;LI&gt;Dumping massive, raw terminal outputs or log dumps directly into the prompt.&lt;/LI&gt;
&lt;LI&gt;Carrying massive, multi-hour chat histories instead of opening fresh sessions.&lt;/LI&gt;
&lt;LI&gt;Disregarding your usage dashboards until the invoice arrives.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2 data-path-to-node="102"&gt;The Bigger Picture: Cost Optimization &lt;EM data-path-to-node="102" data-index-in-node="38"&gt;is&lt;/EM&gt; Workflow Optimization&lt;/H2&gt;
&lt;P data-path-to-node="103"&gt;Optimizing Copilot costs isn't about using AI less—&lt;STRONG data-path-to-node="103" data-index-in-node="51"&gt;it’s about using AI better.&lt;/STRONG&gt; Smaller context windows, tighter tool scoping, cleaner prompts, and proper model selection don't just lower the bill; &lt;STRONG data-path-to-node="103" data-index-in-node="197"&gt;they make the model significantly more accurate.&lt;/STRONG&gt; When you overwhelm a model with irrelevant files and messy logs, you introduce noise that degrades the quality of the output.&lt;/P&gt;
&lt;P data-path-to-node="104"&gt;Cost optimization is an engineering maturity practice, not a restriction. When managed through a smart operating model, GitHub Copilot will continue to securely accelerate our delivery, eliminate boilerplate toil, and maximize our focus on the work that truly matters.&lt;/P&gt;
&lt;P data-path-to-node="104"&gt;&lt;STRONG&gt;References:&lt;/STRONG&gt;&lt;/P&gt;
&lt;P data-path-to-node="104"&gt;&lt;A href="https://techcommunity.microsoft.com/blog/azuredevcommunityblog/token-economics-the-new-finops-for-agentic-ai/4533743" target="_blank" rel="noopener"&gt;Token Economics: The New FinOps for Agentic AI | Microsoft Community Hub&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Contributors:&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;This article is maintained by Microsoft. It was originally written by the following contributors.&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://www.linkedin.com/in/gaurav-bhardwaj-33312854/" target="_blank" rel="noopener"&gt;Gaurav Bhardwaj&lt;/A&gt; | Senior Cloud Solution Architect&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://www.linkedin.com/in/dustinellis/" target="_blank" rel="noopener"&gt;Dustin Ellis&lt;/A&gt; | Senior Cloud Solution Architect&amp;nbsp;&lt;/LI&gt;
&lt;/OL&gt;</description>
      <pubDate>Tue, 07 Jul 2026 15:43:49 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/optimizing-github-copilot-cost-in-the-usage-based-billing-era/ba-p/4534171</guid>
      <dc:creator>gauravbhardwaj</dc:creator>
      <dc:date>2026-07-07T15:43:49Z</dc:date>
    </item>
    <item>
      <title>Revolutionizing Document Intelligence: Scaling Construction Industries with AI-Driven Extraction</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/revolutionizing-document-intelligence-scaling-construction/ba-p/4522393</link>
      <description>&lt;H4&gt;&lt;STRONG&gt;Introduction&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P&gt;Generative AI (GenAI) is poised to transform the construction industry by addressing chronic challenges such as low productivity, cost overruns, schedule delays, and labor shortages. By automating the analysis of drawings, specifications, contracts, and project documentation, GenAI can reduce manual effort, accelerate decision-making, and improve coordination across architects, engineers, contractors, and suppliers. Industry studies indicate that AI-powered workflows can increase productivity by 20–40% in planning, engineering, and administrative functions while reducing costly rework and errors. The result is faster project delivery, improved resource utilization, lower costs, and more predictable project outcomes.&lt;/P&gt;
&lt;P&gt;A major opportunity for GenAI in construction lies in its ability to unlock the vast amount of information trapped within AutoCAD drawings, architectural plans, BIM models, specifications, and engineering documents. Today, project teams spend countless hours manually reviewing drawings, performing quantity takeoffs, identifying dependencies, and translating design intent into actionable work packages for downstream trades. GenAI can automate this process by extracting and interpreting dimensions, materials, quantities, assemblies, and building components directly from design artifacts, then intelligently distributing that information to foundation, framing, roofing, insulation, MEP, and finish teams. This creates a digital thread from design through execution, eliminating manual handoffs, reducing human error, and ensuring every stakeholder works from a single source of truth. The impact extends beyond productivity gains—GenAI enables more accurate material forecasting, streamlined procurement, reduced waste, faster response to design changes, fewer change orders, and greater confidence that the architect's vision is executed precisely in the field. In an industry where margins are tight and inefficiencies are costly, GenAI has the potential to fundamentally redefine how construction projects are planned, coordinated, and delivered.&lt;/P&gt;
&lt;P&gt;This article specifically demonstrates how organizations can leverage Azure AI services—including &lt;STRONG&gt;Azure Content Understanding, Azure foundry, Azure Blob Storage, Azure Open AI&lt;/STRONG&gt;—to extract, understand, and operationalize information from construction drawings and project documentation. The solution illustrates how Azure's AI platform can transform unstructured design artifacts into actionable intelligence that improves productivity, reduces risk, accelerates procurement, and enables more efficient execution across the entire construction lifecycle.&lt;/P&gt;
&lt;P data-path-to-node="6"&gt;This transformation is now achievable through a hybrid AI architecture. By combining structured layout understanding models with Generative AI &lt;STRONG&gt;reasoning capabilities&lt;/STRONG&gt;, organizations can build highly scalable, intelligent extraction systems that meet the rigorous safety and compliance standards of the construction sector.&lt;/P&gt;
&lt;H4&gt;&lt;STRONG&gt;The Evolution from GenAI Approach to Deterministic Precision&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P&gt;Starting with a Generative AI–driven approach to extract structured fields from documents is a fundamentally more effective initial strategy. It accelerates early-stage extraction without requiring large, labeled datasets, while simultaneously enabling structured data collection needed to train deterministic models—which typically require thousands of annotated samples.&lt;/P&gt;
&lt;P&gt;This approach delivers immediate value by rapidly identifying relevant data patterns in documents and uncovering key factors that influence extraction accuracy, such as document quality, layout complexity, and multi-section ambiguity. At the same time, it naturally builds the dataset necessary to transition toward a more scalable and repeatable solution.&lt;/P&gt;
&lt;P&gt;However, while powerful for contextual reasoning across document sections, Generative AI is inherently probabilistic and sensitive to input variability. For enterprise-grade reliability, precision, and repeatable structured document extraction, a complementary approach is required.&lt;/P&gt;
&lt;P&gt;The optimal solution is a hybrid model that combines the strengths of both:&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Azure Content Understanding&lt;/STRONG&gt;&amp;nbsp;provides precise, consistent field extraction with per-field confidence scores at scale.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Azure OpenAI GPT-5.2 &lt;/STRONG&gt;(generative) adds contextual reasoning, validates ambiguous fields, fills extraction gaps, and interprets complex multi-section relationships.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;AI Agent &lt;/STRONG&gt;(bounded triage) handles exception cases with structured CORRECT/ACCEPT/ESCALATE decisions before human escalation. Together, they form a superior system—delivering higher accuracy, reduced ambiguity, bounded AI cost, and stronger auditability in complex real-world conditions.&lt;/P&gt;
&lt;P&gt;&lt;EM&gt;Note : AI cannot compensate for inconsistent input data. Standardized document schemas and operational discipline remain prerequisites for reliable automation.&lt;/EM&gt;&lt;/P&gt;
&lt;H4&gt;&lt;STRONG&gt;Solution Components and Architecture&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P&gt;The solution follows a modular, event-driven architecture that combines deterministic document understanding and Generative AI to enable scalable, intelligent extraction workflows. At a high level, documents are ingested, deduplicated, processed through Azure Content Understanding for primary extraction, enhanced with GPT-5.2 for gap-fill verification, validated against business rules, and routed through a confidence-based decision system before persistence. The code repository for the solution can be found &lt;A class="lia-external-url" href="https://github.com/lifesawesome/Document_Extraction_Repo" target="_blank" rel="noopener"&gt;here&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Conceptual Architecture&lt;/STRONG&gt;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;STRONG&gt;Azure Architecture: -&lt;/STRONG&gt;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;The pipeline execution follows this flow: a document is uploaded to Azure Blob Storage, triggering the orchestrator. The pipeline checks for duplicates via SHA-256 hash against Cosmos DB. New documents are submitted to Azure Content Understanding, which returns structured fields with per-field confidence scores. The AI Schema Mapper then identifies gaps—fields that are missing or have confidence below 0.70—and sends only those to GPT-4.1 for verification. Results are normalized, validated against cross-field business rules, and routed based on aggregate confidence.&lt;/P&gt;
&lt;P&gt;Throughout the pipeline, built-in feedback loops—quality filtering, validation checks, and confidence gates—ensure that only high-confidence results are persisted automatically, enabling a reliable and production-ready extraction system.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure Blob Storage&lt;/STRONG&gt; — Primary storage for source PDFs and extraction artifacts. Standard_LRS, Hot tier, HTTPS-only with SAS-secured access for Content Understanding.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure Content Understanding&lt;/STRONG&gt; — Primary deterministic extractor with custom analyzer supporting 100+ configurable fields. Returns per-field confidence scores (0.0–1.0) plus raw markdown text. Non-LLM, repeatable, and auditable.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure AI Foundry / OpenAI (GPT-5.2) &lt;/STRONG&gt;— Bounded gap-fill verifier invoked only for missing or low-confidence fields (typically 10–20% of total). Temperature 0.0, JSON response format enforced, schema-aware prompting with domain rules.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure Cosmos DB (Serverless)&lt;/STRONG&gt;— Document persistence with SHA-256 deduplication, version increment on re-processing, and partition-by-document-type for efficient querying. Pay-per-request scales from zero.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure Service Bus (Basic)&lt;/STRONG&gt; — Event-driven queue integration with `document-processing` and `human-review` queues for processing triggers and escalation routing.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Application Insights + OpenTelemetry &lt;/STRONG&gt;— End-to-end observability with per-stage telemetry events, custom metrics (fill_rate, record_confidence, extraction_duration_ms), and distributed tracing&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Cost Impact of Hybrid Approach&lt;/STRONG&gt;&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 100%; height: 155.2px; border-width: 1px;"&gt;&lt;colgroup&gt;&lt;col style="width: 25%" /&gt;&lt;col style="width: 25%" /&gt;&lt;col style="width: 25%" /&gt;&lt;col style="width: 25%" /&gt;&lt;/colgroup&gt;&lt;tbody&gt;&lt;tr style="height: 38.8px;"&gt;&lt;td style="height: 38.8px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Metric&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 38.8px;"&gt;
&lt;P&gt;&lt;STRONG&gt;CU-Only&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 38.8px;"&gt;
&lt;P&gt;&lt;STRONG&gt;GPT-Only&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 38.8px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Hybrid (This Architecture)&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 38.8px;"&gt;&lt;td style="height: 38.8px;"&gt;
&lt;P&gt;Cost per document&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 38.8px;"&gt;
&lt;P&gt;&amp;nbsp;~$0.01&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 38.8px;"&gt;
&lt;P&gt;$0.15–0.30&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 38.8px;"&gt;
&lt;P&gt;$0.03–0.05&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 38.8px;"&gt;&lt;td style="height: 38.8px;"&gt;
&lt;P&gt;Determinism&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 38.8px;"&gt;100%&lt;/td&gt;&lt;td style="height: 38.8px;"&gt;Variable&lt;/td&gt;&lt;td style="height: 38.8px;"&gt;95%+&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 38.8px;"&gt;&lt;td style="height: 38.8px;"&gt;
&lt;P&gt;Accuracy&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 38.8px;"&gt;75-80%&lt;/td&gt;&lt;td style="height: 38.8px;"&gt;
&lt;P&gt;80–90%&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 38.8px;"&gt;90-95%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Auditability&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;Full&lt;/td&gt;&lt;td&gt;Limited&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Per-field source attribution&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;Cost savings: 60–80% reduction compared to GPT-only by limiting LLM to gap fields.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Security and Enterprise Considerations&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Azure Blob Storage&lt;/STRONG&gt;:&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Storage accounts can be secured by minimizing public exposure, enforcing strong identity‑based access, protecting data, and continuously monitoring for threats. Organizations should use &lt;A href="https://learn.microsoft.com/en-us/azure/private-link/private-endpoint-overview" target="_blank" rel="noopener"&gt;Private Endpoints&lt;/A&gt; and disable public network access wherever possible, authenticate users and applications with Microsoft Entra ID instead of shared keys, and apply least‑privilege Azure RBAC with managed identities. Data should be encrypted in transit (TLS 1.2+) and at rest using Microsoft‑managed or customer‑managed keys stored in &lt;A href="https://learn.microsoft.com/en-us/azure/key-vault/general/overview" target="_blank" rel="noopener"&gt;Azure Key Vault&lt;/A&gt;, while &lt;A href="https://learn.microsoft.com/en-us/azure/defender-for-cloud/defender-for-storage-introduction" target="_blank" rel="noopener"&gt;Microsoft Defender for Storage&lt;/A&gt;, logging, soft delete, backups, and Azure Policy should be enabled to &lt;A href="https://learn.microsoft.com/en-us/azure/defender-for-cloud/enable-defender-for-storage-data-sensitivity" target="_blank" rel="noopener"&gt;detect threats&lt;/A&gt;, support recovery, and enforce compliance at scale. &lt;A href="https://learn.microsoft.com/en-us/azure/ai-services/content-safety/quickstart-image?tabs=visual-studio%2Cwindows&amp;amp;pivots=programming-language-foundry-portal" target="_blank" rel="noopener"&gt;Content Safety&lt;/A&gt; can be called from the application layer to block uploads based on &lt;A href="https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/harm-categories?tabs=warning" target="_blank" rel="noopener"&gt;image content&lt;/A&gt;. Staging containers can be used to isolate untrusted uploads. Content Safety provides signals; your app enforces policy.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Azure Content Understanding / AI Vision&lt;/STRONG&gt;:&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-teams="true"&gt;Azure AI services support enterprise-grade security through Microsoft Entra ID–based authentication and &lt;A href="https://learn.microsoft.com/en-us/azure/role-based-access-control/overview" aria-label="Link Azure RBAC" target="_blank"&gt;Azure RBAC&lt;/A&gt;, ensuring only authorized applications can access extraction models. Network isolation can be enforced using &lt;A href="https://learn.microsoft.com/en-us/azure/ai-services/content-understanding/concepts/secure-communications" aria-label="Link Virtual Network (VNet)" target="_blank"&gt;Virtual Network (VNet)&lt;/A&gt; integration and &lt;A href="https://learn.microsoft.com/en-us/azure/private-link/private-link-overview" aria-label="Link Private Link" target="_blank"&gt;Private Link&lt;/A&gt; to restrict public internet exposure. All data transmitted is encrypted in transit and at rest. Microsoft Defender for Cloud provides continuous security posture visibility across these AI workloads.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Azure OpenAI&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-teams="true"&gt;&lt;A href="https://learn.microsoft.com/en-us/security/benchmark/azure/mcsb-v2-artificial-intelligence-security" aria-label="Link Govern" target="_blank"&gt;Govern&lt;/A&gt;&amp;nbsp;which models are approved for use and protect model artifacts and training data from unauthorized access through strong identity, network, encryption, and logging controls. AI applications should be designed with &lt;A href="https://learn.microsoft.com/en-us/security/benchmark/azure/baselines/azure-openai-security-baseline" aria-label="Link layered defenses" target="_blank"&gt;layered defenses&lt;/A&gt;, including multi‑stage content filtering, safety meta‑prompts, and least‑privilege permissions for agents and plugins to reduce the risk of prompt injection, data leakage, and unintended actions. High‑risk AI operations should include human‑in‑the‑loop review to prevent autonomous execution of harmful or incorrect outcomes. Organizations must continuously monitor AI systems for misuse, anomalous behavior, and data exfiltration, and they should perform ongoing &lt;A href="https://learn.microsoft.com/en-us/azure/foundry/concepts/ai-red-teaming-agent" aria-label="Link AI red teaming" target="_blank"&gt;AI red teaming&lt;/A&gt; to identify vulnerabilities such as jailbreaking, adversarial inputs, and model manipulation before they can be exploited.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Azure Cosmos DB&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Azure Cosmos enhances network security by supporting access restrictions via &lt;A href="https://www.google.com/url?sa=E&amp;amp;q=https%3A%2F%2Flearn.microsoft.com%2Fen-us%2Fazure%2Fcosmos-db%2Fhow-to-configure-vnet-service-endpoint" target="_blank" rel="noopener"&gt;Virtual Network (VNet) integration&lt;/A&gt;and secure access through &lt;A href="https://www.google.com/url?sa=E&amp;amp;q=https%3A%2F%2Flearn.microsoft.com%2Fen-us%2Fazure%2Fcosmos-db%2Fhow-to-configure-private-endpoints" target="_blank" rel="noopener"&gt;Private Link&lt;/A&gt;. Data protection is reinforced by integration with &lt;A href="https://www.google.com/url?sa=E&amp;amp;q=https%3A%2F%2Flearn.microsoft.com%2Fen-us%2Fazure%2Fpurview%2Foverview" target="_blank" rel="noopener"&gt;Microsoft Purview&lt;/A&gt;, which helps classify and label sensitive data, and &lt;A href="https://www.google.com/url?sa=E&amp;amp;q=https%3A%2F%2Ftechcommunity.microsoft.com%2FOverview%2520of%2520Defender%2520for%2520Azure%2520Cosmos%2520DB%2520-%2520Microsoft%2520Defender%2520for%2520Cloud%2520%7C%2520Microsoft%2520Learn" target="_blank" rel="noopener"&gt;Defender for Cosmos DB&lt;/A&gt;to detect threats and exfiltration attempts. Cosmos DB ensures all data is encrypted in transit using TLS 1.2+ (mandatory) and at rest using Microsoft-managed or &lt;A href="https://www.google.com/url?sa=E&amp;amp;q=https%3A%2F%2Flearn.microsoft.com%2Fen-us%2Fazure%2Fsecurity%2Ffundamentals%2Fencryption-atrest%23customer-managed-keys" target="_blank" rel="noopener"&gt;customer-managed keys (CMKs)&lt;/A&gt;.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Azure Functions / Compute&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-teams="true"&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/azure-functions/security-concepts" aria-label="Link Secured" target="_blank"&gt;Secured&lt;/A&gt; with Entra ID authentication and managed identities, least-privilege &lt;A href="https://learn.microsoft.com/en-us/azure/role-based-access-control/overview" aria-label="Link RBAC" target="_blank"&gt;RBAC&lt;/A&gt;, HTTPS-only access, private endpoints, VNet integration, and &lt;A href="https://learn.microsoft.com/en-us/azure/app-service/app-service-key-vault-references?tabs=azure-cli" aria-label="Link Key Vault" target="_blank"&gt;Key Vault&lt;/A&gt; for secrets. Hardened with Azure Policy, Defender for Cloud, and centralized logging.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Microsoft Foundry&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Microsoft&lt;STRONG&gt; &lt;/STRONG&gt;Foundry supports robust identity management using &lt;A href="https://learn.microsoft.com/en-us/azure/role-based-access-control/overview" target="_blank" rel="noopener"&gt;Azure Role-Based Access Control (RBAC)&lt;/A&gt;&lt;U&gt; &lt;/U&gt;to assign roles within&lt;U&gt; &lt;/U&gt;&lt;A href="https://learn.microsoft.com/en-us/entra/identity/" target="_blank" rel="noopener"&gt;Microsoft Entra ID&lt;/A&gt;, and it supports&lt;U&gt; &lt;/U&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/active-directory/managed-identities-azure-resources/overview" target="_blank" rel="noopener"&gt;Managed Identities&lt;/A&gt; for secure resource access.&lt;U&gt; &lt;/U&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/active-directory/conditional-access/overview" target="_blank" rel="noopener"&gt;Conditional Access&lt;/A&gt; policies allow organizations to enforce access based on location, device, and risk level. For network security, Azure AI Foundry supports &lt;A href="https://learn.microsoft.com/en-us/azure/private-link/private-link-overview" target="_blank" rel="noopener"&gt;Private Link&lt;/A&gt;, Managed Network Isolation, and &lt;A href="https://learn.microsoft.com/en-us/azure/virtual-network/network-security-groups-overview" target="_blank" rel="noopener"&gt;Network Security Groups (NSGs)&lt;/A&gt; to restrict resource access. Data is encrypted in transit and at rest using Microsoft-managed keys or optional &lt;A href="https://learn.microsoft.com/en-us/azure/security/fundamentals/encryption-atrest#customer-managed-keys" target="_blank" rel="noopener"&gt;Customer-Managed Keys (CMKs)&lt;/A&gt;&lt;U&gt;.&lt;/U&gt; &lt;A href="https://learn.microsoft.com/en-us/azure/governance/policy/overview" target="_blank" rel="noopener"&gt;Azure Policy&lt;/A&gt;&lt;U&gt; &lt;/U&gt;enables auditing and enforcing configurations for all resources deployed in the environment. Additionally, &lt;A href="https://learn.microsoft.com/en-us/entra/agent-id/identity-professional/microsoft-entra-agent-identities-for-ai-agents" target="_blank" rel="noopener"&gt;Microsoft Entra Agent ID&lt;/A&gt;, which extends identity management and access capabilities to AI agents. AI agents created within Microsoft Foundry are automatically assigned identities in a Microsoft Entra directory centralizing agent and user management in one solution. &lt;A href="https://learn.microsoft.com/en-us/azure/defender-for-cloud/ai-security-posture" target="_blank" rel="noopener"&gt;AI Security Posture Management&lt;/A&gt; can be used to assess the security posture of AI workloads. &lt;A href="https://learn.microsoft.com/en-us/azure/defender-for-cloud/ai-onboarding" target="_blank" rel="noopener"&gt;Defender for AI Services&lt;/A&gt; provides threat protection and insights for you AI resources. &lt;A href="https://learn.microsoft.com/en-us/purview/developer/secure-ai-with-purview" target="_blank" rel="noopener"&gt;Purview APIs&lt;/A&gt; enable Azure AI Foundry and developers to integrate data security and compliance controls into custom AI apps and agents. This includes enforcing policies based on how users interact with sensitive information in AI applications. &lt;A href="https://learn.microsoft.com/en-us/purview/ai-microsoft-purview" target="_blank" rel="noopener"&gt;Purview&lt;/A&gt; Sensitive Information Types can be used to detect sensitive data in user prompts and responses when interacting with AI applications.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;DevOps Security&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Security is further “shifted left” by integrating automated controls directly into CI/CD pipelines.&lt;STRONG&gt; &lt;/STRONG&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/devops/repos/security/configure-github-advanced-security-features?view=azure-devops&amp;amp;tabs=yaml&amp;amp;pivots=standalone-ghazdo" target="_blank" rel="noopener"&gt;GitHub Advanced Security for Azure DevOps&lt;/A&gt;, which provides dependency scanning, &lt;A href="https://codeql.github.com/" target="_blank" rel="noopener"&gt;CodeQL&lt;/A&gt;-based static application security testing (SAST), and secret scanning&lt;STRONG&gt; &lt;/STRONG&gt;to identify vulnerabilities and exposed credentials in code and third-party libraries. Infrastructure-as-code templates can be validated with &lt;A href="https://learn.microsoft.com/azure/governance/policy/overview" target="_blank" rel="noopener"&gt;Azure Policy&lt;/A&gt; and &lt;A href="https://learn.microsoft.com/azure/defender-for-cloud/" target="_blank" rel="noopener"&gt;Microsoft Defender for Cloud&lt;/A&gt;, while pipeline protections such as protected branches and approvals reduce the risk of unauthorized changes. DevOps environments can be hardened using &lt;A href="https://learn.microsoft.com/azure/key-vault/general/overview" target="_blank" rel="noopener"&gt;Azure Key Vault&lt;/A&gt; for secrets management, &lt;A href="https://learn.microsoft.com/azure/active-directory/managed-identities-azure-resources/overview" target="_blank" rel="noopener"&gt;Managed Identities&lt;/A&gt; and &lt;A href="https://learn.microsoft.com/entra/fundamentals/whatis" target="_blank" rel="noopener"&gt;Microsoft Entra ID&lt;/A&gt; for least-privilege access, and monitoring through &lt;A href="https://learn.microsoft.com/azure/azure-monitor/overview" target="_blank" rel="noopener"&gt;Azure Monitor&lt;/A&gt; . &lt;A href="https://learn.microsoft.com/en-us/azure/defender-for-cloud/defender-for-devops-introduction" target="_blank" rel="noopener"&gt;Microsoft Defender for Cloud DevOps Security&lt;/A&gt; provides centralized code‑to‑cloud visibility across Azure DevOps, GitHub, and GitLab, identifying risks in code, secrets, dependencies, and IaC and helping teams prioritize fixes early in CI/CD pipelines&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Related and Future Scenarios&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Although document extraction serves as the initial use case, this architecture establishes a scalable pattern for many applications:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Insurance Claims Processing&lt;/STRONG&gt;: Swap schema to claim fields; update CU analyzer for claim forms&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Legal Contract Analysis&lt;/STRONG&gt;: Schema for clauses, parties, dates; add NER in normalization&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Healthcare Medical Records&lt;/STRONG&gt;: HIPAA-compliant Cosmos; schema for diagnoses, medications, vitals&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Financial Document Processing&lt;/STRONG&gt;: Schema for transactions, accounts; add currency normalization&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Engineering/Construction Plans&lt;/STRONG&gt;: Schema for dimensions, materials, specifications&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Digital Twin Integration&lt;/STRONG&gt;: Feed extracted data into asset models for real-time facility visualization&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Predictive Analytics&lt;/STRONG&gt;: Track extracted values over time for trend detection and forecasting&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Conclusion&lt;/STRONG&gt;&lt;/P&gt;
&lt;P data-path-to-node="49"&gt;Modernizing document extraction is not simply about applying AI—it requires aligning technology, operational discipline, and data quality. Early exploration using Generative AI enabled rapid learning and feasibility validation. However, a production-grade solution must be built on structured layout understanding models supported by standardized schema definitions and operational controls.&lt;/P&gt;
&lt;P data-path-to-node="50"&gt;By combining primary structured extraction with Generative AI reasoning for bounded gap-fill verification, organizations can achieve scalable, repeatable, and auditable extraction processes. This hybrid approach enables reduced manual effort, lower error rates, and the transition from batch manual processing to intelligent, automated workflows.&lt;/P&gt;
&lt;P data-path-to-node="51"&gt;The result is not just an automated extraction tool, but a scalable AI architecture for modern document intelligence—adaptable to any industry, any document type, and any structured data need.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Contributors:&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;This article is maintained by Microsoft. It was originally written by the following contributors.&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://www.linkedin.com/in/gaurav-bhardwaj-33312854/" target="_blank" rel="noopener"&gt;Gaurav Bhardwaj&lt;/A&gt; | Senior Cloud Solution Architect – US Customer Success &lt;/LI&gt;
&lt;LI&gt;&lt;A style="font-style: normal; font-weight: 400; background-color: rgb(255, 255, 255);" href="https://www.linkedin.com/in/trmanasa" target="_blank" rel="noopener"&gt;Manasa Ramalinga&lt;/A&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt; | Senior Principal Cloud Solution Architect – US Customer Success &lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A style="font-style: normal; font-weight: 400; background-color: rgb(255, 255, 255);" href="https://www.linkedin.com/in/abed-sau/" target="_blank" rel="noopener"&gt;Abed Sau&lt;/A&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt; | Principal Cloud Solution Architect – US Customer Success&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/OL&gt;</description>
      <pubDate>Thu, 18 Jun 2026 20:52:34 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/revolutionizing-document-intelligence-scaling-construction/ba-p/4522393</guid>
      <dc:creator>gauravbhardwaj</dc:creator>
      <dc:date>2026-06-18T20:52:34Z</dc:date>
    </item>
    <item>
      <title>Announcing the Path to Production for Agents Webinar Series</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/announcing-the-path-to-production-for-agents-webinar-series/ba-p/4526560</link>
      <description>&lt;P&gt;Many organizations have made significant progress exploring AI—building pilots, prototypes, and proofs of concept. Yet a common challenge remains: how do you move from promising experiments to production-ready systems that are secure, scalable, and trusted? Join us for the Path to Production Webinar Series on July 27-28, a two-day deep dive designed to help technical teams operationalize AI and agent-based solutions using proven architecture patterns, governance models, and engineering practices.&lt;/P&gt;
&lt;P&gt;This simulive event will be available in two time zones, making it easier to participate.&lt;/P&gt;
&lt;P&gt;Use this link to register: &lt;A class="lia-external-url" href="https://aka.ms/AccelerateThePathToProduction" target="_blank" rel="noopener"&gt;https://aka.ms/AccelerateThePathToProduction&lt;/A&gt;&lt;/P&gt;
&lt;H1&gt;Why this series matters&lt;/H1&gt;
&lt;P&gt;A large percentage of AI initiatives never make it to production—not because of lack of ambition, but because organizations struggle to establish trustworthy, governed AI systems; build scalable architectural foundations; manage risk, cost, and operational complexity; and ensure reliability in non-deterministic systems. This webinar series addresses those challenges head-on with concrete, actionable guidance spanning the full lifecycle of production AI systems.&lt;/P&gt;
&lt;H1&gt;What attendees will learn&lt;/H1&gt;
&lt;P&gt;This series delivers an implementation-focused roadmap for building, deploying, and operating AI agents at enterprise scale. Each session dives deep into governance, architecture patterns, orchestration, security, evaluation, and observability - with reference architectures and real-world engineering examples. Learn how to design scalable agent systems, integrate with enterprise data and services, and apply best practices for reliability and performance.&amp;nbsp;&lt;BR /&gt;&lt;BR /&gt;Attendees will leave with practical techniques and proven patterns to confidently ship production-grade agent solutions. After the workshop, customers who have Unified Contracts are eligible for a packaged set of engagements that will implement this guidance with your Microsoft cloud solution architects. Otherwise, contact your partner to learn more about taking advantage of the Frontier Transformation Offer through the Frontier Accelerate program.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H1&gt;Session overview&lt;/H1&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Day&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Session&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Focus&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Speaker&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;July 27&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Frontier Center of Excellence (CoE) &amp;amp; Governance&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Create a governance framework with quality gates that helps organizations deliver secure, responsible, trustworthy AI at scale.&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Akriti Mehta&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;Divye Sheth&lt;/P&gt;
&lt;img /&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;July 27&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;AI Landing Zones&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Build a production-ready reference architecture for AI applications and agents with guardrails for networking, identity, security, and cost governance.&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Nadeem Ishqair&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;Bilal Amjad&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;July 27&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Agentic Architecture&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Adopt a governance-first, multi-agent architecture blueprint that embeds controls from user channels and orchestration through integration layers, data, and models.&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Yeliz Kilinc&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;Nour Shaker&lt;/P&gt;
&lt;img /&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;July 28&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;AgentOps&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Apply DevOps principles to production AI, including evaluation, CI/CD quality gates, observability, monitoring, red teaming, and incident response.&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Paulo Lacerda&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;Richard Healy&lt;/P&gt;
&lt;img /&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;July 28&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;AI Security, Trust &amp;amp; Observability&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Address prompt injection, data leakage, autonomous tool misuse, and AI-specific observability requirements for traceability, safety, and auditability.&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Yuening Chen&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;Raaid Mahbub&lt;/P&gt;
&lt;img /&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;July 28&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Solution Optimization&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Reduce token cost, cut latency, tune RAG, optimize multi-agent coordination, and apply FinOps practices for sustainable scale.&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Tanuja Bhamidipati&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;Fatos Ismali&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H1&gt;Day 1: Establishing the foundation for production AI&lt;/H1&gt;
&lt;P&gt;&lt;STRONG&gt;July 27&lt;/STRONG&gt;&lt;/P&gt;
&lt;H2&gt;AI Center of Excellence (CoE) &amp;amp; Governance&lt;/H2&gt;
&lt;P&gt;Why do so many AI initiatives die in the PoC graveyard? Because organizations cannot trust the AI. This session shows how an AI CoE plus governance framework creates a uniform quality gate at every layer of your AI application in an organization for delivering a single, organization-wide view of secure, responsible, trustworthy AI that is ready to scale.&lt;/P&gt;
&lt;H2&gt;AI Landing Zones&lt;/H2&gt;
&lt;P&gt;Scaling AI from experimentation to production demands a secure, governed, and scalable foundation. This session explores how AI Landing Zones provide a production-ready reference architecture for deploying AI applications and agents with the right guardrails for networking, identity, security, and cost governance, aligned with the Cloud Adoption Framework and Well-Architected best practices. Attendees will learn how to design AI platforms that balance innovation with compliance, accelerate time-to-production using validated architectures and infrastructure-as-code, and integrate AI services into enterprise environments.&lt;/P&gt;
&lt;H2&gt;Agentic Architecture&lt;/H2&gt;
&lt;P&gt;Many enterprise AI pilots stall not for lack of technology, but because they lack a trustworthy architecture. This session introduces a governance-first, multi-agent architecture blueprint that closes the trust gap by embedding uniform controls and quality checks at every level, from user channels and agent orchestration through integration layers to core data and models, under a common governance and security framework. Attendees will learn how this layered agentic architecture creates a reliable, enterprise-wide AI fabric that organizations can adopt with confidence, aligning AI initiatives with high standards of trust, interoperability, and scale.&lt;/P&gt;
&lt;H1&gt;Day 2: Operating and scaling AI in production&lt;/H1&gt;
&lt;P&gt;&lt;STRONG&gt;July 28&lt;/STRONG&gt;&lt;/P&gt;
&lt;H2&gt;AgentOps&lt;/H2&gt;
&lt;P&gt;This session covers the full lifecycle of deploying and operating agentic AI solutions in production. We will explore how teams can move from successful prototypes to production-ready agents using evaluation, CI/CD quality gates, observability, continuous monitoring, scheduled red teaming, and incident response practices. We will also cover how to apply DevOps principles to the unique challenges of AI systems, including non-deterministic behavior, prompt regression, model drift, tool-calling risk, and changing user behavior. Attendees will learn a practical AgentOps operating model for improving release confidence, detecting regressions earlier, and connecting agent operations back to Microsoft Foundry and Azure Monitor.&lt;/P&gt;
&lt;H2&gt;AI Security, Trust &amp;amp; Observability&lt;/H2&gt;
&lt;P&gt;This session focuses on securing AI systems in production, addressing risks beyond traditional application security such as prompt injection, data leakage, and autonomous tool misuse. It applies a defense-in-depth approach across identity, data protection, orchestration, and runtime controls. It also introduces AI-specific observability for trust and compliance, including traceability, safety and security monitoring, and auditability, ensuring AI systems are secure, controllable, and compliant at scale.&lt;/P&gt;
&lt;H2&gt;Solution Optimization&lt;/H2&gt;
&lt;P&gt;Getting AI to production is only half the battle. Once agentic workloads are live, organizations face compounding challenges including rising token costs, latency that degrades user trust, RAG pipelines that return noise instead of signal, and orchestration overhead that multiplies with every agent added to the mesh. This session provides a practical engineering playbook for optimizing agentic AI across the full stack, from model selection and inference routing through prompt compression, RAG tuning, caching strategies, and multi-agent coordination. It also covers the FinOps discipline required to control cost at scale, including capacity sizing, batch processing, and intelligent model routing. Attendees will leave with actionable patterns for reducing inference cost, cutting latency, and scaling reliably across regions.&lt;/P&gt;
&lt;H1&gt;Who should attend&lt;/H1&gt;
&lt;UL&gt;
&lt;LI&gt;Cloud and solution architects&lt;/LI&gt;
&lt;LI&gt;AI and ML engineers and developers&lt;/LI&gt;
&lt;LI&gt;Platform engineering and infrastructure teams&lt;/LI&gt;
&lt;LI&gt;Technical decision-makers driving AI transformation initiatives&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;If your team is working to move AI beyond prototypes into production-scale systems, this series will provide directly applicable guidance for architecture, governance, operations, and optimization.&lt;/P&gt;
&lt;H1&gt;Next steps after the webinar series:&lt;/H1&gt;
&lt;P&gt;We will conduct a personalized assessment of your organization’s readiness to adopt AI agents at scale.&lt;/P&gt;
&lt;H1&gt;Call to action&lt;/H1&gt;
&lt;P&gt;Join us on July 27-28 to accelerate your path from AI experimentation to trusted, enterprise-scale production systems. Registration details can be added to this announcement before publication.&lt;/P&gt;
&lt;P&gt;Use this link to register: &lt;A class="lia-external-url" href="https://aka.ms/AccelerateThePathToProduction" target="_blank" rel="noopener"&gt;Path to Production for Agents&lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 07 Jul 2026 14:54:02 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/announcing-the-path-to-production-for-agents-webinar-series/ba-p/4526560</guid>
      <dc:creator>brauerblogs</dc:creator>
      <dc:date>2026-07-07T14:54:02Z</dc:date>
    </item>
    <item>
      <title>Cloud Native Platforms: Build</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/cloud-native-platforms-build/ba-p/4519605</link>
      <description>&lt;P&gt;&lt;STRONG&gt;Audience:&lt;/STRONG&gt; Cloud architects, platform engineers, engineering leaders making design decisions&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Reading time:&lt;/STRONG&gt; 8 minutes&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Series:&lt;/STRONG&gt; Cloud Native Platforms. Build, Run, Evolve. This is Part 1 of 3.&lt;/P&gt;
&lt;HR /&gt;
&lt;P&gt;Most engineering teams can build systems.&lt;/P&gt;
&lt;P&gt;Few can scale them without rebuilding them.&lt;/P&gt;
&lt;P&gt;As platforms grow, complexity does not increase linearly. It multiplies across users, services, tenants, regions, and integrations. The systems that struggle and the systems that scale are rarely separated by which cloud they run on. They are separated by a handful of design choices made early and applied consistently.&lt;/P&gt;
&lt;P&gt;This post is about those choices.&lt;/P&gt;
&lt;H2&gt;The differentiator is not the cloud&lt;/H2&gt;
&lt;P&gt;Scalable platforms are not built with the right tools. They are built with the right design choices.&lt;/P&gt;
&lt;P&gt;Cloud services have closed the gap on infrastructure. The differentiator is no longer which managed service a team picks. It is whether the platform is designed to absorb change, tolerate failure, and support visibility from day one. Five engineering disciplines determine whether a platform scales gracefully or collects technical debt while it grows.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;EM&gt;Figure 1. The five disciplines compound into platform scale. Any one neglected becomes the constraint that forces a rewrite later.&lt;/EM&gt;&lt;/P&gt;
&lt;H2&gt;1. Flexibility is the foundation of scale&lt;/H2&gt;
&lt;P&gt;Hard-coded systems work until they do not. The first request to add a tenant, a region, a SKU (a sellable product variant), or a regulatory variant is the moment a rigid design starts to bend. Each subsequent request adds weight.&lt;/P&gt;
&lt;P&gt;Scalable platforms move behavior out of code:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Configuration replaces conditional logic&lt;/LI&gt;
&lt;LI&gt;Feature flags enable safer, tenant-scoped rollouts&lt;/LI&gt;
&lt;LI&gt;APIs evolve through versioning, not breaking changes&lt;/LI&gt;
&lt;LI&gt;Schemas evolve additively. Breaking changes go through versioned contracts with a deprecation window long enough that consumers can migrate without downtime.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: configuration in a managed store, feature flags with tenant scope, and APIs versioned per consumer contract. Cost is the discipline of treating configuration as code (versioned, reviewed, audited). The return is that releases stop being events and start being routine. A change that previously needed a coordinated deployment can be executed in minutes, gated to a single tenant for verification, and rolled out broadly only after the signal is clean. Most platforms reach this state by retrofit, not by design. Doing it earlier costs less than waiting.&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;EM&gt;If a change requires a redeploy, it should require a very good reason.&lt;/EM&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;2. Failures are normal. Resilience is a choice.&lt;/H2&gt;
&lt;P&gt;Distributed systems will fail in unpredictable ways. The real question is not how to prevent failure. It is how the system responds when failure happens.&lt;/P&gt;
&lt;P&gt;Resilience is engineered, not inherited from the platform. The patterns that move the needle are well known and consistently applied:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Idempotent operations&lt;/STRONG&gt; (safe to call multiple times with the same result) that make retries safe&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Reliable messaging patterns such as the transaction outbox&lt;/STRONG&gt; (writing the message to the same database transaction as the business change, then publishing asynchronously) to avoid lost or duplicated events&lt;/LI&gt;
&lt;LI&gt;Decoupled services that contain &lt;STRONG&gt;blast radius&lt;/STRONG&gt; (the scope of damage when one component fails)&lt;/LI&gt;
&lt;LI&gt;Timeouts, retries, and &lt;STRONG&gt;circuit breakers&lt;/STRONG&gt; (a wrapper around a dependency that stops calling it for a cool-off window after repeated failures) tuned per dependency&lt;/LI&gt;
&lt;LI&gt;Bulkheads (isolation pools, often a separate compute or queue lane per workload class) that keep noisy neighbours from starving critical paths of resources&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: every write that can be retried carries an idempotency key, every queue consumer is safe to replay, every event published goes through an outbox in the same transactional unit as the business change. When peak load triggers retries, duplicates collapse cleanly instead of producing duplicate orders, double-charged customers, or split-brain state. The contract changes outwards: callers can retry without thinking, queues can be at-least-once instead of exactly-once, and recovery moves from a manual cleanup task to a property of the system. Most teams that adopt this pattern stop seeing certain classes of incident entirely.&lt;/P&gt;
&lt;H3&gt;Implementation note&lt;/H3&gt;
&lt;P&gt;An idempotent API is not just a design preference. It changes how the rest of the system can be built. Once writes are safe to repeat, retries become cheap, queues become trustworthy, and recovery becomes automatic.&lt;/P&gt;
&lt;P&gt;The naive implementation (read the key, if absent process and save) has a race. Two concurrent requests with the same key both miss the lookup, both call the processor, and both attempt to save. That is the failure mode idempotency exists to prevent. The pattern that survives production is an atomic reserve-then-execute: insert a row keyed by the idempotency key with a unique constraint before doing any work. The first writer wins. Concurrent callers either wait for the original to complete and read its result, or they receive a conflict response.&lt;/P&gt;
&lt;LI-CODE lang="csharp"&gt;// Contract for the idempotency store. The two key methods are TryReserveAsync
// (atomic insert with unique-key constraint) and CompleteAsync (record the
// result of the first writer). GetCompletedResultAsync polls until the first
// writer commits or returns 409 Conflict if the in-flight window exceeds the
// configured deadline.
public interface IIdempotencyStore
{
    Task&amp;lt;Reservation&amp;gt; TryReserveAsync(
        string idempotencyKey, string requestHash, CancellationToken ct);

    Task CompleteAsync(
        string idempotencyKey, OrderResult result, CancellationToken ct);

    Task&amp;lt;OrderResult&amp;gt; GetCompletedResultAsync(
        string idempotencyKey, CancellationToken ct,
        TimeSpan? maxWait = null);
}

public readonly record struct Reservation(
    bool IsFirstWriter, string RequestHash);

// Idempotency via atomic reserve-then-execute.
// First writer wins; replays return the original result; concurrent
// duplicates lose the race and read the winner's outcome (or get 409).
public async Task&amp;lt;OrderResult&amp;gt; CreateOrderAsync(
    Order order, string idempotencyKey, CancellationToken ct)
{
    var requestHash = StableHash(order); // canonical content hash

    // Atomic insert: succeeds for the first caller, fails for the rest.
    var reserved = await _store.TryReserveAsync(
        idempotencyKey, requestHash, ct);

    if (!reserved.IsFirstWriter)
    {
        if (reserved.RequestHash != requestHash)
            throw new IdempotencyKeyReusedException();

        // A previous run committed (return its result) or is in-flight
        // (poll with a bounded deadline; 409 if exceeded).
        return await _store.GetCompletedResultAsync(
            idempotencyKey, ct, maxWait: TimeSpan.FromSeconds(5));
    }

    // We are the first writer. Execute, persist, mark complete.
    var result = await _processor.ProcessAsync(order, ct);
    await _store.CompleteAsync(idempotencyKey, result, ct);
    return result;
}

&lt;/LI-CODE&gt;
&lt;P&gt;Three production details matter:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;TTL or compaction on the idempotency record.&lt;/STRONG&gt; Without it, the store grows forever. Most teams retain records for the request retry window plus a safety margin (commonly 24 to 72 hours).&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Stable content hash, not the default object hash code.&lt;/STRONG&gt; The request hash detects key reuse with a different body, so a client that reuses an idempotency key with a different payload receives &lt;CODE&gt;IdempotencyKeyReusedException&lt;/CODE&gt; rather than silently getting the wrong result. Canonicalise field ordering, locale, and null handling before hashing.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Bound the in-flight window explicitly.&lt;/STRONG&gt; The genuinely hard case is when the processor succeeded but the store write failed. Production-grade implementations either run the side-effect and the store write in the same transaction (when the processor and store share a database) or use the transaction outbox pattern to bridge them. The poll-with-deadline in &lt;CODE&gt;GetCompletedResultAsync&lt;/CODE&gt; handles the duplicate-arrives-mid-flight case; the transactional boundary handles everything else.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;3. Observability is not optional&lt;/H2&gt;
&lt;P&gt;Without observability, teams operate blind. As systems grow, the price of guessing rises faster than the price of seeing.&lt;/P&gt;
&lt;P&gt;At build time, observability is a design property. The decisions made before the system reaches production are what determine whether it can be operated at all. The dashboards, alerts, and incident practices covered in Part 2 of this series rely on instrumentation choices made here.&lt;/P&gt;
&lt;P&gt;The build-time work that pays off in production:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Request identifiers propagated through every service hop, every queue, every async boundary, so a single user action can be traced end to end&lt;/LI&gt;
&lt;LI&gt;Structured logging with a consistent schema (event name, correlation id, tenant, severity) rather than free-form strings&lt;/LI&gt;
&lt;LI&gt;Metrics emitted at the boundaries that matter (every external call, every queue read or write, every database operation), not only at the entry point&lt;/LI&gt;
&lt;LI&gt;Tracing libraries integrated at the framework or middleware layer so coverage is automatic, not opt-in&lt;/LI&gt;
&lt;LI&gt;Schemas designed so business signals (orders, sessions, transactions) and system signals (CPU, latency, errors) share the same identifiers and can be correlated later&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: a single request id flowing through every service hop, every queue, every async boundary, propagated automatically at the framework layer rather than per-call. Add one structured logging schema across services (event name, correlation id, tenant, severity), so that a single query joins business events with system events. The investment is hours of upfront framework wiring. The return is that production diagnosis stops being archaeology. Cross-service questions become single dashboards; postmortems shrink from days to hours; and the dashboards in Part 2 actually work because the data underneath is shaped to support them.&lt;/P&gt;
&lt;H2&gt;4. Delivery practices set the ceiling&lt;/H2&gt;
&lt;P&gt;Scaling teams requires scaling delivery. Small inefficiencies in pipelines, environments, and release coordination compound into measurable drag.&lt;/P&gt;
&lt;P&gt;Delivery maturity that pays off at scale:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Pipelines as code, reviewed and versioned like application code&lt;/LI&gt;
&lt;LI&gt;Parallel deployments across services and regions where dependencies allow&lt;/LI&gt;
&lt;LI&gt;Infrastructure as code with shared modules, not hand-managed environments&lt;/LI&gt;
&lt;LI&gt;Automated quality gates: tests, security scans, dependency checks&lt;/LI&gt;
&lt;LI&gt;Trunk-based development (developers commit to a single shared branch many times a day) with short-lived feature branches and progressive delivery. &lt;STRONG&gt;Important caveat:&lt;/STRONG&gt; trunk-based works only when test automation and feature flags are already in place. Adopting it before those foundations exist tends to amplify production incidents rather than reduce them.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: pipelines run in parallel where dependencies allow, infrastructure provisioning is templated rather than per-environment, and quality gates run automatically rather than as discretionary steps. Sequential deployment of a multi-service platform across three environments takes hours; parallelised deployment of the same change takes minutes. The payback is not only release speed. It is the compounding cost reduction of every wait state for every engineer on every release. Teams that treat pipelines as a product feature, not an afterthought, ship more confidently and recover from bad changes faster because the rollback path was exercised, not invented during an incident.&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;EM&gt;Slow pipelines are not a tooling problem. They are a design problem.&lt;/EM&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;5. Cost discipline is engineering work&lt;/H2&gt;
&lt;P&gt;Cloud platforms can become expensive quickly when cost is treated as someone else's problem. Cost is a property of the design, not a quarterly review.&lt;/P&gt;
&lt;P&gt;The teams that get this right treat cost the same way they treat performance:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Elastic compute and storage tiers chosen per workload pattern&lt;/LI&gt;
&lt;LI&gt;Non-production environments with automated scale-down windows (the easiest savings to leave on the table)&lt;/LI&gt;
&lt;LI&gt;Tagging discipline so cost can be attributed to a service, a feature, a tenant&lt;/LI&gt;
&lt;LI&gt;Egress and data-tier choices, not compute, dominate cloud bills past a certain scale. Right-size storage tiers (hot vs cool vs archive), eliminate cross-region chatter, and watch egress on the data plane more closely than compute on the request path.&lt;/LI&gt;
&lt;LI&gt;Budgets and usage alerts wired into the same channels as reliability alerts&lt;/LI&gt;
&lt;LI&gt;Cost reviews built into design discussions, not deferred to FinOps (Financial Operations: the practice of managing cloud spend as an engineering concern)&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: non-production environments scale down automatically outside business hours, storage tiers match access patterns (hot, cool, archive), and tagging is enforced so every dollar can be attributed to a service or feature. Cost reviews happen at design time, not after the bill arrives. The biggest savings come from data plane decisions, not compute: cross-region egress, oversized storage tiers, and forgotten test environments dominate cloud bills past a certain scale. Treat cost as a first-class non-functional requirement, alongside latency and availability, and the discipline compounds in every design discussion that follows.&lt;/P&gt;
&lt;H2&gt;A scenario that ties it together&lt;/H2&gt;
&lt;img /&gt;
&lt;P&gt;&lt;EM&gt;Figure 2. A reference architecture that puts the disciplines into one shape. The request path is decoupled, the data layer is purpose-fit, identity is brokered by managed identity throughout, private endpoints isolate the data tier from public networks, and observability runs as a first-class lane.&lt;/EM&gt;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Picture a multi-tenant platform at a growth inflection. Onboarding a new tenant takes weeks because tenant-specific behaviour is hard-coded across services. Every release carries risk because there is no way to roll out a change to one tenant without affecting the rest. Incidents linger because logs and metrics live in different tools and nobody can correlate them in production.&lt;/P&gt;
&lt;P&gt;Do not start with a rewrite. Start with the smallest set of changes that unlocks the next year of growth: extract configuration out of code, introduce tenant-aware feature flags, wire a unified observability view into the existing services, and parallelise the pipelines. None of these are architectural revolutions. They are design choices applied with discipline, in the order the disciplines compound.&lt;/P&gt;
&lt;P&gt;Eighteen months in, onboarding a tenant takes hours instead of weeks. Releases move from monthly events to weekly increments. Incidents are caught earlier and resolved faster. The platform did not get bigger. It got more capable. The five disciplines did the work; the team made the choice to apply them.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;What teams get wrong&lt;/H2&gt;
&lt;P&gt;The common pattern is &lt;STRONG&gt;architecting for the system you have, not the system you are growing into.&lt;/STRONG&gt; It looks like progress because the current sprint ships. Pillars get postponed because they feel like overhead.&lt;/P&gt;
&lt;P&gt;The cost surfaces later. Each shortcut becomes a constraint. The constraints compound, and three releases later the team is debating a rewrite.&lt;/P&gt;
&lt;P&gt;The fix is not premature abstraction. It is small, deliberate investments in flexibility, resilience, observability, delivery, and cost from day one. The discipline is to make these investments before they are urgent.&lt;/P&gt;
&lt;H2&gt;Where to start when you cannot do everything at once&lt;/H2&gt;
&lt;P&gt;Five disciplines is a wall, and real teams cannot fund all five at once. The right order depends on whether the platform is being built fresh or already running.&lt;/P&gt;
&lt;P&gt;For a system &lt;STRONG&gt;already in production and already in pain&lt;/STRONG&gt;, the SRE community's &lt;A href="https://sre.google/sre-book/part-III-practices/" target="_blank"&gt;hierarchy of reliability needs&lt;/A&gt; gives the most defensible starting order: &lt;EM&gt;monitoring and observability first&lt;/EM&gt; (you cannot fix what you cannot see), &lt;EM&gt;then incident response&lt;/EM&gt; (close the bleeding cleanly), &lt;EM&gt;then resilience patterns&lt;/EM&gt; (idempotency, retries, decoupling) so the bleeding has fewer reasons to start, &lt;EM&gt;then flexibility and delivery&lt;/EM&gt; so safe change can travel at speed. Cost discipline runs alongside throughout, never as the headline.&lt;/P&gt;
&lt;P&gt;For a system &lt;STRONG&gt;being built fresh&lt;/STRONG&gt;, the order in this post (flexibility, resilience, observability, delivery, cost) reflects the &lt;A href="https://learn.microsoft.com/azure/well-architected/" target="_blank"&gt;Azure Well-Architected Framework's&lt;/A&gt; emphasis on designing for change, failure, and visibility before scaling teams or workloads. Both orders are defensible. What is not defensible is leaving any of the five for later.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;The most concrete starter from this post: request id propagation.&lt;/STRONG&gt; A single correlation identifier travelling through every service hop, every queue, every async boundary, costs hours up front and pays back every time someone has to debug production for the rest of the platform's life. It is the smallest unit of the observability discipline and the foundation that the dashboards, traces, and incident response in Part 2 all depend on.&lt;/P&gt;
&lt;H2&gt;The shift&lt;/H2&gt;
&lt;P&gt;The most important transformation in scaling a platform is not technical. It is mindset.&lt;/P&gt;
&lt;P&gt;The shift is from &lt;STRONG&gt;project thinking to platform thinking&lt;/STRONG&gt;:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Build reusable capabilities, not one-off solutions&lt;/LI&gt;
&lt;LI&gt;Design systems for long-term evolution, not the next release&lt;/LI&gt;
&lt;LI&gt;Enable other teams, not just deliver for one team&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Tools change. Cloud services evolve. The architectural fashions of this year will not be the architectural fashions of the next. What persists is the discipline behind the choices. Scalable systems are not built by tools. They are built by teams that treat design as continuous work. The same discipline shows up again in Part 2 (operating these systems) and Part 3 (using AI to augment that work). The tools change. The disciplines do not.&lt;/P&gt;
&lt;HR /&gt;
&lt;P&gt;&lt;STRONG&gt;Want to discuss?&lt;/STRONG&gt; What single design choice has paid the most dividends in the platforms you run? Drop a comment with patterns you have seen in your environment. Every reply gets read.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Next in this series:&lt;/STRONG&gt; &lt;EM&gt;Running Cloud Native Platforms: Why Day 2 Decides Everything.&lt;/EM&gt; Building is half the journey. The next post looks at what it takes to operate these platforms once they are in production.&lt;/P&gt;
&lt;!--
  Taxonomy:
    Primary product: Azure
    Secondary: .NET, GitHub
    Tags: Cloud Architecture, Platform Engineering, Microservices, Reliability, FinOps

  Visuals embedded:
    1. assets/diagram-1-pillar-map.png  (source: assets/diagram-1-pillar-map.mmd, Mermaid)
    2. assets/diagram-2-reference-architecture.png (source: assets/diagram-2-reference-architecture.py, matplotlib)

  Re-render command:
    &amp; "$env:USERPROFILE\.agents\skills\technical-blog-writer\tools\render-visuals.ps1" `
        -ArticleFolder &lt;part-1-build folder&gt; -Force
    python &lt;part-1-build folder&gt;\assets\diagram-2-reference-architecture.py
--&gt;</description>
      <pubDate>Thu, 21 May 2026 22:41:29 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/cloud-native-platforms-build/ba-p/4519605</guid>
      <dc:creator>KishoreKumarPattabiraman</dc:creator>
      <dc:date>2026-05-21T22:41:29Z</dc:date>
    </item>
    <item>
      <title>Cloud Native Platforms: Run</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/cloud-native-platforms-run/ba-p/4520188</link>
      <description>&lt;P&gt;&lt;STRONG&gt;Audience:&lt;/STRONG&gt; SREs (Site Reliability Engineers), platform engineers, engineering managers running production systems&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Reading time:&lt;/STRONG&gt; 8 minutes&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Series:&lt;/STRONG&gt; Cloud Native Platforms. Build, Run, Evolve. This is Part 2 of 3.&lt;/P&gt;
&lt;HR /&gt;
&lt;P&gt;Most systems are designed thoughtfully.&lt;/P&gt;
&lt;P&gt;Most operations are inherited reactively.&lt;/P&gt;
&lt;P&gt;The systems that survive are not the ones built with the most care. They are the ones operated with the most discipline. Production has a way of revealing every shortcut taken during design and every assumption left unverified.&lt;/P&gt;
&lt;P&gt;This post is about what it takes to operate a platform once the build is done.&lt;/P&gt;
&lt;H2&gt;How they are run, not how they are built&lt;/H2&gt;
&lt;P&gt;Systems are not defined by how they are built. They are defined by how they are run.&lt;/P&gt;
&lt;P&gt;A well-designed system that is operated reactively will fail in production. A modestly designed system that is operated with discipline will outperform it. Five operational disciplines decide which side of that line a platform lives on. Each one is engineering work, not a checklist for someone else to handle.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;EM&gt;Figure 1. The incident lifecycle as a state machine. The states are not optional steps. They are the contract between the team and the system.&lt;/EM&gt;&lt;/P&gt;
&lt;H2&gt;1. Observability is the backbone of reliability&lt;/H2&gt;
&lt;P&gt;Without observability, every operation becomes a guess. As systems grow, the cost of guessing rises faster than the cost of seeing.&lt;/P&gt;
&lt;P&gt;Part 1 of this series argued that observability is a design property: instrumentation contracts, request id propagation, structured logging schemas. Production is where those design choices either pay off or do not. Strong observability in production is a contract that lets any engineer answer three questions in minutes: what failed, why it failed, and what the impact was. The shape of that contract matters more than the tool that implements it. (This three-question framing is community-popularised through the SRE community and writers such as Charity Majors. See &lt;A href="https://www.honeycomb.io/what-is-observability" target="_blank"&gt;Honeycomb's &lt;EM&gt;What is Observability&lt;/EM&gt;&lt;/A&gt; for the canonical articulation of the three-pillars and question framing; the substance is older than the framing.)&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Dashboards organised around user journeys, not infrastructure components&lt;/LI&gt;
&lt;LI&gt;Service level indicators (SLIs: the specific measurements you care about, e.g., success rate, p99 latency) chosen from the user's perspective, not the database's&lt;/LI&gt;
&lt;LI&gt;Alerts that page only on burn-rate against an SLO (Service Level Objective: the target value of an SLI, e.g., 99.9% of requests complete in under 800ms over a rolling month) using a multi-window strategy. A short window catches fast burns; a long window catches slow drifts. This is what makes SLOs operational rather than decorative.&lt;/LI&gt;
&lt;LI&gt;Sampling and retention tuned for cost, but never for blind spots&lt;/LI&gt;
&lt;LI&gt;The distinction between MTTA (mean time to acknowledge: how fast someone notices) and MTTR (mean time to restore: how fast service returns) tracked separately. Conflating them hides whether the team's bottleneck is detection, response, or fix.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: rebuild the operational view around two or three user journeys (sign-in, place order, view history) rather than per-component charts. Tie alerts to error budget burn rather than raw threshold crossings. Track MTTA and MTTR separately so the team's actual bottleneck (detection, response, or fix) is visible. The investment is rethinking what to measure, not buying a new tool. The return is that incidents stop being discovered by customer complaints first. Teams that make this shift typically find their existing telemetry was sufficient; only the questions being asked of it were wrong.&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;EM&gt;If a dashboard cannot answer "what is the user experiencing right now", it is not an observability dashboard. It is decoration.&lt;/EM&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;2. Alerts are signals, not notifications&lt;/H2&gt;
&lt;P&gt;More alerts do not mean better monitoring. In practice, the opposite is true. Once alerts outpace the team's ability to act, important signals start getting missed.&lt;/P&gt;
&lt;P&gt;Effective alerting works to a small set of rules:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Severity that maps to action, not to technical category&lt;/LI&gt;
&lt;LI&gt;Ownership baked in, never inferred at runtime&lt;/LI&gt;
&lt;LI&gt;Thresholds tied to user impact, not raw metric values&lt;/LI&gt;
&lt;LI&gt;Noise treated as a defect, with a regular review cadence&lt;/LI&gt;
&lt;LI&gt;Suppression and grouping for known multi-alert patterns&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: audit every alert against one test, "what action would I take in the next five minutes if this fires now?" Demote alerts with no answer to dashboards. Remove alerts where the answer is the same as another alert's. Group related alerts so one incident produces one page, not twelve. Most teams discover their alert volume drops by an order of magnitude after a thorough audit, and the alerts that remain start getting trusted again. Trust is the precondition for every other operational practice. Without it, on-call rotations decay into noise filtering and the real signals get missed.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;EM&gt;Figure 2. From raw events to pages, in approximate orders of magnitude. The numbers vary by team and workload; what does not vary is that each stage needs to remove one to two orders of magnitude of noise. Teams that page on raw events end up with on-call rotations nobody trusts.&lt;/EM&gt;&lt;/P&gt;
&lt;H2&gt;3. Incident response is a practiced muscle&lt;/H2&gt;
&lt;P&gt;Failures are inevitable. Unstructured response is not.&lt;/P&gt;
&lt;P&gt;The teams that recover quickly do not improvise during incidents. They follow a structure that has been practiced when nothing was on fire. The structure is intentionally simple, because incident time is the worst time to negotiate roles.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Clear roles: incident lead, communications lead, scribe, subject matter expert (the RACI model, Responsible-Accountable-Consulted-Informed, adapted for incident response)&lt;/LI&gt;
&lt;LI&gt;Defined escalation paths with clear handoff criteria. Escalation means re-paging to a higher tier or specialist, not returning to detection. The lifecycle diagram in Figure 1 makes the distinction explicit.&lt;/LI&gt;
&lt;LI&gt;Runbooks for the top failure modes, kept short enough to actually be read&lt;/LI&gt;
&lt;LI&gt;Status communication on a fixed cadence, even when there is nothing new to say. Customer comms and internal comms are tracked separately.&lt;/LI&gt;
&lt;LI&gt;Blameless postmortems (focus on the system that allowed the failure, not the person who pushed the button) that produce action items the team actually completes&lt;/LI&gt;
&lt;LI&gt;Game days: scheduled exercises that simulate failure modes (region outage, dependency unavailability, traffic spike) under controlled conditions, so gaps in runbooks are found before incidents do&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: name the incident lead and the comms lead before the first message goes out. Write runbooks short enough to be scannable at 3 AM. Run blameless postmortems with action items that actually get tracked to completion. Schedule game days quarterly so the runbooks are exercised before real incidents. Teams that operate with this structure do not have more engineers; they have engineers who are not single points of failure during recovery. The deepest experts stay the deepest experts, but the platform stops depending on whether they happen to be online.&lt;/P&gt;
&lt;H3&gt;Implementation note&lt;/H3&gt;
&lt;P&gt;A short, well-structured runbook outperforms a long, exhaustive one. The goal during an incident is not to think. It is to act on a procedure that has been thought through in calmer times.&lt;/P&gt;
&lt;LI-CODE lang="yaml"&gt;
# Runbook header pattern (keep it scannable in incident time)
title: High latency on order API
slo_protected:                  # this runbook protects two SLOs
  - order-completion-success
  - order-completion-latency
severity:                       # derived from burn rate, not declared
  fast_burn: P1                 # 14.4x budget burn over 1 hour =&amp;gt; page now
  slow_burn: P2                 # 6x budget burn over 6 hours =&amp;gt; investigate
owner: payments-team
indicators:                     # triggers for evaluation, not severity
  - p99 (99th-percentile) latency exceeds the SLO target for 5 min
  - error rate exceeds the SLO target for 3 min on order-completion
first_actions:
  - Open the order-journey dashboard. Confirm impact in business terms.
  - Check Service Bus queue depth and dead-letter rate (the most common
    cause of API latency under load is downstream backpressure)
  - Verify Cosmos DB RU/s saturation and partition hotspots
  - Inspect the most recent deployment for behavioural changes
escalate_if:
  - Latency does not recover in 15 min
  - Error rate exceeds 5% (fast burn against the SLO)
  - Customer reports arrive before our own signals do
rollback_path:
  - Feature flag "new-order-pipeline" can be disabled per-tenant
  - Last known good deployment id is in the release tracker
note_on_scaling:
  # CPU is rarely the cause of latency in this service. Scale only after
  # confirming the bottleneck is compute, not a downstream dependency or
  # queue depth. Adding capacity to a saturated downstream amplifies the
  # incident; it does not resolve it.
&lt;/LI-CODE&gt;
&lt;P&gt;The general principle behind that last note travels beyond this runbook: scale-out is the right remediation for compute saturation, not for downstream saturation. When latency rises because a database, queue, or external dependency is saturated, adding capacity in front of the bottleneck moves more requests into the bottleneck and makes the incident worse. This is one of the most common operational mistakes when the dashboard shows red and the on-call instinct says "add more".&lt;/P&gt;
&lt;H2&gt;4. Release confidence is engineered&lt;/H2&gt;
&lt;P&gt;Releases get harder as systems grow. The platforms that ship confidently at scale have engineered the path, not learned to fear it.&lt;/P&gt;
&lt;P&gt;The patterns that change the math:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Feature flags that allow change without deploy&lt;/LI&gt;
&lt;LI&gt;Canary deployments (releasing the new version to a small slice of traffic first, watching error budget burn before continuing) that surface problems on a small slice&lt;/LI&gt;
&lt;LI&gt;Gradual rollouts with automated rollback triggers&lt;/LI&gt;
&lt;LI&gt;Database migrations split from application releases&lt;/LI&gt;
&lt;LI&gt;Release coordination that scales with services, not with team size&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: every change ships behind a feature flag, canary deployments take a small slice of traffic first, and rollback is a one-click step in the pipeline rather than a procedure to be invented during an incident. The cost is the discipline of building rollback paths and exercising them. The return is releases that stop being events. Issues that previously triggered full rollbacks get isolated to a slice and rolled back automatically before they reach most users. The willingness to ship smaller, more frequent changes follows directly from the confidence that bad changes can be undone fast.&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;EM&gt;Big releases feel safe because they are rare. They are actually risky because every change rides together.&lt;/EM&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;5. Reliability is continuous, not a milestone&lt;/H2&gt;
&lt;P&gt;Reliability is not achieved through tools alone. It requires continuous refinement, feedback-driven improvement, and a budget that the team can spend on operational work without negotiating each time.&lt;/P&gt;
&lt;P&gt;The disciplines that keep systems reliable over years are codified well in the SRE-book framing of service level objectives and error budgets (the canonical reference is the &lt;A href="https://sre.google/sre-book/service-level-objectives/" target="_blank"&gt;Google SRE Book chapter on Service Level Objectives&lt;/A&gt;, with the operational follow-up in the &lt;A href="https://sre.google/workbook/alerting-on-slos/" target="_blank"&gt;SRE Workbook chapter on alerting on SLOs&lt;/A&gt;). The names matter less than the practice they enable.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;SLOs&lt;/STRONG&gt; chosen from the user's perspective, with two or three per service rather than ten. More SLOs means none of them shape behaviour.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Error budgets&lt;/STRONG&gt;: the inverse of the SLO, expressing how much unreliability the team is willing to spend in a window. Used up early in the month means slow down on releases. Healthy means feature work keeps moving.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Multi-window burn-rate alerting&lt;/STRONG&gt; turns SLOs from dashboards into pages: short window catches catastrophic failures, long window catches slow drift. Without burn-rate alerting, SLOs are observation, not operation. (The pattern is documented in the &lt;A href="https://sre.google/workbook/alerting-on-slos/" target="_blank"&gt;SRE Workbook&lt;/A&gt;.)&lt;/LI&gt;
&lt;LI&gt;Reliability work has its own backlog, prioritised against features. Not a wishlist after every incident.&lt;/LI&gt;
&lt;LI&gt;Regular game days that exercise failure modes (region failover, dependency outage, traffic spike) before they happen for real&lt;/LI&gt;
&lt;LI&gt;Capacity planning informed by data, not by anxiety&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: define two or three SLOs per service, expressed from the user's perspective. Compute the error budget weekly. When the budget is healthy, ship feature work. When the budget is burning fast, slow down and fix the cause. The conversation about which incidents matter and which can wait becomes possible because there is a shared number to point at. Reliability becomes a quantified property of the platform, not an opinion debated at every retrospective. Teams that adopt this discipline stop having the recurring "how reliable do we need to be?" argument and start having data-grounded trade-off discussions instead.&lt;/P&gt;
&lt;H2&gt;A scenario that ties it together&lt;/H2&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;A platform was launching a new region. The build had gone well. Day 1 was clean. Two weeks in, latency started creeping up during peak hours. Alerts fired on raw thresholds, but no one could tell which ones to trust. Incident calls turned into long debugging sessions because three different teams owned overlapping pieces of the request path.&lt;/P&gt;
&lt;P&gt;The team did not start by buying a new tool. They started by treating operations as engineering work. The dashboard was redesigned around the user journey. Alerts were audited and most were demoted or removed. Roles for incident response were written down. A short runbook covered the top failure modes. Releases were broken into canary slices behind feature flags.&lt;/P&gt;
&lt;P&gt;None of this was new. It was discipline applied consistently to work that was previously assumed to be someone else's. The next region launch took half the effort, and the team's mean time to restore on the failures that did happen was measurably lower.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;What teams get wrong&lt;/H2&gt;
&lt;P&gt;The common pattern is &lt;STRONG&gt;treating Day 2 as the cost of Day 1.&lt;/STRONG&gt; Teams design beautifully, ship fast, then quietly absorb the operational debt. Dashboards proliferate. Alerts grow louder. Postmortems pile up.&lt;/P&gt;
&lt;P&gt;The fix is not more dashboards. It is treating operations as engineering work with the same rigour as feature delivery. Operability is a property the system either has or does not. It is not earned by adding monitoring. It is earned by designing for visibility and operating with discipline.&lt;/P&gt;
&lt;H2&gt;Where to start&lt;/H2&gt;
&lt;P&gt;&lt;STRONG&gt;The most concrete starter from this post: an alert audit.&lt;/STRONG&gt; List every alert that fires in the next week and apply a single test to each one: "what action would I take in the next five minutes?" Demote the alerts that have no answer. Remove the alerts where the answer is the same as another alert's. The audit takes a morning. The result usually halves alert volume and lifts trust on what remains, which is the precondition for every other operational practice in this post.&lt;/P&gt;
&lt;H2&gt;The shift&lt;/H2&gt;
&lt;P&gt;The most important shift in maturity is not technical. It is in stance.&lt;/P&gt;
&lt;P&gt;The shift is from &lt;STRONG&gt;shipping software to operating systems&lt;/STRONG&gt;:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Operations is not a phase that follows engineering. It is engineering.&lt;/LI&gt;
&lt;LI&gt;Reliability is not a milestone reached. It is a discipline practiced.&lt;/LI&gt;
&lt;LI&gt;Incidents are not interruptions to the work. They are the work.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The teams that internalise this shift run platforms that are smaller, calmer, and more trusted. They do not have fewer incidents because their systems are more advanced. They have fewer incidents because their operational discipline is more consistent. Part 3 of this series argues that the same discipline applies again, in a different domain: the practices that make platforms operable are the practices that make AI useful in delivery.&lt;/P&gt;
&lt;HR /&gt;
&lt;P&gt;&lt;STRONG&gt;Want to discuss?&lt;/STRONG&gt; What is the one operational practice your team adopted that changed how you sleep at night? Drop a comment with patterns you have seen in your environment. Every reply gets read.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Previously in this series:&lt;/STRONG&gt; &lt;EM&gt;Building Cloud Native Platforms That Scale: Patterns That Actually Work.&lt;/EM&gt; The first post covered the design choices that make scale possible.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Next in this series:&lt;/STRONG&gt; &lt;EM&gt;AI-First Platform Engineering: From Copilot to Agentic Delivery.&lt;/EM&gt; Cloud helped us scale infrastructure. The next post looks at how AI is now changing how we build and run platforms.&lt;/P&gt;
&lt;!--
  Taxonomy:
    Primary product: Azure
    Secondary: Azure Monitor, Application Insights
    Tags: Site Reliability Engineering, Observability, Incident Management, DevOps, Platform Engineering

  Visuals embedded:
    1. assets/diagram-1-incident-lifecycle.png  (source: assets/diagram-1-incident-lifecycle.mmd, Mermaid)
    2. assets/diagram-2-alert-funnel.png        (source: assets/diagram-2-alert-funnel.py, matplotlib)
--&gt;</description>
      <pubDate>Thu, 21 May 2026 22:41:09 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/cloud-native-platforms-run/ba-p/4520188</guid>
      <dc:creator>KishoreKumarPattabiraman</dc:creator>
      <dc:date>2026-05-21T22:41:09Z</dc:date>
    </item>
    <item>
      <title>Cloud Native Platforms: Evolve</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/cloud-native-platforms-evolve/ba-p/4520195</link>
      <description>&lt;P&gt;&lt;STRONG&gt;Audience:&lt;/STRONG&gt; Engineering leaders, platform architects, senior developers exploring how to operationalise AI in their teams&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Reading time:&lt;/STRONG&gt; 8 minutes&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Series:&lt;/STRONG&gt; Cloud Native Platforms. Build, Run, Evolve. This is Part 3 of 3.&lt;/P&gt;
&lt;HR /&gt;
&lt;P&gt;Cloud helped us scale infrastructure.&lt;/P&gt;
&lt;P&gt;AI is starting to do the same thing for the work around the code: the planning, the testing, the release communication, the incident triage, the writing that surrounds writing software.&lt;/P&gt;
&lt;P&gt;The conversation about AI in software has narrowed too quickly to "Copilot in the editor". The bigger story is happening across the lifecycle. Planning, design, development, testing, release, and operations are all being augmented at once. The platforms that adopt AI well are not the ones with the most usage. They are the ones with the clearest discipline around how it is used.&lt;/P&gt;
&lt;P&gt;This post is about that discipline.&lt;/P&gt;
&lt;H2&gt;AI is changing how we engineer, not how we type&lt;/H2&gt;
&lt;P&gt;AI is not changing how we write code. It is changing how we engineer software.&lt;/P&gt;
&lt;P&gt;Code generation is the surface. Underneath it, AI is reshaping the unit of leverage. The question is no longer how fast a developer can type. It is how well a workflow can be expressed as a reusable engineering asset. Six disciplines determine whether AI moves the needle on outcomes or just adds another tool to the stack.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;EM&gt;Figure 1. AI across the SDLC. Each phase has clear AI assist points and clear human-owned validations. The boundary is not negotiable. It is the design.&lt;/EM&gt;&lt;/P&gt;
&lt;H2&gt;1. From assistance to augmentation&lt;/H2&gt;
&lt;P&gt;Early AI tools focused on assisting individual developers. Code suggestions. Autocomplete. Quick refactors. The value was real but bounded by the editor.&lt;/P&gt;
&lt;P&gt;The shift now is into structured workflows that span the lifecycle. The unit of leverage is no longer a single suggestion. It is a sequence of actions executed reliably across phases. ("Agentic" later in this post means a system that makes its own next-step decisions inside guardrails. A workflow follows a fixed sequence; an agent chooses the path.)&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Code generation has become baseline, not differentiator&lt;/LI&gt;
&lt;LI&gt;Workflow generation is where the largest gains live&lt;/LI&gt;
&lt;LI&gt;Multi-step assistance with explicit human checkpoints&lt;/LI&gt;
&lt;LI&gt;Context that travels across tools, not just within one&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: start with the single highest-volume writing task on the team (commit messages, code review comments, release notes, postmortem first drafts) and turn the AI assist for that task into a shared workflow rather than each individual's private trick. The cost is one engineer's afternoon documenting the workflow and the eval set. The return is that every engineer on the team inherits the work, and the task that used to consume an engineer's morning every two weeks becomes a background step in the release process. Workflow generation, not faster typing, is where the gains compound across a team.&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;EM&gt;Code suggestions help one developer. Reusable workflows help the next ten.&lt;/EM&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;2. AI across the SDLC, with guardrails&lt;/H2&gt;
&lt;P&gt;AI now has a useful role at every phase of delivery. The role is different at each phase, and the guardrails are different too.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;th&gt;Phase&lt;/th&gt;&lt;th&gt;What AI helps with&lt;/th&gt;&lt;th&gt;What humans must validate&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Plan&lt;/td&gt;&lt;td&gt;Breaking down requirements, drafting acceptance criteria&lt;/td&gt;&lt;td&gt;Domain context, business priorities, customer impact&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Build&lt;/td&gt;&lt;td&gt;Code generation, refactoring, scaffolding&lt;/td&gt;&lt;td&gt;Architectural fit, security boundaries, performance&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Test&lt;/td&gt;&lt;td&gt;Test case generation, edge case discovery&lt;/td&gt;&lt;td&gt;Coverage of business-critical paths, regulatory cases&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Release&lt;/td&gt;&lt;td&gt;Release notes, changelog summaries, communication drafts&lt;/td&gt;&lt;td&gt;Accuracy, tone, customer-facing claims&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Operate&lt;/td&gt;&lt;td&gt;Log triage, incident summaries, runbook drafts&lt;/td&gt;&lt;td&gt;Root cause attribution, action item ownership&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;The guardrails are not optional decoration. They are the design.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: stage AI assists for release communication (changelog drafting, customer-facing release notes, internal release announcements) and require a human review before anything goes out. The draft arrives consistently, faster than a human could produce, and easier to compare across releases. The reviewer is not eliminated; the reviewer is moved from author to editor, which is where their judgment actually matters. Teams that adopt this pattern stop missing release-note deadlines and stop publishing inconsistent communication across products.&lt;/P&gt;
&lt;H2&gt;3. From prompts to reusable assets&lt;/H2&gt;
&lt;P&gt;Many teams begin with prompt experimentation. Individuals find techniques that work for their tasks. The result is a patchwork of personal practices that do not survive a team change.&lt;/P&gt;
&lt;P&gt;The compounding value comes when prompts mature into reusable engineering assets.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;EM&gt;Figure 2. The maturity model from prompts to agents. The value compounds at the workflow stage and accelerates at the agent stage. The disciplines that make agents safe are the same ones that made workflows reliable.&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;The maturity stages, in order of leverage:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Prompts&lt;/STRONG&gt;: ad-hoc, individual, hard to share&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Templates&lt;/STRONG&gt;: parameterised prompts versioned with the project&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Workflows&lt;/STRONG&gt;: multi-step sequences with clear inputs, outputs, checkpoints&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Agents&lt;/STRONG&gt;: autonomous task chains operating within explicit guardrails&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The diagram is a maturity ladder, not a graduation. In practice teams operate at all four stages simultaneously for different tasks. A senior engineer may use a one-off prompt to explore a refactor, run a versioned template for commit messages, hand off to a workflow for release notes, and trigger an agent for routine PR triage, all in the same hour. The point of the ladder is not to leave earlier stages behind. It is to know which stage a given task belongs to and to invest accordingly.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: pick the three prompts your team uses every week, codify them as parameterised templates in the same repository as the application code, and treat them as engineering artefacts (reviewed, versioned, owned). New engineers inherit the team's accumulated practice instead of building their own from scratch. Quality becomes consistent because the variance between individuals shrinks. Investment pays back in weeks, not quarters, and the maturity ladder keeps producing returns as the team moves from templates to workflows to agents.&lt;/P&gt;
&lt;H2&gt;4. Agentic delivery, with guardrails that survive a security review&lt;/H2&gt;
&lt;P&gt;The next stage is agentic. AI executes sequences of tasks within a defined scope. The risk is not that the agent will fail. It is that the system around the agent will not catch the failure, and that the failure modes are different in kind from traditional automation. Agents are non-deterministic, they can be manipulated through their inputs, and their actions can have side effects in systems the team does not own.&lt;/P&gt;
&lt;P&gt;Five guardrails make agentic delivery safe. The first four are necessary. The fifth is what carries the agent through a security review at a regulated enterprise.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Identity and scope&lt;/STRONG&gt;: the agent runs as a managed identity (or scoped service principal) with the smallest set of permissions that lets it do its job. Permissions are expressed as allowlists, not denylists. Tools fetched at runtime are subject to the same identity boundary as the agent itself.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Input quarantine&lt;/STRONG&gt;: anything the agent reads from a user-controlled source (work item bodies, PR descriptions, customer tickets) is treated as untrusted text. The agent does not execute instructions found in fetched content, and tool calls are validated against an output schema before execution. This is the prompt-injection mitigation, and it is the most common gap in agentic systems shipped today.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Cost and blast-radius caps&lt;/STRONG&gt;: every run has a maximum token budget, a maximum number of tool calls, and a maximum spend. Exceeding any cap aborts the run cleanly. Without caps, scoped credentials are not enough to bound the damage.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Evaluations and traceability&lt;/STRONG&gt;: agents are evaluated against a fixed test set before deployment, and on every prompt or model change. Every action is logged with inputs, outputs, the model and prompt versions used, and the reasoning trace where the model exposes one. Logs are redacted for secrets and personally identifiable information at write time.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Reversibility taxonomy&lt;/STRONG&gt;: actions are categorised by reversibility, not asserted to be reversible in general. A draft write to a private store is reversible. A post to a customer-facing channel is not reversible (deletion does not unsend). A database update may be reversible by a compensating transaction or not at all. Irreversible actions require human approval at the boundary, before they happen, not after. The agent is allowed to draft and stage. The human is the only one who is allowed to make the move that cannot be undone.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: start with one low-risk agent (release-notes drafter, PR triage assistant) running on read-only inputs, write-only-to-drafts permissions, and a hard cost cap per run. Require explicit human approval at the irreversible step. Wire up an evaluation set on day one, and rerun it on every prompt or model change. Treat regressions as failures, not warnings. The first agent the team ships is rarely the most valuable; it is the rehearsal that establishes the controls every later agent inherits. Teams that skip this rehearsal end up with an agent in production that no one feels safe extending.&lt;/P&gt;
&lt;H3&gt;Implementation note&lt;/H3&gt;
&lt;P&gt;An agent without a reversibility taxonomy and a regression eval set is a liability. The discipline is the same one that made workflows reliable: scoped identity, idempotency, traceability, and a clear boundary between machine action and human decision. The YAML below is illustrative, not a runtime contract; it is meant to show the shape of the controls a real agent definition would carry, not the syntax of any specific platform.&lt;/P&gt;
&lt;LI-CODE lang="yaml"&gt;# Agent run definition (illustrative; not a specific platform's syntax)
name: release-notes-drafter
trigger: pre-release
identity:
  type: managed-identity
  scope: tenant=&amp;lt;tenant-id&amp;gt; resource=release-tools/&amp;lt;app-id&amp;gt;
permissions:
  allow:
    - read: work-items in milestone (filter: state=Done)
    - read: pull-requests in milestone (filter: merged)
    - write: drafts/release-notes/${run-id}
  # Production channels are NOT in the allowlist. The agent cannot post.
limits:
  max_tokens_per_run: 80000
  max_tool_calls_per_run: 20
  max_runtime_seconds: 300
  max_cost_usd: 0.40
  on_exceeded: abort_with_partial_artifact
input_handling:
  treat_fetched_content_as: untrusted
  # Indirect prompt injection is mitigated by the layered discipline below,
  # not by a single feature flag. Each item is a separate control.
  enforce_instruction_hierarchy: true
  validate_tool_args_against_schema: true
  validate_outputs_against_schema: true
steps:
  - fetch: completed work items in milestone
  - draft: release notes from items
  - validate: required fields present
  - request-review:
      from: release-manager
      idempotency_key: ${milestone-id}-${draft-hash}
  - on-approval:
      action: post-to-internal-channel
      reversibility: not-reversible
      requires: explicit-human-click  # the agent does NOT click this
audit:
  log_inputs: true
  log_outputs: true
  redact:
    - secrets
    # Pattern-based: handles structured PII like emails, phones, IDs.
    - pii_patterns: [email, phone, national-id, payment-card, ip-address]
    # Entity-based: required for unstructured PII like names. Pattern alone
    # cannot redact a customer name without an entity-recognition step.
    - pii_entities: ner-based  # names, locations, organisations
  retain: 365_days  # tune to your audit policy, not to the demo
evaluation:
  test_set: tests/release-notes/eval-v3.jsonl
  on_prompt_change: rerun
  on_model_change: rerun
  fail_threshold: 5_percent_regression&lt;/LI-CODE&gt;
&lt;H2&gt;5. Where AI still needs human judgment&lt;/H2&gt;
&lt;P&gt;AI has clear boundaries. The boundaries are not embarrassing. They are the design.&lt;/P&gt;
&lt;P&gt;What must stay human-owned:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Architectural trade-offs and design decisions&lt;/LI&gt;
&lt;LI&gt;Security validation and threat modelling&lt;/LI&gt;
&lt;LI&gt;Correctness for business-critical and regulatory paths&lt;/LI&gt;
&lt;LI&gt;Domain context that has not been written down&lt;/LI&gt;
&lt;LI&gt;Accountability for outcomes, not just outputs&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The goal is collaboration, not replacement. The teams that get the most value from AI are not the ones with the most automation. They are the ones with the clearest sense of where automation ends and judgment begins.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: name the human-owned items explicitly in the team's working agreement (architecture, security, regulatory correctness, accountability) and audit every AI workflow against that list. When a workflow asks the AI to make a decision in any of those categories, redesign it so the AI prepares the analysis and a human makes the call. Most teams over-trust AI for one of these areas in their first six months and learn the hard way. Naming the boundary up front prevents the lesson from being paid in production. The clarity is the value; the model behind the workflow is interchangeable.&lt;/P&gt;
&lt;H2&gt;6. Responsible AI is engineering work&lt;/H2&gt;
&lt;P&gt;The first five disciplines decide whether AI moves the needle. The sixth decides whether the platform can defend the choices it makes with AI. Responsible AI is the engineering practice of building systems whose AI behaviour is fair, transparent, accountable, and safe by design, not by audit after the fact. Treating it as a compliance checkbox at the end of the project is how teams end up shipping AI workflows that fail security review, embarrass the company, or harm users.&lt;/P&gt;
&lt;P&gt;Six controls turn responsible AI from a policy into engineering work. These map directly onto the practices Microsoft and the broader industry have converged on, but the names matter less than the practice they enable.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Fairness in inputs and outputs.&lt;/STRONG&gt; The training data, eval set, and prompts are reviewed for systematic bias against any group the system serves. The eval set covers under-represented cases by design, not by accident, and regressions on those cases fail the build.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Transparency to end users.&lt;/STRONG&gt; When a user sees AI-generated content, they are told. When a decision is AI-assisted, the path from input to output is explainable in plain language, not just in a model card buried in documentation.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Content safety filters.&lt;/STRONG&gt; Inputs and outputs pass through safety classifiers (prompt injection, prohibited content, jailbreak patterns) before reaching the model and before reaching the user. Filtering decisions are logged and reviewable.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Accountability ownership.&lt;/STRONG&gt; Every AI workflow has a named owner who is accountable for its outcomes, not just its uptime. The owner has the authority to pause or roll back the workflow when harm is detected.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Data minimisation and residency.&lt;/STRONG&gt; The AI sees only the data it needs to do the task. Personally identifiable information and customer data are scoped, redacted, and kept inside the boundary the customer agreed to. Cross-tenant leakage is treated as a P1 incident, not a feature request.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Harm evaluation alongside quality evaluation.&lt;/STRONG&gt; The eval set measures harm potential (toxicity, hallucination on factual queries, leakage of confidential context) with the same rigour as it measures correctness. Both must pass for a release to ship.&lt;/LI&gt;
&lt;/UL&gt;
&lt;img /&gt;
&lt;P&gt;&lt;EM&gt;Figure 3. Responsible AI as a set of engineering controls around the AI workflow. The six controls fall into four categories: data discipline (fairness, data minimisation), model discipline (content safety, harm evaluation), deployment discipline (transparency to users), and governance (accountability ownership). All six are necessary; none is sufficient on its own.&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;In practice&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The pattern that works: write the responsible AI plan before the first agent ships, not after the first incident. Pick one workflow that touches user data or generates customer-facing content, and use it as the reference implementation: fairness review on the eval set, content safety filters wrapping the model call, transparency annotation in the UI, redaction of identifying details in logs, harm evals running alongside quality evals on every change, and a named owner with explicit pause authority. The first such workflow takes longer to ship than the unconstrained version. Every workflow after it inherits the controls and ships faster than it would have without them. Teams that defer responsible AI to a future quarter end up retrofitting it under pressure, which is the most expensive way to do it.&lt;/P&gt;
&lt;H2&gt;A scenario that ties it together&lt;/H2&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Picture a platform team several months into using Copilot. Adoption is high. Productivity dashboards show gains. But defect rates are not improving and lead time is flat. Leadership asks the obvious question: is AI actually helping, or just feeling like help?&lt;/P&gt;
&lt;P&gt;The answer is not to stop using AI. It is to change how AI is measured. Move adoption metrics to the background. Move outcome metrics to the front: defect escape rate, lead time for change, change failure rate, mean time to recovery. In parallel, promote the individual prompts that have proved themselves to shared templates, and the templates to versioned workflows. Retrofit responsible AI controls onto the workflows that shipped first: content safety filters, harm evaluations alongside quality evaluations, transparency annotations on customer-facing output, and a named owner for each workflow.&lt;/P&gt;
&lt;P&gt;Six months later, the picture is different. Defect rate improves on the parts of the codebase where reusable workflows were introduced. Onboarding for new engineers is visibly faster. Release notes are consistent across teams. The shift is from celebrating use to tracking outcomes, and once the team measures what matters, the tooling decisions start making themselves.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;What teams get wrong&lt;/H2&gt;
&lt;P&gt;The common pattern is &lt;STRONG&gt;measuring AI by usage, not by outcome&lt;/STRONG&gt;. Adoption metrics tell you who tried Copilot. They do not tell you whether defects dropped, lead time improved, or release notes got better.&lt;/P&gt;
&lt;P&gt;The fix is not less AI. It is better measurement. The four metrics named in the scenario above (defect escape rate, lead time for change, change failure rate, mean time to recovery) come from the &lt;A href="https://dora.dev/" target="_blank" rel="noopener"&gt;DORA research on software delivery performance&lt;/A&gt; and have become a useful default. Two warnings travel with them. First, attribution is hard: an AI workflow rolled out alongside a test refactor and a CI pipeline change cannot claim credit cleanly. Second, baselines matter more than headlines: a single quarter's improvement is not a trend, and a single team's gain is not the platform's gain. Outcome measurement done well needs a baseline window, an attribution discipline, and a kill criterion for workflows that are not paying back. Done poorly, it is just adoption metrics with better names.&lt;/P&gt;
&lt;P&gt;There is also the question of cost. AI usage carries a per-run token bill, an evaluation bill on every change, and (for agents) a cost cap that limits damage when something goes wrong. None of these are large compared to the engineering time saved when the workflow works. All of them are visible enough that a finance-aware reader will ask. Track them.&lt;/P&gt;
&lt;H2&gt;Where to start&lt;/H2&gt;
&lt;P&gt;&lt;STRONG&gt;The most concrete starter from this post: promote one personal prompt to a shared template.&lt;/STRONG&gt; Pick the prompt that gets used most often (commit messages, code reviews, release notes, debugging assist), move it from someone's notes into the repository where the team versions everything else, and watch what changes when the next person on the team runs it. That is the smallest unit of the workflow shift this post argues for, and it is the step where prompts stop being individual practice and start becoming engineering assets.&lt;/P&gt;
&lt;H2&gt;The shift&lt;/H2&gt;
&lt;P&gt;The shift is from &lt;STRONG&gt;building systems to building smarter systems&lt;/STRONG&gt;:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;AI does not replace engineers. It changes what an engineer's leverage looks like.&lt;/LI&gt;
&lt;LI&gt;The unit of value is the workflow, not the suggestion.&lt;/LI&gt;
&lt;LI&gt;The discipline that made platforms operable is the same discipline that makes AI useful.&lt;/LI&gt;
&lt;LI&gt;Responsible AI is not a compliance step. It is the sixth engineering discipline that lets the other five compound safely.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The series ends here, but the arc is consistent across all three posts. The disciplines that make platforms scale are the same disciplines that make AI useful. Build with discipline. Run with discipline. Evolve with discipline. The tools change. The disciplines do not.&lt;/P&gt;
&lt;HR /&gt;
&lt;P&gt;&lt;STRONG&gt;Want to discuss?&lt;/STRONG&gt; Where has AI moved the needle most in your delivery, and where has it disappointed you? Drop a comment with patterns you have seen in your environment. Every reply gets read.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Previously in this series:&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;EM&gt;Building Cloud Native Platforms That Scale: Patterns That Actually Work&lt;/EM&gt;. Part 1 covered the design choices that make scale possible.&lt;/LI&gt;
&lt;LI&gt;&lt;EM&gt;Running Cloud Native Platforms: Why Day 2 Decides Everything&lt;/EM&gt;. Part 2 covered the operational disciplines that decide production outcomes.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;This is the third and final post in the series.&lt;/P&gt;
&lt;!--
  Taxonomy:
    Primary product: GitHub Copilot, Azure
    Secondary: Microsoft 365 Copilot, AI Foundry
    Tags: AI Engineering, GitHub Copilot, Agentic AI, Developer Productivity, Platform Engineering

  Visuals embedded:
    1. assets/diagram-1-sdlc-with-ai.png  (source: assets/diagram-1-sdlc-with-ai.mmd, Mermaid)
    2. assets/diagram-2-maturity-model.png (source: assets/diagram-2-maturity-model.mmd, Mermaid)
--&gt;</description>
      <pubDate>Thu, 21 May 2026 22:40:33 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/cloud-native-platforms-evolve/ba-p/4520195</guid>
      <dc:creator>KishoreKumarPattabiraman</dc:creator>
      <dc:date>2026-05-21T22:40:33Z</dc:date>
    </item>
    <item>
      <title>WAR, Azure Advisor, and Us (Azure Arch Diagram Builder): Three Ways to Score an Azure Architecture</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/war-azure-advisor-and-us-azure-arch-diagram-builder-three-ways/ba-p/4521611</link>
      <description>&lt;P&gt;&lt;STRONG&gt;Author:&lt;/STRONG&gt; Arturo Quiroga, Azure AI services Engineer - Senior Partner Solutions Architect — Microsoft&lt;/P&gt;
&lt;HR /&gt;
&lt;P&gt;A few days ago I published &lt;A href="https://techcommunity.microsoft.com/blog-draft-azure-architecture-blog.md" target="_blank"&gt;&lt;EM&gt;From Prompt to Production: Building Azure Architecture Diagrams with AI&lt;/EM&gt;&lt;/A&gt;, introducing the open-source &lt;A href="https://aka.ms/diagram-builder" target="_blank"&gt;Azure Architecture Diagram Builder&lt;/A&gt;. One feature got more follow-up questions than any other: the &lt;STRONG&gt;Well-Architected Framework (WAF) validation&lt;/STRONG&gt;. Architects from partners and customers — many of whom already use Azure Advisor and the Well-Architected Review — wanted to know exactly what scoring algorithm we use, how it compares to Microsoft's official tools, and whether they should be using all three.&lt;/P&gt;
&lt;P&gt;This post is that answer. It's a deep dive into how design-time WAF validation works, how Microsoft's two official WAF assessment algorithms work, and where each fits in the architecture lifecycle.&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;TL;DR.&lt;/STRONG&gt; Microsoft ships two WAF assessment vehicles — the &lt;STRONG&gt;Well-Architected Review&lt;/STRONG&gt; (questionnaire, scored from human answers) and the &lt;STRONG&gt;Azure Advisor score&lt;/STRONG&gt; (healthy-resources-÷-applicable-resources weighted per subcategory, with Defender Secure Score for Security and cost-weighted math for Cost). Both require either a human filling in a form or live Azure telemetry. Our app runs &lt;STRONG&gt;at design time on a diagram&lt;/STRONG&gt;, before anything is deployed, using a hybrid pipeline: a deterministic rule pre-scan followed by an LLM refinement pass. Same five WAF pillars, different lifecycle stage. Complementary, not competitive.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;HR /&gt;
&lt;H2 id="why-design-time-validation-matters"&gt;Why design-time validation matters&lt;/H2&gt;
&lt;P&gt;Every cost overrun, reliability gap, and security incident I've ever debugged was cheaper to fix on a whiteboard than in production. Yet most WAF tooling assumes the architecture already exists — either because there are deployed resources to scan (Advisor) or because someone has built enough of it to answer 60 specific questions about it (WAR).&lt;/P&gt;
&lt;P&gt;That leaves a gap. &lt;STRONG&gt;Between "rough sketch" and "deployed resource group" there is no algorithmic WAF feedback loop.&lt;/STRONG&gt; That's the gap the Diagram Builder fills.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="microsofts-two-official-waf-assessment-algorithms"&gt;Microsoft's two official WAF assessment algorithms&lt;/H2&gt;
&lt;P&gt;Before describing our approach, it's worth being precise about what Microsoft already ships, because the term "WAF assessment algorithm" can mean either of two very different things.&lt;/P&gt;
&lt;H3 id="1-azure-well-architected-review-war--questionnaire-based"&gt;1. Azure Well-Architected Review (WAR) — questionnaire-based&lt;/H3&gt;
&lt;P&gt;The &lt;A href="https://learn.microsoft.com/assessments/azure-architecture-review/" target="_blank"&gt;Well-Architected Review&lt;/A&gt; is a free self-assessment hosted on Microsoft Learn.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table class="lia-border-color-custom-d6dee6 lia-border-style-solid" border="1" style="width: 100%; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Aspect&lt;/th&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Detail&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;STRONG&gt;Input&lt;/STRONG&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Human answers to ~60 questions mapped to the WAF pillar checklists&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;STRONG&gt;Workload variants&lt;/STRONG&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Core WAR, plus AI/ML, IoT, SAP on Azure, Azure Stack Hub, SaaS, Mission Critical&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;STRONG&gt;Scoring&lt;/STRONG&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Derived from the answers — each "no" or unanswered question subtracts from the pillar score&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;STRONG&gt;Output&lt;/STRONG&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Per-pillar maturity score + prioritized recommendations + optional Advisor integration&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;STRONG&gt;Improvement tracking&lt;/STRONG&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;"Milestones" (point-in-time snapshots)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;STRONG&gt;When to use&lt;/STRONG&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Periodic deep reviews; greenfield design baselining; brownfield audits&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;WAR is human-driven. The algorithm is essentially &lt;EM&gt;"how many of the recommended practices have you confirmed you do?"&lt;/EM&gt; — which is exactly the right algorithm when the assessor is the workload team itself.&lt;/P&gt;
&lt;H3 id="2-azure-advisor-score--telemetry-based"&gt;2. Azure Advisor Score — telemetry-based&lt;/H3&gt;
&lt;P&gt;The &lt;A href="https://learn.microsoft.com/azure/advisor/advisor-score#calculation-of-advisor-score" target="_blank"&gt;Advisor score&lt;/A&gt; is the closest thing Microsoft ships to a real, deterministic WAF &lt;EM&gt;algorithm&lt;/EM&gt;. It runs continuously over your deployed Azure resources.&lt;/P&gt;
&lt;P&gt;The math:&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;Pillar-specific overrides:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Security&lt;/STRONG&gt; uses Microsoft Defender for Cloud's &lt;STRONG&gt;Secure Score&lt;/STRONG&gt; model.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Cost&lt;/STRONG&gt; weights by retail $ cost of healthy resources, plus age-of-recommendation weighting; postponed/dismissed items are removed from the denominator.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Reliability / Performance / Operational Excellence&lt;/STRONG&gt; use the healthy-resources ratio above.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Key terms:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;EM&gt;Healthy resource&lt;/EM&gt; — a deployed resource with no open Advisor recommendation against it for that pillar.&lt;/LI&gt;
&lt;LI&gt;&lt;EM&gt;Total applicable&lt;/EM&gt; — resources Advisor was able to evaluate (excludes dismissed/snoozed).&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Advisor is the right tool once you're in production. It cannot help you before deployment, because there is nothing to count as "healthy" or "applicable."&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="the-missing-stage-design-time"&gt;The missing stage: design time&lt;/H2&gt;
&lt;P&gt;Here's the lifecycle, with each tool's domain shaded:&lt;/P&gt;
&lt;PRE class="mermaid"&gt;&amp;nbsp;&lt;/PRE&gt;
&lt;PRE class="mermaid"&gt;&amp;nbsp;&lt;/PRE&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;PRE class="mermaid"&gt;&lt;CODE&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;/CODE&gt;&lt;STRONG&gt;Design / Diagram&lt;/STRONG&gt; — &lt;EM&gt;Diagram Builder validation&lt;/EM&gt; runs here.&lt;/PRE&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Operate / Observe&lt;/STRONG&gt; — &lt;EM&gt;Azure Advisor&lt;/EM&gt; runs here continuously.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Periodic Review&lt;/STRONG&gt; — &lt;EM&gt;WAR&lt;/EM&gt; runs here, typically quarterly or at major milestones.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;These three stages are sequential and complementary. Our app does not replace Advisor or WAR — it adds a feedback loop earlier in the lifecycle, where corrections are cheapest.&lt;/STRONG&gt;&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="how-design-time-validation-works-in-the-diagram-builder"&gt;How design-time validation works in the Azure Architecture Diagram Builder&lt;/H2&gt;
&lt;P&gt;The validator is a &lt;STRONG&gt;two-phase hybrid pipeline&lt;/STRONG&gt;: deterministic local rules first, then LLM refinement. The full source lives in three files:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://techcommunity.microsoft.com/../src/services/architectureValidator.ts" target="_blank"&gt;&lt;CODE&gt;src/services/architectureValidator.ts&lt;/CODE&gt;&lt;/A&gt; — orchestrator and prompt&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://techcommunity.microsoft.com/../src/services/wafPatternDetector.ts" target="_blank"&gt;&lt;CODE&gt;src/services/wafPatternDetector.ts&lt;/CODE&gt;&lt;/A&gt; — topology + service rule engine&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://techcommunity.microsoft.com/../src/data/wafRules.ts" target="_blank"&gt;&lt;CODE&gt;src/data/wafRules.ts&lt;/CODE&gt;&lt;/A&gt; — the rule knowledge base&lt;/LI&gt;
&lt;/UL&gt;
&lt;PRE class="mermaid"&gt;&amp;nbsp;&lt;/PRE&gt;
&lt;img /&gt;
&lt;PRE class="mermaid"&gt;&amp;nbsp;&lt;/PRE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;PRE class="mermaid"&gt;&amp;nbsp;&lt;/PRE&gt;
&lt;H3 id="phase-1--deterministic-rule-pre-scan-1-ms-no-llm"&gt;Phase 1 — Deterministic rule pre-scan (~1 ms, no LLM)&lt;/H3&gt;
&lt;P&gt;When you click &lt;STRONG&gt;Validate Architecture&lt;/STRONG&gt;, the validator runs a fully client-side rule engine against the diagram's services, connections, and groups. There are two kinds of rules:&lt;/P&gt;
&lt;H4 id="architecture-pattern-rules"&gt;Architecture-pattern rules&lt;/H4&gt;
&lt;P&gt;&lt;STRONG&gt;These fire when a topology anti-pattern is detected:&lt;/STRONG&gt;&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table class="lia-border-color-custom-d6dee6 lia-border-style-solid" border="1" style="width: 100%; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Pattern&lt;/th&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Detection trigger&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;CODE&gt;single-region&lt;/CODE&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;No global LB (Traffic Manager / Front Door) with ≥3 services&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;CODE&gt;single-database&lt;/CODE&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Exactly one database service, no replication signal&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;CODE&gt;no-cache&lt;/CODE&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Compute + database present, no Redis/CDN&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;CODE&gt;no-monitoring&lt;/CODE&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;No Azure Monitor / App Insights / Log Analytics&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;CODE&gt;no-identity&lt;/CODE&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;No Microsoft Entra ID&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;CODE&gt;no-waf&lt;/CODE&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Public web tier without WAF / Front Door / App Gateway&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;CODE&gt;direct-db-access&lt;/CODE&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;An edge from a frontend service directly into a database&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;CODE&gt;no-key-vault&lt;/CODE&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;4+ services and no Key Vault&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;CODE&gt;no-backup&lt;/CODE&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Database present, no Azure Backup / Recovery Services&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;CODE&gt;no-api-gateway&lt;/CODE&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;2+ compute services and no APIM / App Gateway / Front Door&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H4 id="service-specific-rules"&gt;Service-specific rules&lt;/H4&gt;
&lt;P&gt;&lt;STRONG&gt;Every service in the in the generated Azure Architecture diagram is matched against&amp;nbsp;&lt;CODE&gt;SERVICE_SPECIFIC_RULES&lt;/CODE&gt; by normalized type — App Service, Functions, AKS, Cosmos DB, SQL Database, Storage, Key Vault, and 22 more.&lt;/STRONG&gt;&lt;/P&gt;
&lt;H4 id="the-knowledge-base-at-a-glance"&gt;The knowledge base at a glance&lt;/H4&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table class="lia-border-color-custom-d6dee6 lia-border-style-solid" border="1" style="width: 100%; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Metric&lt;/th&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Count&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Total rules&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;STRONG&gt;73&lt;/STRONG&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Architecture-pattern rules&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;STRONG&gt;10&lt;/STRONG&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Service-specific rules&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;STRONG&gt;63&lt;/STRONG&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Distinct Azure services covered&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;STRONG&gt;29&lt;/STRONG&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Rules tagged &lt;EM&gt;Reliability&lt;/EM&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;18&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Rules tagged &lt;EM&gt;Security&lt;/EM&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;34&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Rules tagged &lt;EM&gt;Cost Optimization&lt;/EM&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;5&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Rules tagged &lt;EM&gt;Operational Excellence&lt;/EM&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;7&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Rules tagged &lt;EM&gt;Performance Efficiency&lt;/EM&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;9&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H4 id="the-preliminary-score"&gt;The preliminary score&lt;/H4&gt;
&lt;P&gt;Each finding has a severity, and severity drives a fixed point deduction from a starting score of 100:&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table class="lia-border-color-custom-d6dee6 lia-border-style-solid" border="1" style="width: 100%; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Severity&lt;/th&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Deduction&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;SPAN style="display: inline-block; padding: 2px 8px; border-radius: 10px; background: #FEE2E2; color: #b91c1c; font-weight: 600; font-size: 12px; font-family: 'Segoe UI',Arial,sans-serif;"&gt;critical&lt;/SPAN&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;−12&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;SPAN style="display: inline-block; padding: 2px 8px; border-radius: 10px; background: #FFEDD5; color: #c2410c; font-weight: 600; font-size: 12px; font-family: 'Segoe UI',Arial,sans-serif;"&gt;high&lt;/SPAN&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;−7&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;SPAN style="display: inline-block; padding: 2px 8px; border-radius: 10px; background: #FEF3C7; color: #a16207; font-weight: 600; font-size: 12px; font-family: 'Segoe UI',Arial,sans-serif;"&gt;medium&lt;/SPAN&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;−3&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;SPAN style="display: inline-block; padding: 2px 8px; border-radius: 10px; background: #DCFCE7; color: #15803d; font-weight: 600; font-size: 12px; font-family: 'Segoe UI',Arial,sans-serif;"&gt;low&lt;/SPAN&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;−1&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;Result is floored at 10 (so even a deliberately bad architecture scores at least 10) and ceilinged at 95 (no findings ≠ perfect — there's always something the model might still catch). This is the &lt;STRONG&gt;deterministic baseline&lt;/STRONG&gt; before the LLM ever sees the architecture, and it's what makes the pipeline reproducible.&lt;/P&gt;
&lt;H3 id="phase-2--llm-contextual-refinement"&gt;Phase 2 — LLM contextual refinement&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;The pre-scan output, the topology, and the optional natural-language description are folded into a focused prompt sent to one of seven Azure OpenAI models (GPT-5.1 through 5.4, GPT-5.x Codex variants, DeepSeek V3.2 Speciale, Grok 4.1 Fast). The system prompt gives the model explicit scoring guardrails:&lt;/STRONG&gt;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;UL&gt;
&lt;LI&gt;Score based on what IS present, not what COULD be added.&lt;/LI&gt;
&lt;LI&gt;A well-connected architecture with appropriate services should score 60–80.&lt;/LI&gt;
&lt;LI&gt;Score below 50 only for critical gaps (no auth, no monitoring, single points of failure).&lt;/LI&gt;
&lt;LI&gt;Findings are improvement suggestions, not reasons to penalize the score severely.&lt;/LI&gt;
&lt;/UL&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;The model returns strict JSON:&lt;/P&gt;
&lt;PRE class="jsonc"&gt;&lt;CODE&gt;{
  "overallScore": 0-100,
  "summary": "2–3 sentence assessment",
  "pillars": [
    {
      "pillar": "Reliability | Security | Cost Optimization | Operational Excellence | Performance Efficiency",
      "score": 0-100,
      "findings": [
        {
          "severity": "critical | high | medium | low",
          "category": "...",
          "issue": "...",
          "recommendation": "...",
          "resources": ["service-name-1", "service-name-2"],
          "source": "rule-based | ai-analysis"
        }
      ]
    }
  ],
  "quickWins": [ /* same shape as findings */ ]
}&lt;/CODE&gt;&lt;/PRE&gt;
&lt;P&gt;Two things to call out:&lt;/P&gt;
&lt;OL type="1"&gt;
&lt;LI&gt;&lt;STRONG&gt;Every finding is tagged &lt;CODE&gt;rule-based&lt;/CODE&gt; or &lt;CODE&gt;ai-analysis&lt;/CODE&gt;.&lt;/STRONG&gt; That tag is the credibility lever. You can always see what the deterministic engine produced versus what the model contributed on top. If you don't trust the AI layer, you can ignore it entirely — the rule layer still stands.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;The LLM is given pattern hints, not the entire rule catalog.&lt;/STRONG&gt; The prompt stays small and focused, which is roughly 3–5× faster and cheaper than asking the LLM to do everything from scratch.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H3 id="what-the-user-sees"&gt;What the user sees&lt;/H3&gt;
&lt;P&gt;On every run the modal reports:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Overall WAF score&lt;/STRONG&gt; (0–100)&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Per-pillar score&lt;/STRONG&gt; × 5 (0–100 each)&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Severity breakdown&lt;/STRONG&gt; — counts of critical / high / medium / low across all findings&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Quick wins&lt;/STRONG&gt; — high-impact, low-effort items the model surfaces separately&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Hybrid metadata&lt;/STRONG&gt; — local findings count, patterns detected, KB rules used, preliminary score, local elapsed ms&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;AI metrics&lt;/STRONG&gt; — model used, reasoning effort, prompt/completion/total tokens, elapsed time&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;App Insights telemetry&lt;/STRONG&gt; — an &lt;CODE&gt;Architecture_Validated&lt;/CODE&gt; event with model, overall score, finding count, elapsed time&lt;/LI&gt;
&lt;/UL&gt;
&lt;HR /&gt;
&lt;H2 id="worked-example"&gt;Worked example&lt;/H2&gt;
&lt;P&gt;Take this prompt, which I've used in demos with partners:&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;EM&gt;"A multi-region web application: Azure Front Door in front of two App Service instances in West US 2 and East US 2, both reading from an Azure SQL Database with geo-replication, with Application Insights for telemetry. No Entra ID, no Key Vault."&lt;/EM&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;After generation, &lt;STRONG&gt;Validate Architecture&lt;/STRONG&gt; runs:&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Phase 1 — pre-scan (deterministic), ~1 ms&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Patterns detected: &lt;CODE&gt;no-identity&lt;/CODE&gt;, &lt;CODE&gt;no-key-vault&lt;/CODE&gt;&lt;/LI&gt;
&lt;LI&gt;Findings produced: 8 (1 critical, 1 high, 3 medium, 3 low)&lt;/LI&gt;
&lt;LI&gt;Preliminary score: &lt;STRONG&gt;100 − 12 − 7 − (3×3) − (1×3) = 69&lt;/STRONG&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Phase 2 — LLM refinement, ~6–9 s depending on model&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The model accepts the two pattern hints, validates them in context, and adds three more findings of its own:&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table class="lia-border-color-custom-d6dee6 lia-border-style-solid" border="1" style="width: 100%; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Finding&lt;/th&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Source&lt;/th&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Pillar&lt;/th&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Severity&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;No Microsoft Entra ID for authentication&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;SPAN style="display: inline-block; padding: 2px 8px; border-radius: 10px; background: #DBEAFE; color: #1e3a8a; font-weight: 600; font-size: 12px; font-family: 'Segoe UI',Arial,sans-serif;"&gt;rule-based&lt;/SPAN&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Security&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;SPAN style="display: inline-block; padding: 2px 8px; border-radius: 10px; background: #FEE2E2; color: #b91c1c; font-weight: 600; font-size: 12px; font-family: 'Segoe UI',Arial,sans-serif;"&gt;critical&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;No Key Vault for secret management&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;SPAN style="display: inline-block; padding: 2px 8px; border-radius: 10px; background: #DBEAFE; color: #1e3a8a; font-weight: 600; font-size: 12px; font-family: 'Segoe UI',Arial,sans-serif;"&gt;rule-based&lt;/SPAN&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Security&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;SPAN style="display: inline-block; padding: 2px 8px; border-radius: 10px; background: #FFEDD5; color: #c2410c; font-weight: 600; font-size: 12px; font-family: 'Segoe UI',Arial,sans-serif;"&gt;high&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;App Service slots not used for safe deploys&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;SPAN style="display: inline-block; padding: 2px 8px; border-radius: 10px; background: #F3E8FF; color: #6b21a8; font-weight: 600; font-size: 12px; font-family: 'Segoe UI',Arial,sans-serif;"&gt;ai-analysis&lt;/SPAN&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Operational Excellence&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;SPAN style="display: inline-block; padding: 2px 8px; border-radius: 10px; background: #FEF3C7; color: #a16207; font-weight: 600; font-size: 12px; font-family: 'Segoe UI',Arial,sans-serif;"&gt;medium&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;SQL DB geo-replication present but RTO/RPO not documented&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;SPAN style="display: inline-block; padding: 2px 8px; border-radius: 10px; background: #F3E8FF; color: #6b21a8; font-weight: 600; font-size: 12px; font-family: 'Segoe UI',Arial,sans-serif;"&gt;ai-analysis&lt;/SPAN&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Reliability&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;SPAN style="display: inline-block; padding: 2px 8px; border-radius: 10px; background: #FEF3C7; color: #a16207; font-weight: 600; font-size: 12px; font-family: 'Segoe UI',Arial,sans-serif;"&gt;medium&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;No CDN for static assets behind Front Door&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;SPAN style="display: inline-block; padding: 2px 8px; border-radius: 10px; background: #F3E8FF; color: #6b21a8; font-weight: 600; font-size: 12px; font-family: 'Segoe UI',Arial,sans-serif;"&gt;ai-analysis&lt;/SPAN&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Performance Efficiency&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;SPAN style="display: inline-block; padding: 2px 8px; border-radius: 10px; background: #DCFCE7; color: #15803d; font-weight: 600; font-size: 12px; font-family: 'Segoe UI',Arial,sans-serif;"&gt;low&lt;/SPAN&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;Final scores returned by the model:&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table class="lia-border-color-custom-d6dee6 lia-border-style-solid" border="1" style="width: 100%; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Pillar&lt;/th&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Score&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Reliability&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;78&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Security&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;52&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Cost Optimization&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;80&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Operational Excellence&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;70&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Performance Efficiency&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;75&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;STRONG&gt;Overall&lt;/STRONG&gt;&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;&lt;STRONG&gt;71&lt;/STRONG&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;The Security score is the lowest because two of the highest-severity findings landed there — exactly what a human reviewer would flag first.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="multi-model-comparison"&gt;Multi-model comparison&lt;/H2&gt;
&lt;P&gt;Because the deterministic floor is identical across runs, the &lt;STRONG&gt;Validation Comparison&lt;/STRONG&gt; view becomes a fair shootout of what each LLM adds on top of the same baseline. The same diagram is scored by all seven models, and the UI surfaces:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Overall score per model&lt;/LI&gt;
&lt;LI&gt;Per-pillar score per model&lt;/LI&gt;
&lt;LI&gt;Severity-count deltas&lt;/LI&gt;
&lt;LI&gt;Number of &lt;CODE&gt;ai-analysis&lt;/CODE&gt; findings each model contributed&lt;/LI&gt;
&lt;LI&gt;Quick wins each model identified&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;This is genuinely useful for two reasons. First, it shows that LLM scores vary — typically by ±5–10 points on the same architecture — which is exactly why we publish the &lt;CODE&gt;rule-based&lt;/CODE&gt; vs &lt;CODE&gt;ai-analysis&lt;/CODE&gt; tag. Second, it lets architects pick the model whose review style matches their own.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="how-we-align-with-microsofts-algorithms"&gt;How we align with Microsoft's algorithms&lt;/H2&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table class="lia-border-color-custom-d6dee6 lia-border-style-solid" border="1" style="width: 100%; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Alignment point&lt;/th&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;What it means&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Same five pillars&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Identical names and scope to the official WAF&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Same source material&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Rules derived from WAF docs and Azure Architecture Center service guides&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Severity-graded findings&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Map conceptually to Advisor's high/medium/low impact recommendations&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Per-pillar + overall scoring&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Mirrors WAR/Advisor output shape, so the results feel familiar&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H2 id="where-we-deliberately-differ--and-why"&gt;Where we deliberately differ — and why&lt;/H2&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table class="lia-border-color-custom-d6dee6 lia-border-style-solid" border="1" style="width: 100%; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Concern&lt;/th&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Microsoft&lt;/th&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Diagram Builder&lt;/th&gt;&lt;th class="lia-border-color-custom-0078d4 lia-border-style-solid" style="border-width: 1px; padding: 10px 12px;"&gt;Why we differ&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Needs deployed resources&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Advisor: yes&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;No — works on a diagram&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;We're a &lt;EM&gt;design-time&lt;/EM&gt; tool; the architecture doesn't exist yet&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Needs human Q&amp;amp;A&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;WAR: yes&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;No — derived from the diagram&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;One-click validation inside the design flow&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Healthy/Applicable ratio&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Advisor: yes&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;No&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;No resource-health signal exists pre-deployment&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Subcategory fixed weights&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Advisor: yes&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;No explicit weights&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Severity is the de-facto weight (12/7/3/1)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Defender Secure Score for Security&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Advisor: yes&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;No&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Defender requires deployed resources&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Cost-weighted scoring&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Advisor: yes&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;No (separate Cost Estimation feature)&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Cost is a separate pipeline in our app&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;AI/LLM refinement&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Neither&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Yes&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Catches context-specific issues a static catalog misses, and explains findings in natural language&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Multi-model comparison&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Neither&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Yes&lt;/td&gt;&lt;td class="lia-border-color-custom-d6dee6 lia-vertical-align-top lia-border-style-solid" style="border-width: 1px; padding: 8px 12px;"&gt;Lets architects see scoring variance across models&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;HR /&gt;
&lt;H2 id="honest-limitations"&gt;Honest limitations&lt;/H2&gt;
&lt;P&gt;I'd rather you hear these from me than discover them in production:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;LLM scores drift.&lt;/STRONG&gt; ±5–10 points across models on the same diagram is normal. Treat the score as directional, the findings as actionable. The &lt;CODE&gt;rule-based&lt;/CODE&gt; tag is your anchor.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;No live telemetry.&lt;/STRONG&gt; We can't know if your App Service is actually using availability zones — only that you have App Service in the diagram. Advisor will tell you the truth post-deployment.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Generic ruleset.&lt;/STRONG&gt; No specialized workload branches yet (AI/ML, IoT, SAP, SaaS). WAR has those.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;No milestone tracking.&lt;/STRONG&gt; Each validation run is independent. Compare runs manually using the Validation Comparison view.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Rule coverage is finite.&lt;/STRONG&gt; 29 services and 73 rules is a strong start but not exhaustive — the LLM layer exists in part to compensate for that gap.&lt;/LI&gt;
&lt;/UL&gt;
&lt;HR /&gt;
&lt;H2 id="how-to-use-all-three-together"&gt;How to use all three together&lt;/H2&gt;
&lt;P&gt;&lt;STRONG&gt;A lifecycle that actually works:&lt;/STRONG&gt;&lt;/P&gt;
&lt;OL type="1"&gt;
&lt;LI&gt;&lt;STRONG&gt;Design&lt;/STRONG&gt; — &lt;STRONG&gt;Use the &lt;A href="https://aka.ms/diagram-builder" target="_blank"&gt;Diagram Builder&lt;/A&gt; to sketch the architecture and validate at design time. &lt;/STRONG&gt;Iterate until the per-pillar scores look reasonable and the critical/high findings are addressed.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Deploy&lt;/STRONG&gt; — &lt;STRONG&gt;Generate Bicep from the diagram, deploy, &lt;/STRONG&gt;and let Azure Advisor start scoring real resources.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Operate&lt;/STRONG&gt; — &lt;STRONG&gt;Use Azure Advisor continuously&lt;/STRONG&gt;. Use Defender Secure Score for security posture.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Periodic review&lt;/STRONG&gt; — &lt;STRONG&gt;Run a Core WAR every quarter&lt;/STRONG&gt; or at major milestones to capture the things only humans know (business context, tradeoffs, planned debt).&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;STRONG&gt;None of these three replace the others. They cover different stages of the same loop.&lt;/STRONG&gt;&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="whats-next"&gt;What's next&lt;/H2&gt;
&lt;P&gt;A few things on the roadmap I'd love feedback on:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Milestone tracking&lt;/STRONG&gt; so design-time scores can be compared over time the way WAR milestones work.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Workload-specific rulesets&lt;/STRONG&gt; mirroring WAR's branches — starting with AI/ML.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Direct Advisor handoff&lt;/STRONG&gt; — once a diagram is deployed, surface the corresponding Advisor recommendations in the same UI to close the loop.&lt;/LI&gt;
&lt;/UL&gt;
&lt;HR /&gt;
&lt;H2 id="try-it-fork-it-tell-me-where-its-wrong"&gt;Try it, fork it, tell me where it's wrong&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Live app:&lt;/STRONG&gt; &lt;A href="https://aka.ms/diagram-builder" target="_blank"&gt;https://aka.ms/diagram-builder&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Source:&lt;/STRONG&gt; &lt;A href="https://github.com/Arturo-Quiroga-MSFT/azure-architecture-diagram-builder" target="_blank"&gt;github.com/Arturo-Quiroga-MSFT/azure-architecture-diagram-builder&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Useful references:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/well-architected/pillars" target="_blank"&gt;Azure Well-Architected Framework pillars&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/assessments/azure-architecture-review/" target="_blank"&gt;Azure Well-Architected Review tool&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/advisor/advisor-score#calculation-of-advisor-score" target="_blank"&gt;Azure Advisor score — calculation&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/advisor/advisor-assessments#create-azure-advisor-waf-assessments" target="_blank"&gt;Use Azure WAF assessments (Advisor)&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/well-architected/design-guides/implementing-recommendations" target="_blank"&gt;Complete an Azure Well-Architected Review assessment&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;If you're a partner or customer architect who's already living in Advisor and WAR, I'd genuinely value your reaction — does the design-time stage feel like a real gap to you, or are you already covering it some other way? Open an issue on the repo or reply on LinkedIn.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;EM&gt;Posted on the Azure Architecture Blog · Comments and issues welcome on the &lt;A href="https://github.com/Arturo-Quiroga-MSFT/azure-architecture-diagram-builder" target="_blank"&gt;repo&lt;/A&gt;.&lt;/EM&gt;&lt;/P&gt;</description>
      <pubDate>Thu, 21 May 2026 18:06:15 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/war-azure-advisor-and-us-azure-arch-diagram-builder-three-ways/ba-p/4521611</guid>
      <dc:creator>arturoqu</dc:creator>
      <dc:date>2026-05-21T18:06:15Z</dc:date>
    </item>
    <item>
      <title>From Prompt to Production: Building Azure Architecture Diagrams with AI</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/from-prompt-to-production-building-azure-architecture-diagrams/ba-p/4520336</link>
      <description>&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Author:&lt;/STRONG&gt; Arturo Quiroga, Senior Partner Solutions Architect — Microsoft&lt;/P&gt;
&lt;HR /&gt;
&lt;P&gt;Cloud architects spend significant time translating ideas into architecture diagrams. They toggle between Visio, draw.io, pricing calculators, and documentation. According to the &lt;A href="https://survey.stackoverflow.co/2024/professional-developers#1-daily-time-spent-searching-for-answers-solutions" target="_blank" rel="noopener"&gt;2024 Stack Overflow Developer Survey&lt;/A&gt;, 61% of developers spend more than 30 minutes a day searching for answers or solutions, time lost to context-switching rather than design. What if you could describe your architecture in plain English and get a diagram, cost estimate, and deployment guide in minutes?&lt;/P&gt;
&lt;H2 id="the-challenge-fragmented-architecture-workflows"&gt;The Challenge: Fragmented Architecture Workflows&lt;/H2&gt;
&lt;P&gt;Designing Azure architectures today typically involves multiple disconnected steps:&lt;/P&gt;
&lt;OL type="1"&gt;
&lt;LI&gt;&lt;STRONG&gt;Sketch&lt;/STRONG&gt; the architecture in a diagramming tool&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Look up&lt;/STRONG&gt; official Azure icons and drag them into place&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Research&lt;/STRONG&gt; pricing across regions using the Azure Pricing Calculator&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Validate&lt;/STRONG&gt; the design against the Well-Architected Framework (WAF)&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Write&lt;/STRONG&gt; deployment documentation and Infrastructure as Code templates&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Compare&lt;/STRONG&gt; alternative designs manually&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Each step lives in a different tool, and keeping them in sync as designs evolve is costly. The Azure Architecture Diagram Builder brings these workflows together in a single browser-based experience.&lt;/P&gt;
&lt;H2 id="how-it-works"&gt;How It Works&lt;/H2&gt;
&lt;P&gt;Describe your architecture in natural language, for example &lt;EM&gt;"A HIPAA-compliant healthcare platform with FHIR APIs, event-driven processing, and multi-region disaster recovery"&lt;/EM&gt;, and the AI generates a diagram with grouped services, data flow connections, and logical organization.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 1.&lt;/STRONG&gt; Enter a natural-language prompt describing your architecture. Curated example prompts help you get started, and you can optionally upload an existing diagram for the AI to analyze.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;The tool uses &lt;STRONG&gt;Azure OpenAI&lt;/STRONG&gt; to power generation across multiple models, enabling you to choose the model that best fits your scenario — from fast iterations to deeper reasoning.&lt;/P&gt;
&lt;H2 id="key-features"&gt;Key Features&lt;/H2&gt;
&lt;H3 id="ai-powered-architecture-generation"&gt;AI-Powered Architecture Generation&lt;/H3&gt;
&lt;P&gt;Describe what you need in plain English, and the AI creates an architecture diagram with:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;714 official Azure service icons&lt;/STRONG&gt; across 29 categories&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Smart grouping&lt;/STRONG&gt;: services are logically organized (Frontend, Backend, Data, Security)&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Data flow connections&lt;/STRONG&gt;: labeled edges showing how data moves through the system&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;13 curated example prompts&lt;/STRONG&gt;: from simple web apps to complex enterprise scenarios like Zero Trust networks, Industrial IoT with 5,000+ sensors, and global multiplayer gaming backends&lt;/LI&gt;
&lt;/UL&gt;
&lt;img /&gt;&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 2.&lt;/STRONG&gt; A generated industrial IoT architecture. &lt;STRONG&gt;Top:&lt;/STRONG&gt; the clean diagram view as initially produced. &lt;STRONG&gt;Bottom:&lt;/STRONG&gt; the same diagram with per-service monthly cost overlays toggled on, plus a running subscription total in the toolbar.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3 id="architecture-image-import"&gt;Architecture Image Import&lt;/H3&gt;
&lt;P&gt;Already have an architecture on a whiteboard or in a screenshot? Upload the image and let the AI analyze it, mapping services to official Azure icons and recreating the architecture as an editable, interactive diagram.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 3.&lt;/STRONG&gt; Upload a photo of a whiteboard sketch (top-right reference panel) and the AI recreates it as an editable diagram with official Azure service icons and labeled data flow connections.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3 id="arm-template-import"&gt;ARM Template Import&lt;/H3&gt;
&lt;P&gt;Import existing ARM templates to visualize your current infrastructure. The AI parses resource definitions and dependencies, groups related resources into logical layers, and produces a meaningful diagram of what you actually have deployed — a fast way to document an inherited environment or sanity-check a template before deployment.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 4.&lt;/STRONG&gt; ARM template import in action. &lt;STRONG&gt;Top:&lt;/STRONG&gt; the parser status banner while resources and dependencies are being analyzed. &lt;STRONG&gt;Bottom:&lt;/STRONG&gt; the resulting diagram, with resources auto-grouped into logical layers (Web Tier, Data Layer, Container Platform, Observability &amp;amp; Logging) and a &lt;EM&gt;Generated from: ARM Template&lt;/EM&gt; badge linking the diagram back to its source file.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3 id="well-architected-framework-validation"&gt;Well-Architected Framework Validation&lt;/H3&gt;
&lt;P&gt;Validate your architecture against all five WAF pillars — Security, Reliability, Performance Efficiency, Cost Optimization, and Operational Excellence. The validator provides:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;An overall WAF score with pillar-level breakdowns&lt;/LI&gt;
&lt;LI&gt;Specific findings with severity levels&lt;/LI&gt;
&lt;LI&gt;Actionable recommendations you can select and apply&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Select the recommendations you agree with, and the AI regenerates an improved architecture incorporating those changes.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 5.&lt;/STRONG&gt; WAF validation results showing the overall score, per-pillar breakdowns, and individual findings with severity badges. Tick the recommendations you want and the AI rebuilds the diagram with those changes applied.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3 id="multi-model-comparison"&gt;Multi-Model Comparison&lt;/H3&gt;
&lt;P&gt;Run the same architecture prompt through multiple AI models side-by-side and compare:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Architecture Comparison&lt;/STRONG&gt;: service counts, connection counts, groups, token usage, and latency&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Validation Comparison&lt;/STRONG&gt;: WAF scores across models, severity breakdowns, and finding counts&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Apply Winner&lt;/STRONG&gt;: pick the best result and apply it to the canvas with one click&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Present Critique&lt;/STRONG&gt;: a talking avatar narrates the AI-generated ranking with live closed captions&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 6.&lt;/STRONG&gt; Multi-model comparison. &lt;STRONG&gt;Top:&lt;/STRONG&gt; select the models and reasoning effort, then enter the prompt. &lt;STRONG&gt;Bottom:&lt;/STRONG&gt; side-by-side results across all selected models with service counts, latency, token usage, and &lt;EM&gt;Fastest / Cheapest / Most Thorough&lt;/EM&gt; badges.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3 id="multi-region-cost-estimation"&gt;Multi-Region Cost Estimation&lt;/H3&gt;
&lt;P&gt;Get cost estimates from the Azure Retail Prices API across &lt;STRONG&gt;8 Azure regions&lt;/STRONG&gt;: East US 2, Australia East, Canada Central, Brazil South, Mexico Central, West Europe, Sweden Central, and Southeast Asia. Features include:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Color-coded cost legend (green / yellow / red thresholds)&lt;/LI&gt;
&lt;LI&gt;SKU and tier information for each service&lt;/LI&gt;
&lt;LI&gt;Export options: CSV, JSON, plain-text summary, and an analysis report with top cost drivers, Reserved Instance flags, and a ranked multi-region comparison table&lt;/LI&gt;
&lt;/UL&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 7.&lt;/STRONG&gt; The cost legend overlay shows per-service pricing with color-coded thresholds. The region selector in the toolbar lets you re-price the entire architecture in any of eight Azure regions.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3 id="deployment-guide-generation-with-bicep"&gt;Deployment Guide Generation with Bicep&lt;/H3&gt;
&lt;P&gt;Generate step-by-step deployment documentation including:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Prerequisites and Azure resource requirements&lt;/LI&gt;
&lt;LI&gt;Step-by-step deployment instructions&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Bicep templates&lt;/STRONG&gt; for each service (Infrastructure as Code)&lt;/LI&gt;
&lt;LI&gt;Post-deployment verification steps&lt;/LI&gt;
&lt;LI&gt;Security configuration recommendations&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 8.&lt;/STRONG&gt; Each generated Deployment Guide opens with the architecture name, an estimated deployment time, and a prerequisites checklist covering subscription roles, CLI versions, Microsoft Entra ID permissions, and region requirements, followed by numbered, copy-ready deployment steps.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 9.&lt;/STRONG&gt; The Infrastructure as Code section produces a &lt;CODE&gt;main.bicep&lt;/CODE&gt; orchestrator plus a per-service module (Log Analytics, Key Vault, Cosmos DB, SQL Database, Event Hubs, Azure Functions, and more). The &lt;STRONG&gt;Download All Templates&lt;/STRONG&gt; button packages everything into a ready-to-deploy folder.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3 id="workflow-animation--avatar-presenter"&gt;Workflow Animation &amp;amp; Avatar Presenter&lt;/H3&gt;
&lt;P&gt;Visualize how data flows through your architecture with step-by-step animations that highlight services on the canvas as each step plays. When the Azure Speech Service is configured, a photorealistic talking avatar can narrate the workflow or present model comparison results, with live word-by-word closed captions in a draggable, resizable panel.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 10.&lt;/STRONG&gt; A workflow step is highlighted on the canvas as the Avatar Presenter narrates that step. Live word-by-word closed captions appear in a draggable, resizable panel, useful for accessibility and stakeholder demos.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3 id="export-options"&gt;Export Options&lt;/H3&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 11.&lt;/STRONG&gt; A single-slide PowerPoint export, available in dark or light theme, ready to drop straight into a stakeholder deck.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Format&lt;/th&gt;&lt;th&gt;Use Case&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;PNG&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Documentation, presentations&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;SVG&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Scalable vector graphics&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;PPTX&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Single PowerPoint slide (dark or light theme)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Draw.io&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Edit in diagrams.net&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;JSON&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Backup, version control&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;CSV / ZIP&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Cost analysis with multi-region comparison&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H2 id="highlights"&gt;Highlights&lt;/H2&gt;
&lt;P&gt;The Azure Architecture Diagram Builder unifies the architecture design lifecycle in a single tool:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;End-to-end workflow&lt;/STRONG&gt;: from natural-language description to deployable Bicep templates without tool switching&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Official Azure icons&lt;/STRONG&gt;: 714 icons across 29 categories, mapped directly from the Azure service catalog&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Live pricing&lt;/STRONG&gt;: queries the Azure Retail Prices API at design time rather than relying on static estimates&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;WAF-integrated validation&lt;/STRONG&gt;: architectural best practices built into the design loop rather than applied after the fact&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Multi-model flexibility&lt;/STRONG&gt;: choose the AI model that best suits each task, with fast models for iteration and reasoning models for complex designs&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Open source&lt;/STRONG&gt;: the source code is available for customization and contribution&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2 id="one-command-deploy-with-azure-developer-cli"&gt;One-Command Deploy with Azure Developer CLI&lt;/H2&gt;
&lt;P&gt;The fastest way to get your own instance running is with &lt;A href="https://aka.ms/azd" target="_blank" rel="noopener"&gt;&lt;CODE&gt;azd&lt;/CODE&gt;&lt;/A&gt;:&lt;/P&gt;
&lt;DIV id="cb1" class="sourceCode"&gt;
&lt;PRE class="sourceCode bash"&gt;&lt;CODE class="sourceCode bash"&gt;&lt;SPAN id="cb1-1"&gt;&lt;SPAN class="co"&gt;# Install azd (once)&lt;/SPAN&gt;&lt;/SPAN&gt;
&lt;SPAN id="cb1-2"&gt;&lt;SPAN class="ex"&gt;brew&lt;/SPAN&gt; tap azure/azd &lt;SPAN class="kw"&gt;&amp;amp;&amp;amp;&lt;/SPAN&gt; &lt;SPAN class="ex"&gt;brew&lt;/SPAN&gt; install azd   &lt;SPAN class="co"&gt;# macOS&lt;/SPAN&gt;&lt;/SPAN&gt;
&lt;SPAN id="cb1-3"&gt;&lt;SPAN class="ex"&gt;winget&lt;/SPAN&gt; install microsoft.azd             &lt;SPAN class="co"&gt;# Windows&lt;/SPAN&gt;&lt;/SPAN&gt;
&lt;SPAN id="cb1-4"&gt;&lt;/SPAN&gt;
&lt;SPAN id="cb1-5"&gt;&lt;SPAN class="co"&gt;# Clone, configure, and deploy&lt;/SPAN&gt;&lt;/SPAN&gt;
&lt;SPAN id="cb1-6"&gt;&lt;SPAN class="fu"&gt;git&lt;/SPAN&gt; clone https://github.com/Arturo-Quiroga-MSFT/azure-architecture-diagram-builder&lt;/SPAN&gt;
&lt;SPAN id="cb1-7"&gt;&lt;SPAN class="bu"&gt;cd&lt;/SPAN&gt; azure-architecture-diagram-builder&lt;/SPAN&gt;
&lt;SPAN id="cb1-8"&gt;&lt;SPAN class="ex"&gt;azd&lt;/SPAN&gt; auth login&lt;/SPAN&gt;
&lt;SPAN id="cb1-9"&gt;&lt;SPAN class="ex"&gt;azd&lt;/SPAN&gt; env set AZURE_OPENAI_ENDPOINT &lt;SPAN class="st"&gt;"https://your-resource.openai.azure.com/"&lt;/SPAN&gt;&lt;/SPAN&gt;
&lt;SPAN id="cb1-10"&gt;&lt;SPAN class="ex"&gt;azd&lt;/SPAN&gt; env set AZURE_OPENAI_API_KEY  &lt;SPAN class="st"&gt;"your-key"&lt;/SPAN&gt;&lt;/SPAN&gt;
&lt;SPAN id="cb1-11"&gt;&lt;SPAN class="ex"&gt;azd&lt;/SPAN&gt; up   &lt;SPAN class="co"&gt;# Provisions infrastructure + builds + deploys (~8 min)&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/CODE&gt;&lt;/PRE&gt;
&lt;/DIV&gt;
&lt;P&gt;&lt;CODE&gt;azd up&lt;/CODE&gt; provisions the following via Bicep:&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Resource&lt;/th&gt;&lt;th&gt;Purpose&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Azure Container Registry&lt;/td&gt;&lt;td&gt;Stores the Docker image&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Azure Container Apps&lt;/td&gt;&lt;td&gt;Runs the app (nginx + token server)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Log Analytics + Application Insights&lt;/td&gt;&lt;td&gt;Monitoring and telemetry&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Azure Speech (S0)&lt;/td&gt;&lt;td&gt;Avatar Presenter (optional, keyless auth via managed identity)&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H2 id="try-it-today"&gt;Try It Today&lt;/H2&gt;
&lt;P&gt;The Azure Architecture Diagram Builder is available now:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Live demo&lt;/STRONG&gt;: &lt;A href="https://aka.ms/diagram-builder" target="_blank" rel="noopener"&gt;https://aka.ms/diagram-builder&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Source code&lt;/STRONG&gt;: &lt;A href="https://github.com/Arturo-Quiroga-MSFT/azure-architecture-diagram-builder" target="_blank" rel="noopener"&gt;GitHub repository&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Documentation&lt;/STRONG&gt;: See the &lt;A href="https://github.com/Arturo-Quiroga-MSFT/azure-architecture-diagram-builder/blob/main/DOCS/getting-started-guide.md" target="_blank" rel="noopener"&gt;Getting Started Guide&lt;/A&gt; for detailed setup instructions&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;We welcome feedback and contributions. Use the GitHub Issues page to report bugs, suggest features, or share your experience.&lt;/P&gt;
&lt;HR /&gt;
&lt;P&gt;&lt;STRONG&gt;Tags:&lt;/STRONG&gt; &lt;CODE&gt;artificial intelligence&lt;/CODE&gt; · &lt;CODE&gt;application&lt;/CODE&gt; · &lt;CODE&gt;apps &amp;amp; devops&lt;/CODE&gt; · &lt;CODE&gt;well architected&lt;/CODE&gt; · &lt;CODE&gt;infrastructure&lt;/CODE&gt;&lt;/P&gt;</description>
      <pubDate>Fri, 22 May 2026 18:35:07 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/from-prompt-to-production-building-azure-architecture-diagrams/ba-p/4520336</guid>
      <dc:creator>arturoqu</dc:creator>
      <dc:date>2026-05-22T18:35:07Z</dc:date>
    </item>
  </channel>
</rss>

