<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>rss.livelink.threads-in-node</title>
    <link>https://techcommunity.microsoft.com/t5/azure/ct-p/Azure</link>
    <description>rss.livelink.threads-in-node</description>
    <pubDate>Sat, 25 Jul 2026 12:56:38 GMT</pubDate>
    <dc:creator>Azure</dc:creator>
    <dc:date>2026-07-25T12:56:38Z</dc:date>
    <item>
      <title>Announcing the Open-Source Release of ML Video Codec (MLVC)</title>
      <link>https://techcommunity.microsoft.com/t5/linux-and-open-source-blog/announcing-the-open-source-release-of-ml-video-codec-mlvc/ba-p/4539875</link>
      <description>&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Video codecs compress video for transmission or storage, reducing bandwidth and storage requirements. MLVC is a modern machine-learning-based codec that uses substantially less bandwidth than conventional codecs, improving streaming and video-call quality—especially on constrained or unreliable networks—while lowering delivery and storage costs.&lt;/P&gt;
&lt;P&gt;MLVC is the product iteration of &lt;A class="lia-external-url" href="https://github.com/microsoft/DCVC" target="_blank" rel="noopener"&gt;DCVC (Deep Contextual Video Compression) family of NVC (Neural Video Codec),&lt;/A&gt; open sourced by Microsoft Research since 2021, with improved compression efficiency, real-time performance on commodity Neural Processing Units (NPUs), and cross-platform support. We recently published this work in the paper &lt;A class="lia-external-url" href="https://arxiv.org/abs/2606.28027" target="_blank" rel="noopener"&gt;MLVC: Multi-platform Learned Video Codec for Real-World Deployment&lt;/A&gt;. We are releasing the source code because we believe the next generation of video coding will be built openly, and we want the broader community — researchers, video codec engineers, platform vendors, product teams, as well as general developer community — to build it with us.&lt;/P&gt;
&lt;H2&gt;Why MLVC&lt;/H2&gt;
&lt;P&gt;Traditional video codecs (e.g., H.264/AVC, H.265/HEVC) have served the industry for a long time, but each generation requires enormous engineering effort for incremental gains and needs dedicated hardware which takes years to become commonly available. MLVC replaces conventional primitives — motion estimation, transforms, entropy modeling — with &lt;STRONG&gt;end-to-end learned neural compression&lt;/STRONG&gt;, trained directly against &lt;A class="lia-external-url" href="https://en.wikipedia.org/wiki/Rate%E2%80%93distortion_theory" target="_blank" rel="noopener"&gt;rate-distortion&lt;/A&gt; objectives, and run on general-purpose NPU devices.&amp;nbsp;&lt;/P&gt;
&lt;P&gt;The table below compares MLVC to popular video codecs, showing its lower bitrate and resulting savings in bandwidth and storage.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Resolution&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;vs H.264&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;vs H.265&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;360p&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;87.8%&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;75.5%&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;540p&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;82.7%&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;65.4%&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;For example, for 360p video at 30 fps, where H.264 requires 1 Mbps, MLVC requires roughly 122 kbps for equivalent quality — about one-eight the bitrate under real-time conditions. The inference compute was kept approximately equal for the 360p and 540p resolutions. These results are based on a &lt;A class="lia-external-url" href="https://www.itu.int/rec/t-rec-p.910/en" target="_blank" rel="noopener"&gt;P.910 subjective test&lt;/A&gt; and are based on the &lt;A class="lia-external-url" href="https://github.com/microsoft/VCD" target="_blank" rel="noopener"&gt;Video Conferencing Dataset (VCD) dataset&lt;/A&gt; that we developed and also recently released as an open-source project.&lt;/P&gt;
&lt;P&gt;The following video demo illustrates the extent of quality enhancement achieved by MLVC relative to H.265/HEVC at the same bitrate of 200kbps (please watch by &lt;A class="lia-external-url" href="https://aka.ms/mct/media/azLinOss/0724261" target="_blank" rel="noopener"&gt;opening in a new window&lt;/A&gt; for better demonstration)&lt;/P&gt;
&lt;FIGURE style="margin: 0; padding: 0;"&gt;
&lt;DIV style="position: relative; width: 100%; height: 0; padding-bottom: 28%; overflow: hidden; border: 0;"&gt;&lt;IFRAME src="https://aka.ms/mct/media/azLinOss/0724261" title="MLVC demo video" allowfullscreen="allowfullscreen" allow="fullscreen; picture-in-picture" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border: 0;" sandbox="allow-scripts allow-same-origin allow-forms"&gt;
&lt;/IFRAME&gt;\&lt;/DIV&gt;
&lt;/FIGURE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Beyond video compression efficiency, MLVC also offers:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;NPU-first design.&lt;/STRONG&gt; MLVC is built to run almost entirely on the AI accelerators already shipping in modern devices — Apple Neural Engine, Qualcomm and Intel NPUs — at no more than 50% NPU utilization, leaving NPU headroom for the rest of the system.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Real-time&lt;/STRONG&gt;&lt;STRONG&gt; execution at the targeted operating points.&lt;/STRONG&gt; Demonstrated&lt;STRONG&gt;&amp;nbsp;&lt;/STRONG&gt;540p at 30 fps on Apple, Intel, and Qualcomm hardware.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;A scaling-law trajectory.&lt;/STRONG&gt; Empirically, MLVC's coding efficiency improves with increased model capacity and additional training compute.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Content-adaptive behavior out of the box,&lt;/STRONG&gt; without hand-tuned &lt;A class="lia-external-url" href="https://en.wikipedia.org/wiki/Rate%E2%80%93distortion_optimization" target="_blank" rel="noopener"&gt;Rate Distortion Optimization&lt;/A&gt; heuristics.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Already running in Microsoft Teams&lt;/H2&gt;
&lt;P&gt;MLVC is more than just a research concept for Microsoft. We are currently rolling it out in&lt;STRONG&gt; &lt;/STRONG&gt;Microsoft Teams, where it is being validated on real peer-to-peer video calls with active telemetry and A/B testing. The integration runs alongside fallback to conventional video codecs for hardware or reliability constraints — the kind of mixed-environment deployment that real products need. The scaling and reliability insights from this rollout are shaping the codec and its roadmap.&lt;/P&gt;
&lt;H2&gt;We welcome your contributions&lt;/H2&gt;
&lt;P&gt;We can not cover every use case, every device class, or every content domain by ourselves. That is why we are open sourcing MLVC. If you work on:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Streaming, Video On Demand (VOD), or live broadcast&lt;/LI&gt;
&lt;LI&gt;Real-time communication and conferencing&lt;/LI&gt;
&lt;LI&gt;Cloud gaming or remote rendering&lt;/LI&gt;
&lt;LI&gt;Surveillance, drones, or robotics&lt;/LI&gt;
&lt;LI&gt;Mobile capture, Augmented Reality (AR) / Virtual Reality (VR), or volumetric video&lt;/LI&gt;
&lt;LI&gt;Codec hardware, NPUs, or inference runtimes&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;We would love your help. Contributions of all kinds are welcome, including model improvements, training recipes, new platform ports and conversion targets, runtime backends, domain-specific fine-tunes, evaluation tooling, bug reports, and feedback on what is missing for your scenario. We particularly welcome &lt;STRONG&gt;platform ports&lt;/STRONG&gt; that expand NPU coverage and &lt;STRONG&gt;efficiency improvements&lt;/STRONG&gt; that push the rate-distortion frontier.&lt;/P&gt;
&lt;H2&gt;What's in the release&lt;/H2&gt;
&lt;P&gt;The MLVC repository is available at &lt;A class="lia-external-url" href="https://github.com/microsoft/mlvc" target="_blank" rel="noopener"&gt;https://github.com/microsoft/mlvc&lt;/A&gt; and shared under the &lt;STRONG&gt;MIT License&lt;/STRONG&gt;. It includes:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;MLVC model source code&lt;/STRONG&gt; of the full network architecture.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Trained model weights&lt;/STRONG&gt; ready to run.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Training scripts&lt;/STRONG&gt; used to produce shipping models.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Training data&lt;/STRONG&gt;&lt;STRONG&gt; collection documentation&lt;/STRONG&gt; to help reproduce and improve MLVC.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Platform conversion scripts&lt;/STRONG&gt; to target different NPUs and runtimes.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Issues and pull requests will be open from day one.&lt;/P&gt;
&lt;P&gt;A follow-up release will add a C++ codec library, simplifying integration into real-world applications.&lt;/P&gt;
&lt;H2&gt;Where we're heading&lt;/H2&gt;
&lt;P&gt;Our long-term goal with MLVC is to create an open, learned video codec that meets or exceeds the coding efficiency of the best conventional codecs across the full range of video content, runs efficiently on the AI hardware already shipping in client and cloud devices, scales with compute the way modern ML systems do, and evolves in the open at the pace of the ML community rather than the pace of standardization cycles.&lt;/P&gt;
&lt;P&gt;In the near term that means stabilizing 540p real-time performance, expanding hardware coverage, and improving loss resilience. In the medium term: higher resolution, e.g., 1080p, and broader streaming scenarios. In the long term: an open video codec ecosystem that meaningfully replaces legacy stacks where it makes sense to.&lt;/P&gt;
&lt;P&gt;We can't create the future of MLVC alone. We're glad you're here, and we are looking forward to building the next-generation video codec with you.&lt;/P&gt;
&lt;H2&gt;Who are we&lt;/H2&gt;
&lt;P&gt;MLVC is brought to you by the following awesome folks working on the project at Microsoft: &lt;EM&gt;Ross Cutler, Ando Saabas, Tanel Pärnamaa, Ardi Loot, Haiyan Xie, Lauri Ehrenpreis, Andrei Znobishchev, Martin Lumiste, Evgenii Indenbom, Yan Lu, Bin Li, Jiahao Li, Naba Kumar, Babak Naderi, Juhee Cho, Badal Yadav, Jinxin Zhou, Tianyu Ding, Patrick Gregory.&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;— The MLVC Team, Microsoft&lt;/P&gt;</description>
      <pubDate>Fri, 24 Jul 2026 21:39:05 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/linux-and-open-source-blog/announcing-the-open-source-release-of-ml-video-codec-mlvc/ba-p/4539875</guid>
      <dc:creator>Naba_Kumar</dc:creator>
      <dc:date>2026-07-24T21:39:05Z</dc:date>
    </item>
    <item>
      <title>Your Entire Agentic AI Workflow, Now Inside VS Code: New Course Available</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-developer-community/your-entire-agentic-ai-workflow-now-inside-vs-code-new-course/ba-p/4539009</link>
      <description>&lt;P&gt;If you build with AI, you know the tax: jumping between a browser portal to pick a model, a terminal to run it, a separate playground to test a prompt, and finally your editor to write the code. Every context switch is a small drain on focus—and it adds up.&lt;/P&gt;
&lt;P&gt;What if the whole loop lived in the one tool you never leave?&lt;/P&gt;
&lt;P&gt;That's the promise of the&amp;nbsp;&lt;STRONG&gt;Foundry Toolkit for Visual Studio Code&lt;/STRONG&gt;, and it's exactly what our new&amp;nbsp;&lt;STRONG&gt;VS Code Learn: Foundry Toolkit&lt;/STRONG&gt;&amp;nbsp;video series is here to show you. Whether you're deploying your first model or wiring up your first agent, the series walks you through doing it all without leaving your editor.&lt;/P&gt;
&lt;P&gt;👉&amp;nbsp;&lt;STRONG&gt;Watch the full playlist:&lt;/STRONG&gt;&amp;nbsp;&lt;A href="https://aka.ms/Learn/Foundry" target="_blank" rel="noopener"&gt;aka.ms/Learn/Foundry&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;What is the Foundry Toolkit?&lt;/H2&gt;
&lt;P&gt;The &lt;A class="lia-external-url" href="https://aka.ms/FoundryToolkit" target="_blank" rel="noopener"&gt;Foundry Toolkit is a Visual Studio Code&lt;/A&gt; extension that brings Microsoft Foundry directly into your development environment. Instead of hopping between the portal and your editor, you get a single, integrated surface to:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Sign in to Microsoft Foundry&lt;/STRONG&gt;&amp;nbsp;and connect your workspace without leaving VS Code.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Browse the model catalog&lt;/STRONG&gt;&amp;nbsp;and work with the models that fit your scenario.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Build, test, and iterate on AI agents&lt;/STRONG&gt;&amp;nbsp;as a natural part of your coding workflow.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Explore the toolkit UI&lt;/STRONG&gt;&amp;nbsp;designed to keep discovery, experimentation, and code in one place.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;In short: it turns your editor into the command center for your agentic AI work.&lt;/P&gt;
&lt;H2&gt;What you'll learn in the series&lt;/H2&gt;
&lt;P&gt;The series is built to take you from "just installed it" to "shipping with it," one focused episode at a time. Across the playlist you'll learn how to:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Get set up&lt;/STRONG&gt;&amp;nbsp;— install the extension, sign in to Microsoft Foundry, and get oriented in the UI.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Work with models&lt;/STRONG&gt;&amp;nbsp;— navigate the toolkit to find and use the right model for your task.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Develop AI agents&lt;/STRONG&gt;&amp;nbsp;— go from an idea to a working agent, all inside VS Code.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Fit AI into your existing flow&lt;/STRONG&gt;&amp;nbsp;— pair the toolkit with GitHub Copilot and the tools you already use every day.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Each video is short, practical, and hands-on—perfect for watching alongside your own editor so you can follow along in real time.&lt;/P&gt;
&lt;H2&gt;Featured episode:&amp;nbsp;&lt;EM&gt;Your Entire Agentic AI Workflow Now Inside VS Code&lt;/EM&gt;&lt;/H2&gt;
&lt;P&gt;Not sure where to jump in? Start here. In this episode we run through&amp;nbsp;&lt;STRONG&gt;installation, signing in to Microsoft Foundry, and a tour of the toolkit UI&lt;/STRONG&gt;—the essential first steps that get you productive fast.&lt;/P&gt;
&lt;P&gt;▶️&amp;nbsp;&lt;STRONG&gt;Watch:&lt;/STRONG&gt;&amp;nbsp;&lt;A href="https://www.youtube.com/watch?v=aQFSDGAk9DA" target="_blank" rel="noopener"&gt;Your Entire Agentic AI Workflow Now Inside VS Code&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;By the end you'll have the extension installed, be connected to Foundry, and know your way around the interface—ready to move on to models and agents.&lt;/P&gt;
&lt;div data-video-id="https://www.youtube.com/watch?v=aQFSDGAk9DA&amp;amp;list=PLJfWOmd-Usr4/1784560456748" data-video-remote-vid="https://www.youtube.com/watch?v=aQFSDGAk9DA&amp;amp;list=PLJfWOmd-Usr4/1784560456748" class="lia-video-container lia-media-is-center lia-media-size-large"&gt;&lt;iframe src="https://cdn.embedly.com/widgets/media.html?src=https%3A%2F%2Fwww.youtube.com%2Fembed%2FaQFSDGAk9DA%3Flist%3DPLJfWOmd-Usr4&amp;amp;display_name=YouTube&amp;amp;url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3DaQFSDGAk9DA&amp;amp;image=https%3A%2F%2Fi.ytimg.com%2Fvi%2FaQFSDGAk9DA%2Fhqdefault.jpg&amp;amp;type=text%2Fhtml&amp;amp;schema=youtube" allowfullscreen="" style="max-width: 100%"&gt;&lt;/iframe&gt;&lt;/div&gt;
&lt;H2&gt;Prefer to read and follow along? There's a hands-on course too&lt;/H2&gt;
&lt;P&gt;The video series has a companion: a brand-new, step-by-step&amp;nbsp;&lt;STRONG&gt;written course on the Foundry Toolkit&lt;/STRONG&gt;&amp;nbsp;on the VS Code docs, authored by Chris Noring. It's the perfect complement to the videos—read the lesson, try it in your editor, then watch the walkthrough to reinforce it (or the other way around).&lt;/P&gt;
&lt;P&gt;📚&amp;nbsp;&lt;STRONG&gt;Read the course:&lt;/STRONG&gt; &lt;A class="lia-external-url" href="https://aka.ms/foundry_toolkit_vscode_learn" target="_blank"&gt;https://aka.ms/foundry_toolkit_vscode_learn&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;Together, the videos and the course give you two ways to learn the same material—watch it, then do it.&lt;/P&gt;
&lt;H2&gt;Who is this for?&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Developers new to AI&lt;/STRONG&gt;&amp;nbsp;who want a guided, low-friction on-ramp into building with models and agents.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Experienced builders&lt;/STRONG&gt;&amp;nbsp;who want to collapse their model-to-code loop into a single tool.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Teams standardizing on VS Code&lt;/STRONG&gt;&amp;nbsp;who want their AI workflow to live where their code already does.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;No deep machine-learning background required—if you're comfortable in VS Code, you're ready.&lt;/P&gt;
&lt;H2&gt;Get started today&lt;/H2&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;Watch&lt;/STRONG&gt;&amp;nbsp;the featured episode to get set up:&amp;nbsp;&lt;A href="https://www.youtube.com/watch?v=aQFSDGAk9DA" target="_blank" rel="noopener"&gt;Your Entire Agentic AI Workflow Now Inside VS Code&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Work through&lt;/STRONG&gt;&amp;nbsp;the full series:&amp;nbsp;&lt;A href="https://aka.ms/Learn/Foundry" target="_blank" rel="noopener"&gt;aka.ms/Learn/Foundry&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Follow the written course&lt;/STRONG&gt; on VS Code docs to build alongside each lesson:&amp;nbsp;
&lt;P&gt;&lt;A class="lia-external-url" href="https://aka.ms/foundry_toolkit_vscode_learn" target="_blank"&gt;aka.ms/foundry_toolkit_vscode_learn&lt;/A&gt;&lt;/P&gt;
&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Bring your models, your agents, and your code into one place. Your editor is about to get a whole lot more capable.&lt;/P&gt;</description>
      <pubDate>Fri, 24 Jul 2026 09:20:31 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-developer-community/your-entire-agentic-ai-workflow-now-inside-vs-code-new-course/ba-p/4539009</guid>
      <dc:creator>carlottacaste</dc:creator>
      <dc:date>2026-07-24T09:20:31Z</dc:date>
    </item>
    <item>
      <title>Changing the engine while the plane is flying: migrating 60,000 apps under live load</title>
      <link>https://techcommunity.microsoft.com/t5/azure-integration-services-blog/changing-the-engine-while-the-plane-is-flying-migrating-60-000/ba-p/4539443</link>
      <description>&lt;H2&gt;Summary&lt;/H2&gt;
&lt;P&gt;Azure Logic Apps Consumption Integration Accounts were served by roughly 60,000 per-customer Azure Functions apps on the end-of-life v1/v2 runtime. Over about two years we migrated the entire fleet to Functions v4 (isolated worker), across all Azure regions, with no migration work required from customers. Here is the whole approach before the deep dive:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Shadow traffic.&lt;/STRONG&gt;&amp;nbsp;Run real production traffic through the new runtime in parallel with the old one and compare every result, without ever letting the unproven path serve the customer (Sections 5 and 6).&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;A parity bar we could defend.&lt;/STRONG&gt;&amp;nbsp;Compare 100% of eligible traffic, separate genuine bugs from naturally nondeterministic workloads, and fix every real divergence before moving anyone (Section 6).&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;A progressive, reversible rollout.&lt;/STRONG&gt;&amp;nbsp;Shift traffic by a hash-based rate, region by region, with a configuration-only rollback that takes effect in minutes (Section 7).&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Careful retirement.&lt;/STRONG&gt;&amp;nbsp;Stop the old apps first, observe, and only then delete, with a long reversible gap (Section 9).&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Why this was possible without customer involvement: we own and operate the compute, the durable customer data (agreements, maps, schemas) never moved, and the actions are pure transforms, so we could run both runtimes on the same traffic and compare them behind our own control plane. The result was an aggregate cutover with no material latency or error regression, with one honest caveat: a small subset of customers hit a brief edge case under high load before we rolled back (Section 10).&lt;/P&gt;
&lt;H2&gt;1. The hard part wasn't the new runtime, it was proving the old one could be replaced&lt;/H2&gt;
&lt;P&gt;Azure Logic Apps Consumption Integration Accounts power critical behind-the-scenes work in enterprise integration: schema and map storage, XML and flat-file transforms, partner and agreement resolution, and the connectors that lean on them. That work is live, multi-tenant, and always on. Customers don't schedule downtime for it, and neither can we.&lt;/P&gt;
&lt;P&gt;Under the hood, those capabilities were served by Azure Functions apps running on runtime v1/v2. Those apps executed the Integration Account actions (transforms, lookups) while the durable data itself (agreements, maps, schemas) lived in storage they didn't own. Those runtime versions have reached the end of their supported life, and staying on them was its own growing risk. Moving to a supported runtime was part of a broader modernization effort: it restored a healthy security- and dependency-update path immediately, and it opened the door to newer .NET down the road. The destination was clear: Azure Functions v4.&lt;/P&gt;
&lt;P&gt;Here's the catch that made this more than a version bump. We couldn't prove compatibility from a test environment. The behavior we had to preserve was the behavior real customer workloads depended on, in production, across many regions. A staging suite gives you confidence; it doesn't give you proof.&lt;/P&gt;
&lt;P&gt;So we built a migration path around a simple idea. Let the old runtime keep serving the live path while the new runtime processes the same real traffic in parallel, and don't move anyone until the data says it's safe. This post is how we did that: shadow traffic, a parity bar we could defend, a progressive rollout with a fast way back, and a guardrailed path to finally retire the old apps.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 1. Before and after.&lt;/STRONG&gt;&amp;nbsp;A new v4 app is provisioned per classic app (one-to-one) and reads the same durable data, which never moves. Customers stay on classic until parity is proven.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;2. Why Functions v4, and why we couldn't just flip a switch&lt;/H2&gt;
&lt;P&gt;The move to v4 is also a move to the isolated worker model. On v1/v2, our code ran in-process with the Functions host. On v4 it runs in an isolated worker process, decoupled from the host. That decoupling is exactly what a modern, supportable runtime looks like, but it's also a breaking change in how the code is hosted, which is the crux of why this couldn't be a quick switch.&lt;/P&gt;
&lt;P&gt;We considered the usual lighter-weight options and ruled each out for this workload:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;In-place upgrade.&lt;/STRONG&gt;&amp;nbsp;Going from v1/v2 to the v4 isolated worker isn't an in-place runtime bump; it's a host-model change. There's no safe "upgrade the live app and hope" here.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Deployment slots or blue-green alone.&lt;/STRONG&gt;&amp;nbsp;Slots solve routing, not proof. Even with staged or split routing, moving customer traffic across a fleet of roughly 60,000 per-customer apps before we'd proven the new app matched the old one on real requests would still be a blind behavioral bet.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Maintenance windows.&lt;/STRONG&gt;&amp;nbsp;Integration Accounts are live customer infrastructure. There's no acceptable downtime budget for this surface.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Which points at the thesis of the whole project: the approach we chose is more complex for us, but simpler for the customer. We absorbed the complexity so customers wouldn't have to. This was only possible because of where the work sat: we own and operate these Functions apps, the durable customer data never moves, and the actions are pure transforms, so the whole migration could happen behind our control plane without customers changing, redeploying, or even knowing which runtime served a given call.&lt;/P&gt;
&lt;H2&gt;3. The shape of the problem&lt;/H2&gt;
&lt;P&gt;A few properties made this harder than a typical app move:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Per-account compute at fleet scale.&lt;/STRONG&gt;&amp;nbsp;Integration Account capabilities were backed by Functions apps at a per-account granularity, so "migrate the app" really meant "migrate a fleet of roughly 60,000 apps" across all Azure regions.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;A small number of shared platform workloads.&lt;/STRONG&gt;&amp;nbsp;Alongside the per-account apps, a few shared platform workloads served cross-cutting functionality. There's only one set of them per region, so they had a different blast radius and needed their own, more cautious targeting (Section 7).&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Stateless compute.&lt;/STRONG&gt;&amp;nbsp;The durable Integration Account data (agreements, maps, schemas) lives in storage and did not move; those stateful assets stayed exactly where they were. We were migrating compute, not customer data, which kept the parity surface smaller than it first appears.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;A runtime decaying under time pressure.&lt;/STRONG&gt;&amp;nbsp;One risk was easy to underrate: the longer we stayed on the deprecated v1/v2 platform, with no security-patch path and no dependency updates, the more its reliability could erode underneath us. "Do nothing" wasn't a safe holding pattern.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The bar we set ourselves was to keep customer-visible disruption as close to invisible as we could make it, and to prove we were hitting that bar with data rather than assert it.&lt;/P&gt;
&lt;H2&gt;4. The confidence ladder&lt;/H2&gt;
&lt;P&gt;Instead of a cutover, we built a progression. Each rung had to hold before we climbed to the next:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;Provision alongside.&lt;/STRONG&gt;&amp;nbsp;Stand up a new v4 app next to the classic app; change nothing customers touch.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Validate parity.&lt;/STRONG&gt;&amp;nbsp;Make the v4 app's behavior equivalent to classic (Section 5).&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Shadow real traffic.&lt;/STRONG&gt;&amp;nbsp;Send real traffic to v4 in parallel while classic still serves the workflow; compare results (Section 6).&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Analyze mismatches.&lt;/STRONG&gt;&amp;nbsp;Triage every divergence, fix forward, and raise the bar until any residual divergence is understood and explainable.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Rate-gated cutover.&lt;/STRONG&gt;&amp;nbsp;Move a controlled, hash-based fraction of traffic to v4, region by region (Section 7).&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Stay rollback-ready.&lt;/STRONG&gt;&amp;nbsp;Keep a fast, fine-grained way back at every step (Section 7).&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Stop, then delete, carefully.&lt;/STRONG&gt;&amp;nbsp;Retire classic only after guardrails prove it's safe (Section 9).&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 2. The confidence ladder.&lt;/STRONG&gt;&amp;nbsp;Each rung had to hold before we climbed. Rollback stayed available throughout, and retirement (stop first, delete much later) came last.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;This ladder is the spine of everything below.&lt;/P&gt;
&lt;H2&gt;5. Parity foundation: what had to be equivalent&lt;/H2&gt;
&lt;P&gt;Because the durable Integration Account data stayed in storage, parity wasn't about copying state. It was about one question: does the v4 code package produce the same observable output as the classic one, for the same input? Every action an app could perform had to behave equivalently between the in-process host (v1/v2) and the isolated worker (v4).&lt;/P&gt;
&lt;P&gt;The subtle part is that the host-model change can shift behavior you never explicitly wrote. The clearest example is JSON serialization. The in-process and isolated worlds start from different serializer defaults, so the same object can render differently: timestamp formats, enum representations, and the like. For customer output to stay reliably equivalent, byte-for-byte where customers observed exact serialized values, those defaults couldn't be left to chance. (The concrete bug this caused, and how shadow caught it, is in Section 6.)&lt;/P&gt;
&lt;P&gt;So "validate parity" meant pinning the things that change observable behavior (runtime and dependency versions, host and app configuration, and serialization settings) and then letting real traffic be the judge of whether we'd actually achieved it.&lt;/P&gt;
&lt;H2&gt;6. Shadow mode: the showpiece&lt;/H2&gt;
&lt;P&gt;This is the heart of the migration. For an app in shadow mode, every eligible call still went to the classic runtime, and the workflow received the classic result. In parallel, we sent the same call to the v4 app and compared the two responses, emitting a match/mismatch signal. The workflow was never given the v4 result while we were still proving it. A v4 shadow error or slow response stayed off the customer path entirely and could only ever influence our telemetry and gating, never the result the workflow received.&lt;/P&gt;
&lt;P&gt;For each shadowed request we compared the status code, the action or outcome code, and the output payload, the signals that define a correct result. We also captured request timings, but as a health and regression signal, not as part of the match/mismatch verdict.&lt;/P&gt;
&lt;P&gt;Why was this safe to do at all? The obvious objection to running every request twice is side effects: wouldn't the parallel call double real-world actions? For this workload, no. Integration Account actions are essentially pure transforms, meaning transforms and metadata lookups that read the durable Integration Account data and produce an output, without writing to any external customer system. The shadow call computes a result and we compare it; nothing is being written anywhere, so there was nothing to double. That property is what made full-traffic shadowing safe by construction.&lt;/P&gt;
&lt;P&gt;Because it was safe, we didn't sample. When an app was in shadow mode we compared 100% of its live traffic, every eligible request. "Eligible" simply excluded apps opted out at the resource level (for example, unsupported special-case apps on an exclusion list); everything else was compared.&lt;/P&gt;
&lt;P&gt;We started strict, requiring a full JSON match, then relaxed to a normalized comparison so that formatting noise wasn't miscounted as a real difference. But some customer workloads are legitimately nondeterministic: they embed random values or timestamps, so classic itself won't return an identical response twice. To separate "v4 is wrong" from "this operation is inherently nondeterministic," we added a deterministic-match mode. For a shadowed request, we also called classic a second time. The logic is simple:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;If v4 disagrees with classic's first response and classic's second response also disagrees with its first, the operation is nondeterministic, and the v4 "mismatch" is discounted.&lt;/LI&gt;
&lt;LI&gt;If v4 disagrees but the two classic calls agree with each other, that's a real v4 divergence to fix.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Crucially, that extra classic call ran off the customer path, and only on the handful of apps already showing mismatches, so it added no customer-visible latency and was never an always-on tax. Shadowing itself (classic plus one v4 call) covered 100% of eligible traffic; the extra classic call existed only for those already-mismatching apps.&lt;/P&gt;
&lt;P&gt;When the two disagreed, the migration team triaged the divergence. Cases sorted, informally, into a few recognizable buckets: serialization and format differences, customer nondeterminism, transient infrastructure blips, and genuine v4 logic bugs. We didn't formalize that into a tracked taxonomy. We also watched the asymmetric cases, where classic succeeds but the v4 shadow call errors, or the reverse. We read these as ordinary infrastructure noise rather than v4 defects because of their shape: they stayed at very low rates, correlated in time with known regional blips, and showed no app-specific concentration, so they didn't materially gate an app's readiness.&lt;/P&gt;
&lt;P&gt;There was no single magic number to advance, but the judgment wasn't loose either. Before an app's traffic could move we wanted a sustained, near-100% normalized match with no unexplained systemic divergence, its asymmetric failures reviewed and understood, and stable v4 error and latency signals through a bake period. Within those guardrails the final call was operational judgment, and we also weighed known quirks like cold starts occasionally causing a failure or a retry alongside the overall match rate.&lt;/P&gt;
&lt;P&gt;The best evidence the method earned its cost is a real bug it surfaced before any customer was affected: the serializer-defaults divergence from Section 5. In-process v1 and isolated v4 serialized certain values differently out of the box, so v4's output didn't match classic's for things like timestamps and enums. Shadow flagged the mismatch; the fix was to explicitly pin v4's JSON serialization options to match the classic runtime, making customer output identical again. Without shadow, that's exactly the kind of quiet difference that ships and generates confused support tickets weeks later.&lt;/P&gt;
&lt;P&gt;Running both worlds isn't free. Shadowing roughly doubles compute and the read-only dependency calls that go with it. We absorbed that by standing up the v4 apps on their own dedicated App Service Plans, isolating their capacity so the shared downstream services stayed healthy. Relative to each region's overall load the added traffic wasn't a large delta, and because it ran on isolated, monitored capacity behind the same switches, we could cut it instantly if it ever strained anything downstream. Per app, shadow typically ran for weeks before cutover, and longer in the early stages while we built confidence that the pattern held across the fleet.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 3. Runtime routing.&lt;/STRONG&gt;&amp;nbsp;Solid lines are the customer path; dashed lines run off it. In shadow, the workflow always gets the classic result while v4 runs in parallel and only emits a match/mismatch signal.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;7. Controlling blast radius, rolling out, and watching it land&lt;/H2&gt;
&lt;P&gt;Parity earns the right to move traffic; control is what lets you move it safely.&lt;/P&gt;
&lt;P&gt;Traffic moved by a hash-based exposure rate keyed on a stable resource or tenant identifier. The hash is deterministic: for a stable identifier, the same account maps to the same bucket, so an app stays put on classic or v4 and never flip-flops between them as the rate ramps or as we roll region to region. That stability matters, because it prevents apps from oscillating across deployments.&lt;/P&gt;
&lt;P&gt;Layered control came from four switches: a master switch, a mode selector (classic, shadow, or v4-only), the percentage-based rate, and resource-level exclusion controls. The exclusion controls were the emergency brake: name an affected resource and it snaps back to classic, with no deployment required, effective within a few minutes.&lt;/P&gt;
&lt;P&gt;We advanced through safe-deployment rings, and we deliberately started where a mistake would hurt least: early canary and test regions, then internal-to-Microsoft accounts, and only then the broader regions. Advancing a ring was a monitor-and-judge decision against rich telemetry. We checked the prior ring's rates, confirmed the target region was healthy, let earlier regions bake, and watched how they behaved under customer-traffic spikes before moving on.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Stage / ring&lt;/th&gt;&lt;th&gt;Advance gate&lt;/th&gt;&lt;th&gt;Signals watched&lt;/th&gt;&lt;th&gt;Halt / rollback action&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Canary + test regions&lt;/td&gt;&lt;td&gt;Sustained near-100% normalized match; no open systemic mismatch; stable v4 error/latency through bake&lt;/td&gt;&lt;td&gt;Match rate, v4 error rate, latency, customer-reported incidents&lt;/td&gt;&lt;td&gt;Exclude resource(s) to classic in minutes&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Internal-to-Microsoft accounts&lt;/td&gt;&lt;td&gt;Prior ring's gates still green at higher scenario variety; asymmetric failures reviewed and understood&lt;/td&gt;&lt;td&gt;same&lt;/td&gt;&lt;td&gt;Hold ring or roll back via config&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Broad regions (safe-deployment rings)&lt;/td&gt;&lt;td&gt;Prior-ring gates held, target region healthy, bake time elapsed&lt;/td&gt;&lt;td&gt;same plus behavior under traffic spikes&lt;/td&gt;&lt;td&gt;Exclusion controls or lower the rate&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;When a signal turned, whether a dip in match rate, a rise in v4-only errors, a latency regression, or a customer-reported incident, we could hold the ring or roll back. The rollback is just a configuration change that takes effect in a few minutes, and it wasn't theoretical: we exercised it for real several times (most notably for a high-load edge case, and on some internal apps), which means it was battle-tested under real load, not just on paper.&lt;/P&gt;
&lt;P&gt;The migration team owned the rollout day to day, driving ring advancement, watching the dashboards, and making the advance, halt, and rollback calls, with standard on-call and incident processes layered on top, so a customer-reported incident paged the on-call engineer through the normal channels.&lt;/P&gt;
&lt;P&gt;None of this is safe if you can't see it. We invested heavily in dashboards, metrics, and counters, and fed mismatch and health signals (non-reversible telemetry, not payload logging) into health and mismatch dashboards. The loop was simple and strict: signal, triage, gate decision. A batch advanced only when the telemetry said it had earned it.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Figure 4. The observability loop.&lt;/STRONG&gt;&amp;nbsp;Signal, triage, then gate decision: a batch advanced only when telemetry (parity and v4 health) said it had earned it; otherwise it held, rolled back, or was fixed forward.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;8. Keeping shadow traffic safe&lt;/H2&gt;
&lt;P&gt;Shadowing real customer traffic is a privacy responsibility, not just an engineering technique, so we designed the comparison to need no customer data at rest. The comparison happened in memory, in the moment: we never logged PII or customer payloads. What we emitted were signals, not data, meaning match/mismatch outcomes and metrics designed not to reveal or reconstruct the underlying content. Because comparisons were transient and no payload was ever logged or retained, the design avoided creating a new customer-data-at-rest surface, so there's no store of sensitive traces to govern. The platform's standard runtime access controls and operational safeguards still applied to the live path itself. And as covered in Section 6, because the actions are pure transforms, the shadow path produced no duplicate customer-visible side effects.&lt;/P&gt;
&lt;P&gt;That privacy choice came with a real tradeoff. Because we never captured payloads, a handful of mismatches simply couldn't be reproduced from the signals alone. For those, we reached out to the affected customers directly and, with their help, collected example maps and inputs that let us reproduce the divergence in a controlled setting, root-cause it, and fix it. Trading a little investigative convenience for keeping customer data out of our telemetry was the right call, and customers were glad to help us close real bugs.&lt;/P&gt;
&lt;H2&gt;9. Retiring the old world: stop, then delete, with guardrails&lt;/H2&gt;
&lt;P&gt;A migration isn't done when traffic moves; it's done when the old system is safely gone. We deliberately separated stopping from deleting, with a long, reversible gap between them.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Confirm, then disable.&lt;/STRONG&gt;&amp;nbsp;After an app cut over, we first confirmed via telemetry that no traffic was still reaching the classic app, then disabled it through a per-region configuration. Disabled but not deleted means quickly restorable if anything surfaced.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Observe.&lt;/STRONG&gt;&amp;nbsp;We waited again after disabling, watching for anything to break, before considering deletion at all.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Delete last, and only in bulk confidence.&lt;/STRONG&gt;&amp;nbsp;We only began deleting classic apps after most apps across most regions had been disabled for a sustained period. Deletion is the final, irreversible stage, gated on a long, successful disabled window rather than done eagerly per app.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Retirement safety, in one line: traffic fully moved, no remaining callers, disabled (reversible), observation window elapsed, and only then delete.&lt;/P&gt;
&lt;H2&gt;10. Results, lessons, and what's next&lt;/H2&gt;
&lt;P&gt;By the numbers (rounded):&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Scale.&lt;/STRONG&gt;&amp;nbsp;Roughly 60,000 Functions apps migrated, across all Azure regions.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Shadow volume.&lt;/STRONG&gt;&amp;nbsp;At peak, the fleet compared on the order of hundreds of millions of requests per week, with 100% of eligible traffic compared, not a sample.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Parity arc.&lt;/STRONG&gt;&amp;nbsp;The match rate climbed in step-changes, each driven by a big, systemic fix: serialization pinning chief among them, alongside alignments in host and DI configuration and dependency versions. After those landed, it was mostly smooth, with only minor or nondeterministic residue.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Cutover impact (fleet level).&lt;/STRONG&gt;&amp;nbsp;Cutover was intentionally uneventful in aggregate: aggregate latency and error-rate telemetry showed no material regression at the moment of switch, because shadow had already proven parity. That's a fleet-level signal, not a promise of zero individual impact; see the honest caveat below.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Timeline.&lt;/STRONG&gt;&amp;nbsp;A little over two years: roughly year one achieving and confirming parity via shadow, and year two moving real traffic and then retiring the old apps.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;An honest caveat, because "zero impact" would be too strong. Safe deployment and shadowing caught the overwhelming majority of issues before customers saw them, and most customers had to do nothing. But it wasn't perfectly invisible: a small subset of customers hit an edge case under high load and experienced brief, real impact before we rolled back and coordinated a minor adjustment. Rollback did its job; we used it several times early on, then rarely once the big fixes were in. The incident trend followed the same shape: a few incidents up front while the systemic fixes landed, then quiet.&lt;/P&gt;
&lt;P&gt;Lessons worth keeping:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;A parity bar is only credible if it accounts for nondeterminism.&lt;/STRONG&gt;&amp;nbsp;The "call classic twice" trick was what let us tell a real v4 bug apart from a workload that simply never repeats itself, and enabling it selectively kept it from costing customers latency.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;The host-model change hides in serialization.&lt;/STRONG&gt;&amp;nbsp;The most important fix wasn't in business logic; it was pinning serializer settings so v4's output matched the classic runtime exactly.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;No side effects, no duplication risk.&lt;/STRONG&gt;&amp;nbsp;Because the actions only read data and compute, with nothing written to any external system, we could shadow all traffic safely. That single property is the biggest reason this approach was even available to us.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Signals beat payloads.&lt;/STRONG&gt;&amp;nbsp;Comparing in memory and emitting only non-reversible signals gave us the confidence of full-traffic comparison without the liability of storing customer data.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Separate "stop" from "delete."&lt;/STRONG&gt;&amp;nbsp;The long, reversible disabled window turned the scariest, irreversible step into a calm one.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;What's next: customer traffic is on v4, and we have already deleted the large majority of the classic v1 apps. What remains is the final cleanup, removing the last disabled classic apps in the trailing regions and finishing the more cautious retirement of the shared platform workloads. The migration itself is done; the remaining deletion is deliberately unhurried.&lt;/P&gt;
&lt;P&gt;The broader takeaway is that shadow traffic plus a progressive, reversible rollout is a repeatable pattern for replacing a runtime under live, multi-tenant load. Customers still carry real risk when the infrastructure under them changes; what this pattern shifts is control. The team that owns the migration absorbs the complexity and holds the levers to prove parity, gate the rollout, and reverse it, so customers don't have to manage the change or its mitigation.&lt;/P&gt;</description>
      <pubDate>Thu, 23 Jul 2026 20:05:25 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-integration-services-blog/changing-the-engine-while-the-plane-is-flying-migrating-60-000/ba-p/4539443</guid>
      <dc:creator>AbodeSaafan</dc:creator>
      <dc:date>2026-07-23T20:05:25Z</dc:date>
    </item>
    <item>
      <title>Introducing Kubernetes-Native Policy Validation with CEL and VAP in Azure Policy</title>
      <link>https://techcommunity.microsoft.com/t5/azure-governance-and-management/introducing-kubernetes-native-policy-validation-with-cel-and-vap/ba-p/4534585</link>
      <description>&lt;P&gt;Azure Policy for Kubernetes now supports&amp;nbsp;&lt;A href="https://open-policy-agent.github.io/gatekeeper/website/docs/validating-admission-policy/" target="_blank" rel="noopener"&gt;Gatekeeper’s integration&lt;/A&gt; with native Kubernetes Validating Admission Policy (VAP) using Common Expression Language (CEL)! This integration leverages the new VAP feature introduced in Kubernetes 1.30, providing a more efficient, reliable and in-process way to enforce policies.&lt;/P&gt;
&lt;P&gt;This integration brings lower latency for admission decisions, simpler constraint template authoring—for the first time – allows you to enable stronger fail-close behavior to prevent non-compliant components getting created or updated in your environment . And you keep all the Azure Policy benefits you already rely on like centralized assignment, scope management, safe rollout control, and governance at scale.&lt;/P&gt;
&lt;P&gt;Previously, Azure Policy for Kubernetes only supported on OPA Rego-based evaluation through an admission webhook flow:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;A request is sent to the Kubernetes API server.&lt;/LI&gt;
&lt;LI&gt;The request is forwarded to Gatekeeper.&lt;/LI&gt;
&lt;LI&gt;Gatekeeper enforces resources using constraint template written in Rego.&lt;/LI&gt;
&lt;LI&gt;The result is returned to the API server.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2&gt;Native Kubernetes Validation with Common Expression Language (CEL)&lt;/H2&gt;
&lt;P&gt;Policy expressions can be written in CEL syntax, a lightweight and expressive language that Kubernetes designed specifically for this validation context. Here's what changes:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt; In-tree evaluation&lt;/STRONG&gt;: Policy rules run inside the Kubernetes API server itself, not in an external webhook service. This means:&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Lower latency&lt;/STRONG&gt;: No round-trip delay to an external service. Admission decisions are made faster because evaluation happens in-process.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Better reliability&lt;/STRONG&gt;: Validation doesn't depend on a separate service remaining healthy. If the VAP controller finds a violating resource, the PUT request will be denied, resulting in a fail-close scenario.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Getting Azure Policy Benefits with CEL&lt;/P&gt;
&lt;P&gt;Kubernetes-native enforcement does not replace Azure Policy governance capabilities. It strengthens them. Azure Policy becomes the governance and deployment layer, while CEL becomes the validation logic. Here's what that means:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;CEL handles the validation&lt;/STRONG&gt;: You write lightweight, readable expressions that run natively in Kubernetes. CEL is simpler than Rego and aligns with Kubernetes' own validation framework.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure Policy handles the governance&lt;/STRONG&gt;: Assignments, versioning, compliance tracking, safe rollouts, overrides, and multi-cluster management all happen through Azure Policy. You get enterprise governance without extra operational overhead.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Your First CEL Policy with Azure Policy&lt;/H2&gt;
&lt;P&gt;Here's the key insight:&amp;nbsp;&lt;STRONG&gt;you don't write Azure Policy definitions in CEL&lt;/STRONG&gt;. Instead, you write a CEL constraint template, package it inside an Azure Policy definition, and deploy it through Azure Policy.&lt;/P&gt;
&lt;P&gt;Think of it as two layers:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Layer 1: CEL Constraint Template&lt;/STRONG&gt; — The Kubernetes validation logic you write&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Layer 2: Azure Policy Definition&lt;/STRONG&gt; — The wrapper that adds assignments, versioning, compliance tracking, and safe deployment&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;Real Example: Limit Deployment Replicas&lt;/H3&gt;
&lt;P&gt;Let's walk through a concrete scenario. Your platform team wants to prevent developers from creating Deployments with more than 5 replicas—a common guard rail for dev/test environments to control resource consumption.&lt;/P&gt;
&lt;P&gt;- replicas: 1 through replicas: 5 — allowed&lt;/P&gt;
&lt;P&gt;- replicas: 6 or higher — denied&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Step 1: Write the CEL constraint template&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;You write a Gatekeeper ConstraintTemplate using the K8sNativeValidation engine. The CEL expression itself is a single, readable line:&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;LI-CODE lang="yaml"&gt;apiVersion: templates.gatekeeper.sh/v1beta1
kind: ConstraintTemplate
metadata:
  name: k8se2etestcelmaxdeployreplicas
spec:
  crd:
    spec:
      names:
        kind: K8sE2ETestCelMaxDeployReplicas
      validation:
       # Schema for the `parameters` field
        openAPIV3Schema:
          type: object
          properties:
            message:
              type: string
  targets:
    - target: admission.k8s.gatekeeper.sh
      code:
      - engine: K8sNativeValidation
        source:
          validations:
          - expression: "object.spec.replicas &amp;lt;= 5"
            message: "message example"&lt;/LI-CODE&gt;
&lt;P&gt;The key part is the engine: K8sNativeValidation block. This tells Gatekeeper to generate a Kubernetes ValidatingAdmissionPolicy resource from this template, so evaluation runs inside the API server rather than through the Gatekeeper webhook. The CEL expression object.spec.replicas &amp;lt;= 5 is evaluated natively at admission time.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Step 2: Package the CEL template into an Azure Policy definition&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;You then will need to base64-encode the constraint template (you can use an online tool or a local tool to do this) and embed the encoded content in an Azure Policy custom definition using sourceType: Base64Encoded. The full custom policy definition structure, matching the real format Azure Policy expects:&lt;/P&gt;
&lt;LI-CODE lang="json"&gt;{
  "policyType": "Custom",
  "mode": "Microsoft.Kubernetes.Data",
  "displayName": "Ensure deployment replicas are less than or equal to 5 in Kubernetes cluster",
  "policyRule": {
    "if": {
      "field": "type",
      "in": [
        "Microsoft.ContainerService/managedClusters"
      ]
    },
    "then": {
      "effect": "[parameters('effect')]",
      "details": {
        "templateInfo": {
          "sourceType": "Base64Encoded",
          "content": "&amp;lt;BASE64_ENCODED_CONSTRAINT_TEMPLATE&amp;gt;"
        },
        "apiGroups": [
          "apps"
        ],
        "kinds": [
          "Deployment"
        ],
        "namespaces": "[parameters('namespaces')]",
        "excludedNamespaces": "[parameters('excludedNamespaces')]",
        "labelSelector": "[parameters('labelSelector')]",
        "values": {
          "message": "[parameters('message')]"
        }
      }
    }
  },
  "parameters": {
    "effect": {
      "type": "String",
      "allowedValues": [
        "audit",
        "deny",
        "disabled"
      ],
      "defaultValue": "audit"
    },
    "namespaces": {
      "type": "Array",
      "defaultValue": []
    },
    "labelSelector": {
      "type": "Object",
      "defaultValue": {}
    },
    "message": {
      "type": "String"
    }
  }
}&lt;/LI-CODE&gt;
&lt;P&gt;A few things worth noting in this structure:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;sourceType: Base64Encoded - Azure Policy stores your constraint template as base64 content inline in the policy definition. When Azure Policy deploys to your cluster, it decodes and installs the template automatically.&lt;/LI&gt;
&lt;LI&gt;apiGroups: ["apps"] and kinds: ["Deployment"] - These scope the policy to only evaluate Kubernetes Deployment resources in the apps API group. Pods, Services, and other resources are unaffected.&lt;/LI&gt;
&lt;LI&gt;excludedNamespaces - System namespaces like kube-system and gatekeeper-system are excluded by default, so the policy doesn't interfere with cluster infrastructure.&lt;/LI&gt;
&lt;LI&gt;effect: audit (default) - The policy starts in audit mode, flagging violations without blocking them. Switch to deny when you're ready to enforce.&lt;/LI&gt;
&lt;/UL&gt;
&lt;LI-SPOILER label="Important"&gt;
&lt;P&gt;Right now only audit, deny, and disabled Policy effects are supported.&lt;/P&gt;
&lt;/LI-SPOILER&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;You can create the policy definition via the Azure portal, Azure CLI, or the Azure Policy VS Code extension.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Step 3: Assign the policy to your AKS clusters&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Now you can assign this policy through Azure Policy to your clusters:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;In the Azure portal, go to Azure Policy &amp;gt; Definitions.&lt;/LI&gt;
&lt;LI&gt;Find your policy by name (e.g., "Ensure deployment replicas are less than or equal to 5").&lt;/LI&gt;
&lt;LI&gt;Click Assign.&lt;/LI&gt;
&lt;LI&gt;Set the Scope to the management group, subscription, or resource group containing your AKS clusters.&lt;/LI&gt;
&lt;LI&gt;Set the effect parameter to audit to start, then graduate to deny once you've validated the policy behavior.&lt;/LI&gt;
&lt;LI&gt;Click Create.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Azure Policy deploys the CEL-based constraint template to all clusters in scope. Within a few minutes, the Azure Policy add-on decodes the template, installs it on each cluster, and Gatekeeper generates the corresponding ValidatingAdmissionPolicy resource. From that point on, the Kubernetes API server evaluates every new Deployment admission request against your CEL rule directly. You can target different clusters or namespaces with different assignments and apply overrides for exceptions.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Step 4: Monitor compliance through Azure Policy&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;In the Azure Policy console, you see real-time compliance across all your clusters:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Non-compliant Deployments (those with replicas &amp;gt; 5) appear in the Compliance view.&lt;/LI&gt;
&lt;LI&gt;You can drill down by policy, cluster, or namespace to see which Deployments are violating the limit.&lt;/LI&gt;
&lt;LI&gt;While in audit mode, existing violations are surfaced without blocking anything—giving you a clear picture before switching to deny.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Availability and prerequisites&lt;/H2&gt;
&lt;OL&gt;
&lt;LI&gt;Supported Kubernetes version: AKS clusters running Kubernetes v1.30 or later (VAP became generally available in K8s v1.30).&lt;/LI&gt;
&lt;LI&gt;Gatekeeper version: Azure Policy add-on includes Gatekeeper v3.17.0 or later, which supports VAP and CEL policy generation.&lt;/LI&gt;
&lt;LI&gt;Azure Policy Add-on: The Azure Policy add-on for Kubernetes must be installed on your clusters. See &lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/governance/policy/concepts/policy-for-kubernetes" target="_blank" rel="noopener"&gt;Azure Policy for Kubernetes documentation&lt;/A&gt;for installation instructions.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;This is currently generally available and can be used to enforce policies in your Kubernetes clusters today.&lt;/P&gt;
&lt;H2&gt;Roadmap&lt;/H2&gt;
&lt;P&gt;We are actively working to include more built in Azure Policies for Azure Kubernetes service with the constraint template built with CEL.&lt;/P&gt;
&lt;H2&gt;Resources&lt;/H2&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/governance/policy/concepts/policy-for-kubernetes" target="_blank" rel="noopener"&gt;Azure Policy for Kubernetes&lt;/A&gt;: Full documentation on how Azure Policy integrates with Kubernetes, including built-in definitions and assignment workflows.&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://kubernetes.io/docs/reference/using-api/cel/" target="_blank" rel="noopener"&gt;Common Expression Language in Kubernetes&lt;/A&gt;: Reference guide for authoring CEL expressions used in Validating Admission Policies.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2&gt;Conclusion&lt;/H2&gt;
&lt;P&gt;Azure Policy for Kubernetes now lets you bring Kubernetes-native CEL validation into your enterprise governance workflow.&lt;/P&gt;
&lt;P&gt;The mental model is simple: &lt;STRONG&gt;CEL handles what to validate, Azure Policy handles how to govern it.&lt;/STRONG&gt; You write clean, readable CEL expressions that run natively in your Kubernetes API server. Azure Policy wraps those expressions and handles the hard parts of enterprise governance—assignments across management hierarchies, versioning, compliance tracking, safe rollouts, and overrides.&lt;/P&gt;
&lt;P&gt;Whether you're enforcing container registries, requiring security contexts, or validating resource configurations, you now have a path to do it natively in Kubernetes with enterprise governance backing it up.&lt;/P&gt;</description>
      <pubDate>Thu, 23 Jul 2026 19:27:12 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-governance-and-management/introducing-kubernetes-native-policy-validation-with-cel-and-vap/ba-p/4534585</guid>
      <dc:creator>stevenbucher</dc:creator>
      <dc:date>2026-07-23T19:27:12Z</dc:date>
    </item>
    <item>
      <title>Announcing public preview of Azure DDoS Protection custom policy</title>
      <link>https://techcommunity.microsoft.com/t5/azure-networking-blog/announcing-public-preview-of-azure-ddos-protection-custom-policy/ba-p/4538963</link>
      <description>&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;We are excited to announce the public preview of Azure DDoS Protection custom policy, a new capability that gives customers more granular control over how Azure DDoS Protection detects and mitigates attacks against protected workloads.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Azure DDoS Protection has always focused on delivering automatic, adaptive protection at scale. With custom policy, customers can now fine-tune mitigation behaviour for supported resources and configure protocol-specific detection thresholds. This allows customers to better align protection settings with any planned or projected changes in their application traffic patterns upfront, while giving them additional control over supported protocol thresholds.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;EM&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-props="{}"&gt;Azure DDoS Protection custom policy is currently in preview. See the&amp;nbsp;&lt;A href="vscode-file://vscode-app/c:/Users/ofirsarfaty/AppData/Local/Programs/Microsoft%20VS%20Code%20Insiders/5e212606d5/resources/app/out/vs/sessions/electron-browser/sessions.html" target="_blank" rel="noopener" data-href="https://azure.microsoft.com/en-us/support/legal/preview-supplemental-terms/"&gt;Supplemental Terms of Use for Microsoft Azure Previews&lt;/A&gt; for legal terms that apply to Azure features that are in beta, preview, or otherwise not yet generally available.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/EM&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-props="{}"&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;Why customers asked for more control&lt;/SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Azure DDoS Protection automatically analyses traffic patterns and applies adaptive mitigation during attacks. While this approach works well for most workloads, some organizations require additional flexibility to support unique traffic characteristics, operational environments, or changes around big releases and events.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Customers running latency-sensitive applications, high-throughput services, gaming platforms, or workloads with predictable traffic spikes often want the ability to customize mitigation trigger behavior for specific protocols.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Example scenarios include:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Applications with known peak traffic periods&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="5" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Workloads that experience sustained high packet-per-second rates&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="5" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="2" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Services requiring protocol-specific tuning&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="9" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Organizations that need separate mitigation policies across environments or applications&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Custom policy allows organizations to align DDoS protection behaviour with their operational traffic baselines while still leveraging Azure’s global-scale mitigation infrastructure.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H3&gt;&lt;SPAN data-ccp-props="{}"&gt;&lt;SPAN data-contrast="auto"&gt;Key benefits:&lt;/SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H3&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="auto"&gt;Granular, protocol-level&amp;nbsp;control&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN data-contrast="auto"&gt;&lt;STRONG&gt;:&amp;nbsp;&lt;/STRONG&gt;Configure custom detection thresholds for TCP, UDP, and TCP SYN traffic, so mitigation triggers reflect how your application behaves rather than generic baselines.&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="auto"&gt;Predictable protection during planned&amp;nbsp;events&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN data-contrast="auto"&gt;&lt;STRONG&gt;:&lt;/STRONG&gt; Align mitigation behaviour with&amp;nbsp;anticipated&amp;nbsp;traffic&amp;nbsp;changes&amp;nbsp;product launches, seasonal peaks, gaming events,&amp;nbsp;so legitimate traffic surges&amp;nbsp;aren't&amp;nbsp;mistaken for attacks.&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="auto"&gt;Flexibility without sacrificing scale&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN data-contrast="auto"&gt;&lt;STRONG&gt;:&lt;/STRONG&gt;&amp;nbsp;Fine-tune&amp;nbsp;triggers for&amp;nbsp;specific workloads while still&amp;nbsp;benefiting&amp;nbsp;from Azure's global-scale mitigation infrastructure and adaptive protection everywhere else.&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="auto"&gt;Per-resource policy management&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN data-contrast="auto"&gt;&lt;STRONG&gt;:&lt;/STRONG&gt;&amp;nbsp;Apply separate policies across environments or applications, giving teams the ability to tune protection independently for dev, test, and production workloads.&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="auto"&gt;Full operational visibility&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN data-contrast="auto"&gt;&lt;STRONG&gt;:&lt;/STRONG&gt;&amp;nbsp;Existing Azure Monitor and DDoS Protection telemetry continue to work with custom policies, so you can&amp;nbsp;validate&amp;nbsp;threshold changes and monitor mitigation activity end to end.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;H3 aria-level="2"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;Public preview walkthrough&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Customers can deploy and manage DDoS custom policies directly from the Azure portal.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;To create a new policy:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;Open the Azure portal.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;Search for&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;DDoS custom policies&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;Select&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;Create&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;Choose the subscription, resource group, and region.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;Select a supported frontend IP configuration.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;Configure protocol detection rules.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;Review and deploy.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Once deployed, the custom policy can be associated with supported frontend IP resources to apply protocol-specific mitigation settings.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P class="lia-clear-both"&gt;&amp;nbsp;&lt;/P&gt;
&lt;P class="lia-clear-both"&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;H3 aria-level="2"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;Current public preview scope&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;During public preview, Azure DDoS Protection custom policy supports:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="8" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Standard Load Balancer frontend IP configurations&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="8" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="2" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;TCP, UDP, and TCP SYN threshold customization&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="8" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="3" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Azure portal management&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="8" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="4" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;ARM and REST API deployment&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Current limitations include:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="7" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Support is currently limited to&amp;nbsp;Standard Load Balancer frontend&amp;nbsp;IPs&amp;nbsp;and no Power Shell support&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;As with all preview features, functionality and supported scenarios may evolve before general availability.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H3 aria-level="2"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;Important considerations&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;When a customer configures a custom threshold for a protocol, Azure disables&amp;nbsp;autotuning&amp;nbsp;triggers&amp;nbsp;for&amp;nbsp;that protocol and uses the configured static value.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Customers should:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="6" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Use&amp;nbsp;anticipated&amp;nbsp;traffic baselines to guide threshold&amp;nbsp;selection&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="6" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="2" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Start with conservative tuning changes&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="6" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="3" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Validate behaviour in lower environments before broad deployment&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="6" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="4" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Monitor mitigation telemetry after configuration changes&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Azure Monitor and existing Azure DDoS Protection telemetry continue to provide visibility into mitigation activity and operational behavior.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H3 aria-level="2"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;Looking ahead&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Azure DDoS Protection continues to evolve to support modern application architectures, large-scale internet exposure, and advanced operational requirements.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Custom policy extends Azure DDoS Protection's automatic, adaptive baseline with optional per-resource control, without changing the turnkey default that protects every workload out of the box. Autotuning stays on everywhere a custom threshold is not explicitly configured.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;We want customers to know that Azure DDoS continues to be a hands-off fully automated service, and that policy customization is not a new operational model, but rather an optional override-on-top automated/adaptive engine.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H3 aria-level="2"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;Get started&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Customers can begin using Azure DDoS Protection custom policy today through the public preview experience.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;To learn more:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="2" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Review the Azure REST API documentation for DDoS Custom Policies&amp;nbsp;&lt;/SPAN&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/rest/api/virtualnetwork/ddos-custom-policies?view=rest-virtualnetwork-2025-05-01" target="_blank" rel="noopener"&gt;here&lt;/A&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="2" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="2" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Deploy a test policy in a supported subscription&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="2" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="3" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Explore the Azure portal experience for policy management&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="2" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="4" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Evaluate protocol-specific threshold tuning for your applications&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;We look forward to&amp;nbsp;hearing&amp;nbsp;your&amp;nbsp;feedback.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Thu, 23 Jul 2026 17:16:28 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-networking-blog/announcing-public-preview-of-azure-ddos-protection-custom-policy/ba-p/4538963</guid>
      <dc:creator>OfirSarfaty</dc:creator>
      <dc:date>2026-07-23T17:16:28Z</dc:date>
    </item>
    <item>
      <title>Design, test, and ship Foundry hosted agents from a canvas in GitHub Copilot App</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-developer-community/design-test-and-ship-foundry-hosted-agents-from-a-canvas-in/ba-p/4539921</link>
      <description>&lt;P&gt;Building a Microsoft Foundry hosted agent has historically meant juggling multiple surfaces — looking up a model deployment in the Foundry portal, wiring toolboxes and skills by hand, copying an endpoint into your editor, then dropping to the CLI to test and deploy. That context-switching adds friction at exactly the moment you want to be experimenting.&lt;/P&gt;
&lt;P&gt;Today we're changing that. &lt;STRONG&gt;The Microsoft Foundry Canvas is now in public preview.&lt;/STRONG&gt; With this release, the full hosted-agent workflow — discover, scaffold, configure, test, deploy — lives right beside the Copilot chat that's already writing your code.&lt;/P&gt;
&lt;P&gt;The Foundry Canvas is a GitHub Copilot App extension that opens a visual canvas in the side panel. You make choices in the canvas, and it hands each step to Copilot with your Foundry project context already attached. Copilot writes the code and runs the commands; the canvas keeps the visual state and your workspace in sync. Here are a few things this unlocks for developers.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;H1&gt;🧭 A project-aware canvas, one prompt away&lt;/H1&gt;
&lt;P&gt;Ask Copilot to create a Foundry hosted agent and the canvas opens automatically, connected to your project. It surfaces what's actually deployed — models, Foundry Toolboxes and their tools, project skills, and account guardrails — so you're choosing from real resources, not a static list.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;To get started:&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;In the GitHub Copilot App, ask Copilot: &lt;EM&gt;Create a Foundry hosted agent using the Foundry Canvas.&lt;/EM&gt;&lt;/LI&gt;
&lt;LI&gt;Open the canvas project menu, sign in to Azure if you're prompted, and pick a subscription.&lt;/LI&gt;
&lt;LI&gt;Select a Foundry project. The canvas remembers your choice across reopens.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H1&gt;✨ Scaffold and configure without leaving the panel&lt;/H1&gt;
&lt;P&gt;Start from a generated idea with &lt;STRONG&gt;Inspire me&lt;/STRONG&gt;, or begin from a &lt;STRONG&gt;Hello world&lt;/STRONG&gt; sample. From there, wire the agent to a deployed model and connect any toolboxes, skills, and guardrails. Every selection becomes a project-aware prompt to Copilot, which updates the agent code for you.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;To build your agent:&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Choose &lt;STRONG&gt;Inspire me&lt;/STRONG&gt; for a generated starting point, or the &lt;STRONG&gt;Hello world&lt;/STRONG&gt; sample prompt.&lt;/LI&gt;
&lt;LI&gt;Select a deployed model, then connect toolboxes, skills, and guardrails from your project.&lt;/LI&gt;
&lt;LI&gt;Let Copilot apply each change in your workspace as you go.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H1&gt;🔍 Test locally with the embedded Agent Inspector&lt;/H1&gt;
&lt;P&gt;When you're ready to try it, &lt;STRONG&gt;Inspect Locally&lt;/STRONG&gt; runs the agent and embeds the Agent Inspector right in the canvas — no separate REST client, no extra tooling. Chat with the agent, and if something breaks, send the error straight back to Copilot as a fix request.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;To test and deploy:&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Click &lt;STRONG&gt;Inspect Locally&lt;/STRONG&gt; — the canvas runs the agent and opens the Agent Inspector on port 8088.&lt;/LI&gt;
&lt;LI&gt;Send a prompt such as &lt;EM&gt;Write a haiku about deploying cloud applications.&lt;/EM&gt; and iterate.&lt;/LI&gt;
&lt;LI&gt;Click &lt;STRONG&gt;Deploy to Foundry&lt;/STRONG&gt; to publish the agent to Foundry Agent Service.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Because &lt;STRONG&gt;Inspect Locally&lt;/STRONG&gt; and &lt;STRONG&gt;Deploy to Foundry&lt;/STRONG&gt; run the underlying Azure Developer CLI (azd) commands, you're never boxed in — drop back to the terminal whenever you want.&lt;/P&gt;
&lt;P&gt;Hosted agents are one of the fastest-growing ways developers build on Microsoft Foundry — from internal copilots to task automation to customer-facing assistants. This release gives you a first-class, visual path to design and ship them faster, without giving up the code-first workflow you already know.&lt;/P&gt;
&lt;H1&gt;🚀 Get Started Today&lt;/H1&gt;
&lt;P&gt;Ready to build your first agent from the canvas? Here's how to jump in:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;In the GitHub Copilot App, open &lt;STRONG&gt;Settings&lt;/STRONG&gt; &amp;gt; &lt;STRONG&gt;Plugins&lt;/STRONG&gt;, search for foundry-agent-canvas, and select &lt;STRONG&gt;Install&lt;/STRONG&gt;. You can also install it from the &lt;A class="lia-external-url" href="https://awesome-copilot.github.com/extension/foundry-agent-canvas/" target="_blank" rel="noopener"&gt;&lt;U&gt;awesome-copilot listing&lt;/U&gt;&lt;/A&gt; or ask Copilot to install it into user scope.&lt;/LI&gt;
&lt;LI&gt;Ask Copilot to create a Foundry hosted agent, then follow the canvas from project pick to deploy.&lt;/LI&gt;
&lt;LI&gt;Walk through it end to end with the &lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/foundry/agents/quickstarts/quickstart-hosted-agent?pivots=canvas" target="_blank" rel="noopener"&gt;&lt;U&gt;Deploy your first hosted agent quickstart&lt;/U&gt;&lt;/A&gt;, and read the &lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/foundry-agent-canvas" target="_blank" rel="noopener"&gt;&lt;U&gt;Foundry Canvas overview&lt;/U&gt;&lt;/A&gt; for the full picture.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;We'd love to hear from you! Whether it's a feature request, a bug report, or feedback on your experience, join the conversation and contribute directly on our &lt;A class="lia-external-url" href="https://github.com/microsoft/foundry-toolkit" target="_blank" rel="noopener"&gt;&lt;U&gt;GitHub repository&lt;/U&gt;&lt;/A&gt;.&lt;/P&gt;
&lt;P&gt;Happy Coding!&lt;/P&gt;</description>
      <pubDate>Fri, 24 Jul 2026 09:19:54 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-developer-community/design-test-and-ship-foundry-hosted-agents-from-a-canvas-in/ba-p/4539921</guid>
      <dc:creator>junjieli</dc:creator>
      <dc:date>2026-07-24T09:19:54Z</dc:date>
    </item>
    <item>
      <title>Cannot run "Get-SolutionUpdate" or "Get-SolutionUpdateEnvironment" after updating to 2604</title>
      <link>https://techcommunity.microsoft.com/t5/azure-stack/cannot-run-quot-get-solutionupdate-quot-or-quot-get/m-p/4539895#M305</link>
      <description>&lt;P&gt;My Azure Local environment had been a bit out of date for the last few months so I was catching it up with the updates, stepping from 2602 to 2603 and then to 2604 last night. Since that update, when I run "Get-SolutionUpdate" or "Get-SolutionUpdateEnvironment" the command hangs for a few minutes and then eventually returns with something like the following:&lt;/P&gt;&lt;LI-CODE lang=""&gt;get-solutionUpdate : A WebException occurred while sending a RestRequest. WebException.Status: SendFailure on
https://FQDN:4900/providers/Microsoft.Update.Admin/updateLocations?api-version=2022-08-01
    + CategoryInfo          : ConnectionError: (:) [Get-SolutionUpdate], SendFailureException
    + FullyQualifiedErrorId : SendFailureException,Microsoft.AzureStack.Lcm.PowerShell.GetSolutionUpdateCmdlet

get-solutionUpdate : Object reference not set to an instance of an object.
    + CategoryInfo          : InvalidOperation: (:) [Get-SolutionUpdate], NullReferenceException
    + FullyQualifiedErrorId : NullReferenceException,Microsoft.AzureStack.Lcm.PowerShell.GetSolutionUpdateCmdlet&lt;/LI-CODE&gt;&lt;P&gt;I have tried &lt;A class="lia-external-url" href="https://github.com/Azure/AzureLocal-Supportability/blob/main/TSG/Update/Get-SolutionUpdate-GatewayTimeout.md" target="_blank"&gt;the guide here&lt;/A&gt; to try resolve it as it seemed similar but it hasn't helped. Any idea what might be going on?&lt;/P&gt;</description>
      <pubDate>Thu, 23 Jul 2026 05:17:13 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-stack/cannot-run-quot-get-solutionupdate-quot-or-quot-get/m-p/4539895#M305</guid>
      <dc:creator>MattENZ</dc:creator>
      <dc:date>2026-07-23T05:17:13Z</dc:date>
    </item>
    <item>
      <title>Hybrid Logic Apps on RKE2: a self-managed cluster with MetalLB</title>
      <link>https://techcommunity.microsoft.com/t5/azure-integration-services-blog/hybrid-logic-apps-on-rke2-a-self-managed-cluster-with-metallb/ba-p/4539846</link>
      <description>&lt;P&gt;Azure Logic Apps Hybrid lets you run the Logic Apps runtime on your own Kubernetes cluster and still author, deploy, and monitor from Azure. The &lt;A href="https://techcommunity.microsoft.com/blog/integrationsonazureblog/hybrid-logic-apps-deployment-on-red-hat-openshift/4534828" target="_blank"&gt;previous post&lt;/A&gt; walked through Red Hat OpenShift. This one is RKE2. We spun up a single-node RKE2 cluster, gave the ingress an IP with MetalLB, connected it to Arc, and shipped a Hybrid Logic App end-to-end. A handful of things needed care along the way, and those are what this post is really about.&lt;/P&gt;
&lt;P&gt;If you've done Hybrid on AKS or K3s, the shape will look familiar. The &lt;A href="https://learn.microsoft.com/azure/logic-apps/set-up-standard-workflows-hybrid-deployment-requirements" target="_blank"&gt;official requirements doc&lt;/A&gt; covers the general flow. RKE2 is close enough to vanilla Kubernetes that most of it just works out of the box. Four things don't. Three of them are DNS. They're called out at the step they matter.&lt;/P&gt;
&lt;H2&gt;What makes RKE2 different?&lt;/H2&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Area&lt;/th&gt;&lt;th&gt;What we found&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Ingress IP&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;RKE2 has an embedded cloud-controller-manager for node lifecycle, but no service load balancer. &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;type: LoadBalancer&lt;/CODE&gt; sits at &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;&amp;lt;pending&amp;gt;&lt;/CODE&gt; until you bring one. We used MetalLB (Step 2).&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Pod security&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;No SCC dance like OpenShift; the extension pods schedule as-is. What we did hit was the node's &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;inotify&lt;/CODE&gt; limit. Enough runtime pods use it that the default runs out and about half the deployment crash-loops (Step 4).&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;CoreDNS&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;RKE2 names its DNS objects &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;rke2-coredns-rke2-coredns&lt;/CODE&gt;, not &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;coredns&lt;/CODE&gt;/&lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;kube-dns&lt;/CODE&gt;. &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;az containerapp arc setup-core-dns&lt;/CODE&gt; doesn't have an RKE2 distro flag yet, so DNS is manual (Step 7).&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Distribution flag&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;RKE2 is upstream Kubernetes. Don't copy the OpenShift install command; the &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;Azure.Cluster.Distribution=openshift&lt;/CODE&gt; and &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;coreDNSVersion&lt;/CODE&gt; overrides don't apply here and will bite you (Step 4).&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H2&gt;Prerequisites&lt;/H2&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;A Linux host&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Anywhere RKE2 runs: bare metal, a VM, an edge appliance. We used a single Ubuntu 22.04 Azure VM (&lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;Standard_D8as_v5&lt;/CODE&gt;) so everything was easy to nuke afterwards. 4 vCPU / 16 GB is the practical floor.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Azure subscription&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Rights to create Arc resources, custom locations, and connected environments.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Azure CLI&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Latest, with the &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;connectedk8s&lt;/CODE&gt;, &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;k8s-extension&lt;/CODE&gt;, &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;customlocation&lt;/CODE&gt;, and &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;containerapp&lt;/CODE&gt; extensions.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;An SMB file share&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Reachable from the cluster, for workflow artifacts.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;A spare IP range&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;A small unused range on your node's L2 network for MetalLB to hand out (Step 2).&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;PRE style="background: #f6f8fa; border: 1px solid #d0d7de; border-radius: 6px; padding: 14px 16px; overflow: auto; font-family: Consolas,Menlo,'Courier New',monospace; font-size: 13px; line-height: 1.55; white-space: pre; color: #24292e;"&gt;az extension add --name connectedk8s
az extension add --name k8s-extension
az extension add --name customlocation
az extension add --name containerapp&lt;/PRE&gt;
&lt;H2&gt;Step 1: Stand up RKE2&lt;/H2&gt;
&lt;P&gt;One line and a couple of minutes:&lt;/P&gt;
&lt;PRE style="background: #f6f8fa; border: 1px solid #d0d7de; border-radius: 6px; padding: 14px 16px; overflow: auto; font-family: Consolas,Menlo,'Courier New',monospace; font-size: 13px; line-height: 1.55; white-space: pre; color: #24292e;"&gt;curl -sfL https://get.rke2.io | sudo sh -
sudo systemctl enable --now rke2-server.service

# kubectl + kubeconfig are dropped in place
export PATH=$PATH:/var/lib/rancher/rke2/bin
export KUBECONFIG=/etc/rancher/rke2/rke2.yaml
kubectl get nodes          # Ready in ~2-5 min after the first image pull&lt;/PRE&gt;
&lt;P&gt;RKE2 servers are schedulable by default, so this single node is your entire cluster. No taint to remove. If you want to run kubectl from your laptop, add the node's public IP or hostname to &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;tls-san&lt;/CODE&gt; in &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;/etc/rancher/rke2/config.yaml&lt;/CODE&gt; &lt;EM&gt;before&lt;/EM&gt; the first start, then copy &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;rke2.yaml&lt;/CODE&gt; out and swap &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;127.0.0.1&lt;/CODE&gt; for that address.&lt;/P&gt;
&lt;BLOCKQUOTE style="border-left: 4px solid #0067b8; background: #f0f7ff; margin: 20px 0; padding: 10px 16px; border-radius: 0 6px 6px 0;"&gt;
&lt;P&gt;&lt;STRONG&gt;Quick sanity check.&lt;/STRONG&gt; Expose a throwaway nginx as &lt;CODE style="background: #e6eefc; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;LoadBalancer&lt;/CODE&gt;: &lt;CODE style="background: #e6eefc; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;kubectl create deploy nginx --image=nginx &amp;amp;&amp;amp; kubectl expose deploy nginx --port=80 --type=LoadBalancer&lt;/CODE&gt;. The &lt;CODE style="background: #e6eefc; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;EXTERNAL-IP&lt;/CODE&gt; will sit at &lt;CODE style="background: #e6eefc; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;&amp;lt;pending&amp;gt;&lt;/CODE&gt; forever. That's the gap MetalLB fills in the next step. Clean up the deployment and service when you're done looking.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;Step 2: Give the ingress an IP with MetalLB&lt;/H2&gt;
&lt;P&gt;All the Logic Apps in this environment share a single Envoy ingress, and something has to give that service an IP. A cloud does it for you; on-prem, you pick one of two options:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;In-cluster load balancer.&lt;/STRONG&gt; MetalLB (or Cilium, kube-vip, OpenELB) assigns an IP from a pool on your node network. Leave &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;envoy.serviceType=LoadBalancer&lt;/CODE&gt; in Step 4. This is what we did.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;External L4 load balancer in front of NodePort.&lt;/STRONG&gt; If you already run an F5, HAProxy, NetScaler, or an Azure Standard Load Balancer, keep using it. Set &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;envoy.serviceType=NodePort&lt;/CODE&gt; (plus &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;envoy.nodeHttpsPort&lt;/CODE&gt; / &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;envoy.nodeHttpPort&lt;/CODE&gt;) and point your LB at those NodePorts. Then pass the LB's VIP as &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;--static-ip&lt;/CODE&gt; in Step 6.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;We went with MetalLB in L2 mode:&lt;/P&gt;
&lt;PRE style="background: #f6f8fa; border: 1px solid #d0d7de; border-radius: 6px; padding: 14px 16px; overflow: auto; font-family: Consolas,Menlo,'Courier New',monospace; font-size: 13px; line-height: 1.55; white-space: pre; color: #24292e;"&gt;kubectl apply -f https://raw.githubusercontent.com/metallb/metallb/v0.14.9/config/manifests/metallb-native.yaml
kubectl wait --for=condition=Ready pods --all -n metallb-system --timeout=180s&lt;/PRE&gt;
&lt;P&gt;Then give it a pool from your node subnet and advertise it over L2:&lt;/P&gt;
&lt;PRE style="background: #f6f8fa; border: 1px solid #d0d7de; border-radius: 6px; padding: 14px 16px; overflow: auto; font-family: Consolas,Menlo,'Courier New',monospace; font-size: 13px; line-height: 1.55; white-space: pre; color: #24292e;"&gt;apiVersion: metallb.io/v1beta1
kind: IPAddressPool
metadata:
  name: hybrid-pool
  namespace: metallb-system
spec:
  addresses:
    - 192.168.1.240-192.168.1.250   # a spare range on your node network
---
apiVersion: metallb.io/v1beta1
kind: L2Advertisement
metadata:
  name: hybrid-l2
  namespace: metallb-system
spec:
  ipAddressPools:
    - hybrid-pool&lt;/PRE&gt;
&lt;P&gt;Two things to watch. The range should be free on your node subnet, and it must not collide with RKE2's cluster (&lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;10.42.0.0/16&lt;/CODE&gt;) or service (&lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;10.43.0.0/16&lt;/CODE&gt;) CIDRs. Re-run the throwaway &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;LoadBalancer&lt;/CODE&gt; test from Step 1; the &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;EXTERNAL-IP&lt;/CODE&gt; should now be one of your pool addresses.&lt;/P&gt;
&lt;BLOCKQUOTE style="border-left: 4px solid #0067b8; background: #f0f7ff; margin: 20px 0; padding: 10px 16px; border-radius: 0 6px 6px 0;"&gt;
&lt;P&gt;&lt;STRONG&gt;Running MetalLB on one cloud VM?&lt;/STRONG&gt; L2 mode ARPs the pool IPs on the node's network, but a cloud SDN won't route an unassigned IP to your NIC from the outside. It works fine for anything inside the cluster (kube-proxy programs the IP locally), but to hit envoy from outside the VM you'll need to &lt;CODE style="background: #e6eefc; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;socat&lt;/CODE&gt;-forward the VM's public IP to the LB IP, or add the pool IP as a secondary ipconfig on the NIC. On real on-prem L2 you don't have to think about any of this.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;Step 3: Connect the cluster to Azure Arc&lt;/H2&gt;
&lt;PRE style="background: #f6f8fa; border: 1px solid #d0d7de; border-radius: 6px; padding: 14px 16px; overflow: auto; font-family: Consolas,Menlo,'Courier New',monospace; font-size: 13px; line-height: 1.55; white-space: pre; color: #24292e;"&gt;az connectedk8s connect --name &amp;lt;cluster&amp;gt; --resource-group &amp;lt;rg&amp;gt; --location &amp;lt;region&amp;gt;
kubectl get pods -n azure-arc      # all agents Running in a couple of minutes&lt;/PRE&gt;
&lt;BLOCKQUOTE style="border-left: 4px solid #0067b8; background: #f0f7ff; margin: 20px 0; padding: 10px 16px; border-radius: 0 6px 6px 0;"&gt;
&lt;P&gt;If your host is behind a strict egress policy, or in our case an internal MS VM with an inbound-deny NRMS policy, remember that Arc onboarding only needs &lt;EM&gt;outbound&lt;/EM&gt; connectivity and a &lt;EM&gt;local&lt;/EM&gt; kubeconfig. We ran the connect command on the VM itself using a system-assigned managed identity (&lt;CODE style="background: #e6eefc; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;az login --identity&lt;/CODE&gt;). No inbound port, no browser device-code prompt.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;Step 4: Install the Container Apps extension&lt;/H2&gt;
&lt;P&gt;This is the extension that turns the cluster into a Logic Apps host. On RKE2 it's the plain command, with &lt;STRONG&gt;no&lt;/STRONG&gt; &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;Azure.Cluster.Distribution&lt;/CODE&gt; or &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;coreDNSVersion&lt;/CODE&gt; overrides. If you copy those from an OpenShift guide, the install will still succeed and then behave in mildly confusing ways.&lt;/P&gt;
&lt;PRE style="background: #f6f8fa; border: 1px solid #d0d7de; border-radius: 6px; padding: 14px 16px; overflow: auto; font-family: Consolas,Menlo,'Courier New',monospace; font-size: 13px; line-height: 1.55; white-space: pre; color: #24292e;"&gt;az k8s-extension create \
  --resource-group &amp;lt;rg&amp;gt; --cluster-name &amp;lt;cluster&amp;gt; \
  --cluster-type connectedClusters --name logicapps-aca-extension \
  --extension-type Microsoft.App.Environment \
  --release-train stable --auto-upgrade-minor-version true --scope cluster \
  --release-namespace logicapps-aca-ns \
  --configuration-settings "Microsoft.CustomLocation.ServiceAccount=default" \
  --configuration-settings "appsNamespace=logicapps-aca-ns" \
  --configuration-settings "clusterName=&amp;lt;env-name&amp;gt;" \
  --configuration-settings "keda.enabled=true" \
  --configuration-settings "keda.logicAppsScaler.enabled=true" \
  --configuration-settings "keda.logicAppsScaler.replicaCount=1" \
  --configuration-settings "containerAppController.api.functionsServerEnabled=true" \
  --configuration-settings "functionsProxyApiConfig.enabled=true" \
  --configuration-settings "envoy.serviceType=LoadBalancer"&lt;/PRE&gt;
&lt;P&gt;Our first install came back "Succeeded", but half the pods were in &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;CrashLoopBackOff&lt;/CODE&gt;. &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;billing&lt;/CODE&gt;, &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;containerapp-controller&lt;/CODE&gt;, &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;keda-logicapps-scaler&lt;/CODE&gt;, &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;mdm&lt;/CODE&gt;, and &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;log-processor&lt;/CODE&gt; all had the same line in their logs:&lt;/P&gt;
&lt;PRE style="background: #f6f8fa; border: 1px solid #d0d7de; border-radius: 6px; padding: 14px 16px; overflow: auto; font-family: Consolas,Menlo,'Courier New',monospace; font-size: 13px; line-height: 1.55; white-space: pre; color: #24292e;"&gt;The configured user limit (128) on the number of inotify instances has been reached&lt;/PRE&gt;
&lt;P&gt;These are .NET services with &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;FileSystemWatcher&lt;/CODE&gt; under the hood. The default &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;fs.inotify.max_user_instances=128&lt;/CODE&gt; on stock Ubuntu is comfortably below what the extension needs. Raise it on the node, delete the pods, done:&lt;/P&gt;
&lt;PRE style="background: #f6f8fa; border: 1px solid #d0d7de; border-radius: 6px; padding: 14px 16px; overflow: auto; font-family: Consolas,Menlo,'Courier New',monospace; font-size: 13px; line-height: 1.55; white-space: pre; color: #24292e;"&gt;sudo sysctl -w fs.inotify.max_user_instances=8192
sudo sysctl -w fs.inotify.max_user_watches=1048576
echo -e "fs.inotify.max_user_instances=8192\nfs.inotify.max_user_watches=1048576" | sudo tee /etc/sysctl.d/99-inotify.conf
kubectl delete pods --all -n logicapps-aca-ns&lt;/PRE&gt;
&lt;P&gt;After that everything came up green, and envoy grabbed a MetalLB address on its own. Confirm:&lt;/P&gt;
&lt;PRE style="background: #f6f8fa; border: 1px solid #d0d7de; border-radius: 6px; padding: 14px 16px; overflow: auto; font-family: Consolas,Menlo,'Courier New',monospace; font-size: 13px; line-height: 1.55; white-space: pre; color: #24292e;"&gt;az k8s-extension show -g &amp;lt;rg&amp;gt; --cluster-name &amp;lt;cluster&amp;gt; \
  --cluster-type connectedClusters --name logicapps-aca-extension \
  --query provisioningState -o tsv          # -&amp;gt; Succeeded
kubectl get svc microsoft-app-environment-k8se-envoy -n logicapps-aca-ns   # EXTERNAL-IP from your pool&lt;/PRE&gt;
&lt;H2&gt;Step 5: Create a custom location&lt;/H2&gt;
&lt;P&gt;Turn on the custom-locations feature, then create the location that ties the cluster, extension, and namespace together:&lt;/P&gt;
&lt;PRE style="background: #f6f8fa; border: 1px solid #d0d7de; border-radius: 6px; padding: 14px 16px; overflow: auto; font-family: Consolas,Menlo,'Courier New',monospace; font-size: 13px; line-height: 1.55; white-space: pre; color: #24292e;"&gt;az connectedk8s enable-features -n &amp;lt;cluster&amp;gt; -g &amp;lt;rg&amp;gt; \
  --features cluster-connect custom-locations \
  --custom-locations-oid $(az ad sp show --id bc313c14-388c-4e7d-a58e-70017303ee3b --query id -o tsv)

EXT_ID=$(az k8s-extension show -g &amp;lt;rg&amp;gt; --cluster-name &amp;lt;cluster&amp;gt; \
  --cluster-type connectedClusters --name logicapps-aca-extension --query id -o tsv)
CC_ID=$(az connectedk8s show -g &amp;lt;rg&amp;gt; -n &amp;lt;cluster&amp;gt; --query id -o tsv)

az customlocation create \
  --resource-group &amp;lt;rg&amp;gt; --name &amp;lt;custom-location&amp;gt; \
  --location &amp;lt;region&amp;gt; --host-resource-id $CC_ID \
  --namespace logicapps-aca-ns --cluster-extension-ids $EXT_ID&lt;/PRE&gt;
&lt;BLOCKQUOTE style="border-left: 4px solid #0067b8; background: #f0f7ff; margin: 20px 0; padding: 10px 16px; border-radius: 0 6px 6px 0;"&gt;
&lt;P&gt;One thing we tripped on: the extension type isn't offered in every Azure region. It's registered in &lt;CODE style="background: #e6eefc; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;eastus&lt;/CODE&gt;, &lt;CODE style="background: #e6eefc; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;westus&lt;/CODE&gt;, &lt;CODE style="background: #e6eefc; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;westeurope&lt;/CODE&gt;, and a handful of others, but not &lt;CODE style="background: #e6eefc; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;eastus2&lt;/CODE&gt; at the time of writing. The &lt;CODE style="background: #e6eefc; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;connectedCluster&lt;/CODE&gt; resource's region is independent of where the node actually runs, so pick a supported one here.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;Step 6: Create the connected environment&lt;/H2&gt;
&lt;PRE style="background: #f6f8fa; border: 1px solid #d0d7de; border-radius: 6px; padding: 14px 16px; overflow: auto; font-family: Consolas,Menlo,'Courier New',monospace; font-size: 13px; line-height: 1.55; white-space: pre; color: #24292e;"&gt;CL_ID=$(az customlocation show -g &amp;lt;rg&amp;gt; -n &amp;lt;custom-location&amp;gt; --query id -o tsv)

az containerapp connected-env create \
  --resource-group &amp;lt;rg&amp;gt; --name &amp;lt;env-name&amp;gt; \
  --location &amp;lt;region&amp;gt; --custom-location $CL_ID&lt;/PRE&gt;
&lt;P&gt;Because MetalLB already put an &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;EXTERNAL-IP&lt;/CODE&gt; on the envoy &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;LoadBalancer&lt;/CODE&gt; service, the platform picks it up automatically. &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;--static-ip&lt;/CODE&gt; is only needed if you went the NodePort + external LB route in Step 2, in which case you pass your LB's VIP here.&lt;/P&gt;
&lt;H2&gt;Step 7: Configure cluster DNS&lt;/H2&gt;
&lt;P&gt;App hostnames like &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;*.&amp;lt;env&amp;gt;.&amp;lt;region&amp;gt;.k4apps.io&lt;/CODE&gt; need to resolve to the Envoy ingress &lt;EM&gt;from inside&lt;/EM&gt; the cluster. On AKS or OpenShift, &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;az containerapp arc setup-core-dns&lt;/CODE&gt; handles this. RKE2 isn't a supported distro on that command yet, so it's manual, but it's a small manual.&lt;/P&gt;
&lt;P&gt;The extension already writes the correct rewrite and forward rules into a &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;coredns-custom&lt;/CODE&gt; ConfigMap. What's missing is CoreDNS mounting that ConfigMap and importing it from the Corefile.&lt;/P&gt;
&lt;P&gt;We tried the direct route first: patch the Corefile ConfigMap, edit the Deployment to add the volume, restart. It works, right up until the next &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;rke2-server&lt;/CODE&gt; restart. RKE2's helm controller re-renders both objects from the chart and wipes your changes. The right way is a &lt;STRONG&gt;&lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;HelmChartConfig&lt;/CODE&gt;&lt;/STRONG&gt;. The &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;rke2-coredns&lt;/CODE&gt; chart exposes &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;extraConfig&lt;/CODE&gt; (config outside the default zone block) and &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;extraVolumes&lt;/CODE&gt;/&lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;extraVolumeMounts&lt;/CODE&gt;, which is exactly what we need:&lt;/P&gt;
&lt;PRE style="background: #f6f8fa; border: 1px solid #d0d7de; border-radius: 6px; padding: 14px 16px; overflow: auto; font-family: Consolas,Menlo,'Courier New',monospace; font-size: 13px; line-height: 1.55; white-space: pre; color: #24292e;"&gt;# /var/lib/rancher/rke2/server/manifests/rke2-coredns-config.yaml
apiVersion: helm.cattle.io/v1
kind: HelmChartConfig
metadata:
  name: rke2-coredns
  namespace: kube-system
spec:
  valuesContent: |-
    extraConfig:
      import:
        parameters: /etc/coredns/custom/*.server
    extraVolumes:
      - name: custom-config-volume
        configMap:
          name: coredns-custom
          optional: true
    extraVolumeMounts:
      - name: custom-config-volume
        mountPath: /etc/coredns/custom&lt;/PRE&gt;
&lt;P&gt;Drop the file in &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;/var/lib/rancher/rke2/server/manifests/&lt;/CODE&gt;, or &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;kubectl apply -f&lt;/CODE&gt; it (the helm controller watches both). Within a minute CoreDNS redeploys with the mount and the top-level &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;import&lt;/CODE&gt;. This time it survives restarts and RKE2 upgrades because the chart is now the source of truth.&lt;/P&gt;
&lt;P&gt;To verify, run &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;nslookup test.&amp;lt;env-domain&amp;gt;&lt;/CODE&gt; from a throwaway pod. You should get back the &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;microsoft-app-environment-k8se-envoy-internal&lt;/CODE&gt; ClusterIP.&lt;/P&gt;
&lt;BLOCKQUOTE style="border-left: 4px solid #0067b8; background: #f0f7ff; margin: 20px 0; padding: 10px 16px; border-radius: 0 6px 6px 0;"&gt;
&lt;P&gt;&lt;STRONG&gt;One last DNS thing: a &lt;CODE style="background: #e6eefc; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;kube-dns&lt;/CODE&gt; alias.&lt;/STRONG&gt; The Logic Apps platform locates cluster DNS by looking for a Service named &lt;CODE style="background: #e6eefc; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;kube-dns&lt;/CODE&gt; in &lt;CODE style="background: #e6eefc; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;kube-system&lt;/CODE&gt;. That's the convention on AKS, K3s, kubeadm, and most other distros. RKE2 uses &lt;CODE style="background: #e6eefc; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;rke2-coredns-rke2-coredns&lt;/CODE&gt; and doesn't ship the alias, so the platform can't find DNS. Symptom: revision calls come back as 502s. Fix: create the alias.&lt;/P&gt;
&lt;PRE style="background: #eef2f6; border: 1px solid #cdd7e1; border-radius: 6px; padding: 12px 14px; overflow: auto; font-family: Consolas,Menlo,'Courier New',monospace; font-size: 13px; line-height: 1.55; white-space: pre; color: #24292e;"&gt;apiVersion: v1
kind: Service
metadata: { name: kube-dns, namespace: kube-system, labels: { k8s-app: kube-dns } }
spec:
  selector: { k8s-app: kube-dns }
  ports:
    - { name: dns, port: 53, protocol: UDP, targetPort: 53 }
    - { name: dns-tcp, port: 53, protocol: TCP, targetPort: 53 }&lt;/PRE&gt;
&lt;P&gt;It's a plain Service. Nothing about it gets templated or overwritten, so it just stays put across reconciles and upgrades.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;Step 8: Create and deploy a Logic App&lt;/H2&gt;
&lt;P&gt;Plumbing's out of the way. In the portal, create a new &lt;STRONG&gt;Logic App (Standard)&lt;/STRONG&gt; with the &lt;STRONG&gt;Hybrid&lt;/STRONG&gt; hosting option, point it at your connected environment, wire storage up to your SMB share, and create. Author the workflow the way you always would. The runs will execute on your RKE2 cluster instead of in Azure.&lt;/P&gt;
&lt;H2&gt;Troubleshooting: the RKE2 things we hit&lt;/H2&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Symptom&lt;/th&gt;&lt;th&gt;Cause &amp;amp; fix&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Test &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;LoadBalancer&lt;/CODE&gt; / envoy stays &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;&amp;lt;pending&amp;gt;&lt;/CODE&gt;&lt;/td&gt;&lt;td&gt;No LB provider on the cluster. Install MetalLB (Step 2). RKE2 ships no ServiceLB.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Runtime pods &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;CrashLoopBackOff&lt;/CODE&gt; with &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;inotify instances ... reached&lt;/CODE&gt;&lt;/td&gt;&lt;td&gt;Node &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;fs.inotify.max_user_instances&lt;/CODE&gt; too low. Raise it to 8192 and restart the pods (Step 4).&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Revision invoke returns &lt;STRONG&gt;502&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Platform can't find a &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;kube-dns&lt;/CODE&gt; Service; RKE2's CoreDNS is under a different name. Add the alias (Step 7).&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;App hostnames don't resolve inside the cluster&lt;/td&gt;&lt;td&gt;CoreDNS isn't loading &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;coredns-custom&lt;/CODE&gt;. Apply the &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;HelmChartConfig&lt;/CODE&gt; (Step 7).&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Extension type "not registered in region"&lt;/td&gt;&lt;td&gt;The Arc cluster is in an unsupported region. Reconnect it in a supported one; it's independent of where the node runs (Step 5).&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Cluster shows Arc &lt;STRONG&gt;Offline&lt;/STRONG&gt; after a reboot; &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;azure-arc&lt;/CODE&gt; pods &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;Unknown&lt;/CODE&gt;&lt;/td&gt;&lt;td&gt;Stale pods from an ungraceful stop. &lt;CODE style="background: #eff1f3; padding: 1px 5px; border-radius: 4px; font-family: Consolas,Menlo,'Courier New',monospace; font-size: .92em; color: #24292e;"&gt;kubectl delete pods --all -n azure-arc --force&lt;/CODE&gt; and they'll recreate and reconnect.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H2&gt;References&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/logic-apps/set-up-standard-workflows-hybrid-deployment-requirements" target="_blank"&gt;Set up Logic Apps Hybrid deployment&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://techcommunity.microsoft.com/blog/integrationsonazureblog/hybrid-logic-apps-deployment-on-red-hat-openshift/4534828" target="_blank"&gt;Hybrid Logic Apps on Red Hat OpenShift&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://docs.rke2.io/" target="_blank"&gt;RKE2 documentation&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://metallb.io/" target="_blank"&gt;MetalLB&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/azure-arc/kubernetes/quickstart-connect-cluster" target="_blank"&gt;Connect a Kubernetes cluster to Azure Arc&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;</description>
      <pubDate>Thu, 23 Jul 2026 05:53:29 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-integration-services-blog/hybrid-logic-apps-on-rke2-a-self-managed-cluster-with-metallb/ba-p/4539846</guid>
      <dc:creator>anandgmenon</dc:creator>
      <dc:date>2026-07-23T05:53:29Z</dc:date>
    </item>
    <item>
      <title>Reminder: Path to Production for Agents Webinar Series Starts Next Week</title>
      <link>https://techcommunity.microsoft.com/t5/azure-architecture-blog/reminder-path-to-production-for-agents-webinar-series-starts/ba-p/4539877</link>
      <description>&lt;P&gt;Next week, join Microsoft for the&amp;nbsp;&lt;STRONG&gt;Path to Production for Agents&lt;/STRONG&gt; webinar series—a free, six-session technical training designed to help organizations move from AI experimentation to secure, scalable, production-ready agent solutions. The series takes place &lt;STRONG&gt;July 27–28&lt;/STRONG&gt; and is aimed at architects, technical leaders, engineers, and AI practitioners looking to operationalize AI at enterprise scale.&lt;/P&gt;
&lt;P&gt;Across six expert-led sessions, you'll learn how to:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Establish governance foundations for AI at scale&lt;/LI&gt;
&lt;LI&gt;Design production-ready AI platforms and landing zones&lt;/LI&gt;
&lt;LI&gt;Build reliable multi-agent architectures&lt;/LI&gt;
&lt;LI&gt;Implement AgentOps practices for deployment and observability&lt;/LI&gt;
&lt;LI&gt;Secure and govern AI systems&lt;/LI&gt;
&lt;LI&gt;Optimize performance, cost, and scalability&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Each session includes proven Microsoft architecture patterns, real-world engineering guidance, and practical techniques you can apply immediately to accelerate your path from prototype to production.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Register today:&lt;/STRONG&gt; &lt;A href="https://aka.ms/AccelerateThePathToProduction" target="_blank"&gt;Path to Production for Agents Registration&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;Don't miss this opportunity to gain the knowledge and frameworks needed to confidently deploy production-grade AI agents at scale. We look forward to seeing you next week!&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Thu, 23 Jul 2026 00:15:07 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-architecture-blog/reminder-path-to-production-for-agents-webinar-series-starts/ba-p/4539877</guid>
      <dc:creator>brauerblogs</dc:creator>
      <dc:date>2026-07-23T00:15:07Z</dc:date>
    </item>
    <item>
      <title>Azure Local expands SAN capabilities with iSCSI support</title>
      <link>https://techcommunity.microsoft.com/t5/azure-arc-blog/azure-local-expands-san-capabilities-with-iscsi-support/ba-p/4531999</link>
      <description>&lt;P&gt;As organizations modernize datacenters and accelerate migration from legacy virtualization platforms, flexibility in storage architecture has become a key requirement. Customers increasingly want to reuse existing storage investments, scale infrastructure independently, and choose the connectivity model that best fits their environment.&lt;/P&gt;
&lt;P&gt;Building on the general availability of Fibre Channel (FC) SAN support, Azure Local now introduces &lt;STRONG&gt;iSCSI SAN integration&lt;/STRONG&gt;, extending disaggregated architecture support to IP-based storage networks. With support for both Fibre Channel and iSCSI, Azure Local provides customers greater flexibility in how they modernize and scale their infrastructure while maintaining an Azure-consistent management experience.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Expanding disaggregated infrastructure&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;iSCSI support enables organizations to:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Leverage existing IP-based storage networks&lt;/LI&gt;
&lt;LI&gt;Deploy cost-efficient disaggregated architectures without FC dependency&lt;/LI&gt;
&lt;LI&gt;Scale compute and storage independently&lt;/LI&gt;
&lt;LI&gt;Maintain an Azure-consistent management experience&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Architecture and deployment&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Azure Local supports &lt;STRONG&gt;6-adapter configurations&lt;/STRONG&gt; to balance cost, performance, and resiliency:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;6 adapters&lt;/STRONG&gt;: enhanced performance and redundancy with dedicated iSCSI paths.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Today, iSCSI follows a &lt;STRONG&gt;manual configuration flow during deployment&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Azure Local Deployment&amp;nbsp;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Azure Local supports two SAN deployment approaches:&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Hybrid deployments (S2D + SAN)&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Customers can attach external SAN storage to existing Azure Local deployments while continuing to use Storage Spaces Direct (S2D) for platform storage. This approach enables organizations to incrementally adopt SAN while reusing existing storage investments.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Disaggregated deployments (SAN-only)&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Customers can also deploy Azure Local using external SAN storage as the primary storage platform for both infrastructure and workloads. This enables:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Independent scaling of compute and storage&lt;/LI&gt;
&lt;LI&gt;Fibre Channel or iSCSI connectivity&lt;/LI&gt;
&lt;LI&gt;Larger-scale infrastructure deployments&lt;/LI&gt;
&lt;LI&gt;Connected and disconnected deployment models&lt;/LI&gt;
&lt;LI&gt;Manage rising disk costs associated with hyperconverged architectures&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Additionally, customers can create local availability zones to align VM placement with physical infrastructure boundaries and support more granular workload placement. The deployment also validates connected SAN arrays against the supported vendor ecosystem, helping ensure a streamlined and fully supported experience.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Accelerating infrastructure modernization&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Azure Migrate now supports migration to Azure Local deployments that use external SAN storage, including NTFS-based volumes. This allows organizations to modernize compute infrastructure while preserving existing storage investments.&lt;/P&gt;
&lt;P&gt;Customers can:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Reuse existing SAN arrays and operational processes&lt;/LI&gt;
&lt;LI&gt;Minimize disruption during modernization projects&lt;/LI&gt;
&lt;LI&gt;Retain familiar storage architectures while adopting Azure Local&lt;/LI&gt;
&lt;LI&gt;Simplify migration from existing virtualization environments&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;What's next&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;iSCSI support represents another step in our broader vision for external storage on Azure Local. Our goal is to provide a comprehensive storage platform that spans deployment, operations, protection, and recovery.&lt;/P&gt;
&lt;P&gt;Looking ahead, we are investing in:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Integration of iSCSI node configuration into cluster deployment to simplify the initial setup.&lt;/LI&gt;
&lt;LI&gt;Business Continuity and Disaster Recovery (BCDR) for SAN-backed workloads, including replication, failover, and failback capabilities between two external SAN attached or disaggregated Azure local clusters.&lt;/LI&gt;
&lt;LI&gt;Day-N storage management experiences that simplify monitoring, troubleshooting, and operational workflows.&lt;/LI&gt;
&lt;LI&gt;Replication management capabilities that provide visibility into recovery readiness, replication health, and workload mobility across environments.&lt;/LI&gt;
&lt;LI&gt;Expanding enterprise storage vendors ecosystem for Azure Local, helping customers adopt Azure Local while preserving their existing storage investments.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Together, these investments will extend Azure Local beyond SAN connectivity and deployment to deliver a unified storage management experience across a broad ecosystem of enterprise storage solutions.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Summary&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;With iSCSI support, Azure Local now delivers a more complete SAN strategy—giving customers the flexibility to choose Fibre Channel or iSCSI, deploy hybrid or fully disaggregated architectures, and modernize infrastructure without abandoning existing storage investments.&lt;/P&gt;
&lt;P&gt;As we continue to invest in SAN management, replication, and disaster recovery, Azure Local is evolving into a comprehensive platform for enterprise storage and infrastructure modernization—from edge deployments to sovereign-scale datacenters.&lt;/P&gt;</description>
      <pubDate>Wed, 22 Jul 2026 22:15:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-arc-blog/azure-local-expands-san-capabilities-with-iscsi-support/ba-p/4531999</guid>
      <dc:creator>saniya0307</dc:creator>
      <dc:date>2026-07-22T22:15:00Z</dc:date>
    </item>
    <item>
      <title>Free Extension: Generate AI Development Prompts from Azure DevOps Work Items — Verity Framework</title>
      <link>https://techcommunity.microsoft.com/t5/azure/free-extension-generate-ai-development-prompts-from-azure-devops/m-p/4539812#M22762</link>
      <description>&lt;P&gt;Hi Azure DevOps community,&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;I wanted to share a free extension we just published to the Visual Studio&lt;/P&gt;&lt;P&gt;Marketplace: the Verity Framework ADO Extension.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;**What it does**&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;It adds a "Verity Prompt" tab to your work items. When you populate five&lt;/P&gt;&lt;P&gt;custom fields on a User Story or Delivery Item, the extension generates a&lt;/P&gt;&lt;P&gt;structured prompt ready to paste into Claude Code, Cursor, or GitHub&lt;/P&gt;&lt;P&gt;Copilot Workspace.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;The five fields:&lt;/P&gt;&lt;P&gt;- VF Intent — the problem and who experiences it&lt;/P&gt;&lt;P&gt;- VF Value — the expected outcome&lt;/P&gt;&lt;P&gt;- VF Appetite — time/resource ceiling (not an estimate)&lt;/P&gt;&lt;P&gt;- VF Trust Criteria — Gate 2 hardening requirements&lt;/P&gt;&lt;P&gt;- VF Failure Scope — Narrow / Moderate / Broad / Systemic&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;**Why it matters**&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Most engineers using AI tools start from a blank context or a vague task&lt;/P&gt;&lt;P&gt;description. The extension carries the team's planning judgment — including&lt;/P&gt;&lt;P&gt;production risk level and hardening requirements — directly into the&lt;/P&gt;&lt;P&gt;implementation context.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;It also includes field validation that flags vague Trust Criteria,&lt;/P&gt;&lt;P&gt;incorrectly formatted Appetite values, and solution-framed Intent&lt;/P&gt;&lt;P&gt;statements before generating the prompt.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;**Setup**&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;About 20 minutes. You add five custom fields to your existing User Story&lt;/P&gt;&lt;P&gt;work item type (no new work item type required). Full instructions are&lt;/P&gt;&lt;P&gt;in the extension tab.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;**Install**&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Search "Verity Framework" on the Visual Studio Marketplace or visit&lt;/P&gt;&lt;P&gt;idearoost.com/verity for the full framework context.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Free. No subscription required.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Happy to answer questions about the field definitions or the setup process.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;— Jacques Steward, IdeaRoost&lt;/P&gt;</description>
      <pubDate>Wed, 22 Jul 2026 18:00:07 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure/free-extension-generate-ai-development-prompts-from-azure-devops/m-p/4539812#M22762</guid>
      <dc:creator>Idearoost1</dc:creator>
      <dc:date>2026-07-22T18:00:07Z</dc:date>
    </item>
    <item>
      <title>MCP in Azure: Using API Management for Authentication, Access, Logging &amp; Governance</title>
      <link>https://techcommunity.microsoft.com/t5/apps-on-azure-blog/mcp-in-azure-using-api-management-for-authentication-access/ba-p/4532139</link>
      <description>&lt;P&gt;If MCP is becoming the “USB-C for AI,” enterprises need something equally familiar on the other side of that port: a&amp;nbsp;&lt;STRONG&gt;control plane&lt;/STRONG&gt;. The Model Context Protocol (MCP) gives AI applications a standard way to discover tools, retrieve context, and call capabilities from external systems. But once you move beyond a single local server and start connecting multiple remote tools across teams, environments, and data domains, the protocol alone is not enough, you also need a way to manage identity, access, traffic, visibility, and lifecycle consistently.&lt;/P&gt;
&lt;P&gt;That is where&amp;nbsp;&lt;STRONG&gt;Azure API Management (APIM)&lt;/STRONG&gt; becomes strategically important. APIM is not “MCP instead of APIs.” It is the layer that can sit&amp;nbsp;in front of&amp;nbsp;MCP servers or even turn existing REST APIs into MCP-compatible tool surfaces so that organizations can apply the same security, policy, and operational controls they already use for business APIs. In other words, MCP standardizes how agents talk to tools; APIM helps standardize how the enterprise governs those tool endpoints at runtime.&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;What MCP actually is and why it changes AI integration&lt;/STRONG&gt;&lt;/H5&gt;
&lt;P&gt;At its core, MCP defines a&amp;nbsp;host / client / server&amp;nbsp;model. The host is the AI application or agent experience, the client manages the MCP connection, and the server exposes capabilities to that client. Those capabilities are commonly expressed as&amp;nbsp;&lt;STRONG&gt;tools&lt;/STRONG&gt;&amp;nbsp;(actions an AI model can call),&amp;nbsp;&lt;STRONG&gt;resources&lt;/STRONG&gt;&amp;nbsp;(context or data an application can read), and&amp;nbsp;&lt;STRONG&gt;prompts&lt;/STRONG&gt;&amp;nbsp;(reusable interaction templates). The same model works whether the server is local or remote, which is one of the reasons MCP is becoming an attractive way to make tool integrations reusable across many AI surfaces.&lt;/P&gt;
&lt;P&gt;That portability matters because most teams do&amp;nbsp;not want to rebuild connector logic for every host, model, or IDE. Microsoft’s Azure guidance explicitly describes two broad ways developers work with MCP: consume existing MCP servers, or build their own when they need custom tools, resources, or prompts.&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;Why MCP sprawl becomes an enterprise problem&lt;/STRONG&gt;&lt;/H5&gt;
&lt;P&gt;MCP is easy to love in a demo because a single server can make an agent instantly smarter. It is harder in production because success creates&amp;nbsp;sprawl. One team wraps an inventory API. Another exposes compliance rules. Another adds DevOps data. Soon, you have multiple remote MCP endpoints, different authentication models, inconsistent rate limits, uneven logging, and no clear answer to “which tools are approved, versioned, and discoverable?” &amp;nbsp;This is the anti-pattern of a&amp;nbsp;monolithic, over-privileged MCP server, and instead one must align server boundaries to business capabilities so access, observability, and ownership remain tractable.&lt;/P&gt;
&lt;P&gt;This is exactly the point at which centralization becomes valuable. Microsoft’s MCP ecosystem organizes the landscape into three layers: a&amp;nbsp;foundation&amp;nbsp;of built-in servers, a&amp;nbsp;build/govern&amp;nbsp;layer that includes Azure Functions, APIM, and API Center, and a&amp;nbsp;consumption layer where agents and runtimes use those tools. That framing is useful because it shows MCP management is not only about building servers; it is equally about establishing how those servers are exposed, governed, and discovered across the enterprise.&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;APIM as the front door for MCP&lt;/STRONG&gt;&lt;/H5&gt;
&lt;P&gt;Azure API Management supports two especially important patterns for MCP centralization. The first is&amp;nbsp;exposing an APIM-managed REST API as a remote MCP server, which lets you turn selected REST operations into MCP tools without rebuilding the backend. The second is&amp;nbsp;putting APIM in front of an existing MCP server, so APIM acts as a facade or proxy that applies policy and governance to a server you already run somewhere else such as Azure Functions, App Service, Container Apps, or another HTTP host.&lt;/P&gt;
&lt;P&gt;This is why APIM is best understood as the&amp;nbsp;&lt;STRONG&gt;front door&lt;/STRONG&gt; rather than the destination. It can route requests, apply transformations, enforce policy, and mediate access while preserving the standardized MCP interaction pattern for clients. MCP server endpoints can be described as “backend APIs” from a governance perspective, which is precisely the enterprise advantage: once a remote MCP endpoint is treated as an API surface, you can bring familiar controls, security, reliability, monitoring, and access policy to an otherwise fast-spreading agent/tool ecosystem.&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;Authentication and authorization: where APIM earns its keep&lt;/STRONG&gt;&lt;/H5&gt;
&lt;P&gt;Remote MCP servers need a strong authorization story because they are not just returning text, they are gateways to systems, data, and actions. MCP authorization guidance uses&amp;nbsp;OAuth 2.1,&amp;nbsp;Protected Resource Metadata (PRM / RFC 9728)&amp;nbsp;for discovery,&amp;nbsp;dynamic client registration (RFC 7591)&amp;nbsp;where supported, and&amp;nbsp;resource indicators (RFC 8707)&amp;nbsp;so tokens remain bound to the intended resource. The practical flow is clear: a client tries the server, gets a&amp;nbsp;401&amp;nbsp;with metadata, discovers the authorization server, requests consent with PKCE and the&amp;nbsp;resource&amp;nbsp;parameter, exchanges for a token, and then includes that bearer token on subsequent calls.&lt;/P&gt;
&lt;P&gt;APIM becomes powerful here because it can be the&amp;nbsp;&lt;STRONG&gt;authorization gateway&lt;/STRONG&gt;&amp;nbsp;in front of the MCP server. Internal security guidance and APIM-focused enablement both converge on the same pattern: use APIM to validate JWTs, apply OAuth/Entra-based access policies, and avoid unsafe token-handling shortcuts such as token passthrough. In practice, this means the caller authenticates against Microsoft Entra ID, APIM validates the token and audience, and only then forwards the call to the MCP backend, optionally with credential mediation patterns such as managed identity, On-Behalf-Of (OBO), or backend credential management depending on the scenario.&lt;/P&gt;
&lt;P&gt;For Azure builders, the authentication options are already familiar. Microsoft’s Foundry guidance for custom MCP tools maps server authentication patterns to&amp;nbsp;key-based auth,&amp;nbsp;Microsoft Entra authentication, and&amp;nbsp;OAuth identity passthrough (OBO). That same guidance explicitly recommends requiring authentication, treating credentials as secrets, applying least privilege to downstream APIs, and logging tool calls for troubleshooting. If you already think in terms of Entra ID, managed identities, audience validation, and OBO, centralizing MCP behind APIM lets you use those same patterns at the agent-tool boundary rather than reinventing them server by server.&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;Logging, monitoring, and observability without losing the plot&lt;/STRONG&gt;&lt;/H5&gt;
&lt;P&gt;Observability is where many MCP deployments discover that a protocol is not the same thing as an operations platform. Internal field guidance is explicit that MCP monitoring today is often&amp;nbsp;product-specific rather than fleet-wide: enterprise MCP usage can show up in Graph activity logs, Foundry uses OpenTelemetry-based tracing through Application Insights, and custom servers hosted in Azure Functions can rely on standard Function/App Insights diagnostics. The practical implication is that you can get strong telemetry, but you still have to stitch it together intentionally rather than expecting a single universal “MCP operations” pane out of the box.&lt;/P&gt;
&lt;P&gt;APIM improves that picture because gateway traffic becomes a consistent choke point for diagnostics. That makes APIM valuable not only for security but also for&amp;nbsp;&lt;STRONG&gt;operational visibility&lt;/STRONG&gt;: which tools are being called, how often, by whom, with which response codes, and where latency or backend errors are emerging.&lt;/P&gt;
&lt;P&gt;There is, however, an important streaming caveat. MCP transports, especially remote ones can depend on&amp;nbsp;&lt;STRONG&gt;streaming-friendly behavior&lt;/STRONG&gt;, and response buffering can break MCP interactions. In other words: yes, centralize observability, but do it in a way that respects streaming transports and does not accidentally make the MCP server stop working while you are trying to observe it.&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;Governance: from “cool tool” to managed platform capability&lt;/STRONG&gt;&lt;/H5&gt;
&lt;P&gt;If authentication answers “who can call this?”, governance answers “what is this server, which version is trusted, and how should anyone discover it?” That is where&amp;nbsp;Azure API Center&amp;nbsp;matters. API Center is centralized location for&amp;nbsp;registering and discovering MCP servers, including those exposed through APIM and those hosted elsewhere. Microsoft’s Foundry guidance similarly describes API Center registration as the path to creating a&amp;nbsp;private organizational tool catalog, where authentication and access settings can be configured consistently and users can discover approved servers rather than trading raw URLs in chat threads.&lt;/P&gt;
&lt;P&gt;This matters more than it sounds. Governance is not only about storage of metadata, it is about&amp;nbsp;versioning, lifecycle, access management, tags, and exposure policy. Recommendation is to version MCP endpoints, document usage patterns, and using API Center to make discovery enterprise-ready.&amp;nbsp;&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;A practical Azure reference pattern&lt;/STRONG&gt;&lt;/H5&gt;
&lt;P&gt;A useful way to think about the architecture is this:&amp;nbsp;&lt;STRONG&gt;MCP hosts and agent runtimes stay at the edge of user experience, while APIM becomes the edge of enterprise control&lt;/STRONG&gt;. The agents or developer tools speak MCP. APIM governs the runtime surface. MCP backends provide tools, and API Center catalogs them. Azure Monitor and Application Insights provide telemetry. Entra ID provides identity. The protocol stays standardized; the operations model becomes centralized.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;Centralizing via APIM vs. letting clients call MCP servers directly&lt;/STRONG&gt;&lt;/H5&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 100%; height: 686.667px; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr style="height: 62.6667px;"&gt;&lt;td style="height: 62.6667px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Dimension&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 62.6667px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Direct access to MCP servers&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 62.6667px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Centralized through APIM&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr style="height: 94.6667px;"&gt;&lt;td style="height: 94.6667px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Authentication model&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 94.6667px;"&gt;
&lt;P&gt;Each server must implement and enforce auth consistently, which can drift across teams.&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 94.6667px;"&gt;
&lt;P&gt;Central token validation, Entra/OAuth policy, and consistent gateway enforcement become possible across servers.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 94.6667px;"&gt;&lt;td style="height: 94.6667px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Access control&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 94.6667px;"&gt;
&lt;P&gt;Per-server controls can be fragmented and harder to audit centrally.&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 94.6667px;"&gt;
&lt;P&gt;Products, subscriptions, access policies, and gateway rules create one place to enforce who can use what.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 94.6667px;"&gt;&lt;td style="height: 94.6667px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Observability&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 94.6667px;"&gt;
&lt;P&gt;Logs depend on each backend platform, and MCP-wide visibility is often fragmented.&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 94.6667px;"&gt;
&lt;P&gt;APIM adds a common telemetry point for traffic, errors, and policy outcomes, while backend telemetry still remains available.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 94.6667px;"&gt;&lt;td style="height: 94.6667px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Reuse of existing APIs&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 94.6667px;"&gt;
&lt;P&gt;Existing REST APIs stay REST APIs unless you build/host a separate MCP layer yourself.&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 94.6667px;"&gt;
&lt;P&gt;APIM can expose existing REST operations as MCP tools, reducing incremental engineering work for standard tool surfaces.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 66.6667px;"&gt;&lt;td style="height: 66.6667px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Registry / discovery&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 66.6667px;"&gt;
&lt;P&gt;Teams often share endpoints ad hoc, which does not scale well organizationally.&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 66.6667px;"&gt;
&lt;P&gt;API Center provides a governed private registry for cataloging, access, and discoverability.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 178.667px;"&gt;&lt;td style="height: 178.667px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Trade-offs&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 178.667px;"&gt;
&lt;P&gt;Fewer moving parts and potentially lower latency for a single, tightly scoped use case.&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 178.667px;"&gt;
&lt;P&gt;Adds a gateway hop and policy/configuration overhead, and APIM’s REST-to-MCP mode focuses on &lt;STRONG&gt;tools&lt;/STRONG&gt; rather than full MCP resources/prompts, so some scenarios still benefit from a dedicated MCP server behind the gateway.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;&lt;STRONG&gt;&amp;nbsp;&lt;/STRONG&gt;&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;Design guidance and best practices&lt;/STRONG&gt;&lt;/H5&gt;
&lt;P&gt;&lt;STRONG&gt;1) Treat MCP servers as business capability boundaries, not as a giant “AI super-endpoint”&lt;/STRONG&gt;&amp;nbsp;Monolithic, over-privileged MCP servers are an anti-pattern; separate servers improve least privilege, ownership clarity, and observability. A focused capability boundary also gives the model clearer tool descriptions, which improves invocation quality and reduces collisions between unrelated tools.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;2) Prefer a no-code APIM path when you already have good REST APIs; use a pro-code MCP server when you need custom orchestration, transformation, or non-API-native logic.&lt;/STRONG&gt;&amp;nbsp;Implement&amp;nbsp;No-Code via APIM&amp;nbsp;for existing REST interfaces and&amp;nbsp;Pro-Code&amp;nbsp;for scenarios requiring more control, custom logic, or deeper integration patterns. This is one of the cleanest ways to make the architecture pragmatic instead of ideological: not every MCP server should be handwritten if the backend capability already exists in API form.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;3) Use Entra ID and avoid token passthrough.&lt;/STRONG&gt;&amp;nbsp;Token passthrough creates audit gaps, bypass risks, and lateral movement opportunities. A better pattern is for APIM to validate and mediate identity at the edge, with audience/resource validation and downstream access configured explicitly. When you need delegated access, use OBO or equivalent identity-aware patterns rather than forwarding arbitrary bearer tokens through the stack unchecked.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;4) Design for observability from day one, but keep streaming in mind.&lt;/STRONG&gt;&amp;nbsp;Log tool invocations, failure paths, latency, and correlation identifiers. Use Azure Monitor / Application Insights on the backend and APIM diagnostics at the gateway. But for remote MCP transports, avoid diagnostics or policy configurations that introduce response buffering and unintentionally break stream-oriented behavior.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;5) Make discovery a governance problem, not a tribal-knowledge problem.&lt;/STRONG&gt;&amp;nbsp;Register internal MCP servers in API Center when you want them discoverable across projects or teams. Use metadata, access management, and lifecycle conventions to decide what is visible, what is approved, and which versions are current. That gives platform teams a catalog to govern and developers a trusted place to discover servers without guessing which endpoint is “the right one.”&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;The bottom line&lt;/STRONG&gt;&lt;/H5&gt;
&lt;P&gt;MCP standardizes how AI applications talk to external capabilities.&amp;nbsp;APIM standardizes how the enterprise protects and operates those capabilities&lt;STRONG&gt;.&lt;/STRONG&gt;&amp;nbsp;Put differently: MCP solves&lt;STRONG&gt; interoperability&lt;/STRONG&gt;, while APIM helps solve&amp;nbsp;&lt;STRONG&gt;operability&lt;/STRONG&gt;. When organizations start treating remote MCP endpoints like first-class API assets with identity, policy, telemetry, registry, and lifecycle attached, tmhey move from isolated agent demos to something that can actually survive production.&lt;/P&gt;
&lt;P&gt;That is why centralizing MCP management in Azure is not really about “putting a gateway in front of servers.” It is about giving your emerging agent ecosystem a&amp;nbsp;&lt;STRONG&gt;governed front door&lt;/STRONG&gt;: one place to authenticate, authorize, observe, version, discover, and evolve the tools your models depend on. And for enterprises already invested in Azure, APIM plus API Center is one of the clearest ways to build that front door without throwing away the APIs and security patterns they already trust.&lt;/P&gt;</description>
      <pubDate>Wed, 22 Jul 2026 16:51:16 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/apps-on-azure-blog/mcp-in-azure-using-api-management-for-authentication-access/ba-p/4532139</guid>
      <dc:creator>Kalaivanan</dc:creator>
      <dc:date>2026-07-22T16:51:16Z</dc:date>
    </item>
    <item>
      <title>CDK Global modernizes automotive CRM on Azure SQL Managed Instance</title>
      <link>https://techcommunity.microsoft.com/t5/customer-innovation-blog/cdk-global-modernizes-automotive-crm-on-azure-sql-managed/ba-p/4539479</link>
      <description>&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;For automotive retailers, &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;most &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;customer relationships &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;don't&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; begin and end with &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;a vehicle purchase&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;. A customer might browse inventory online, visit a dealership weeks later, return for service months after that, and eventually &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;purchase&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; another vehicle years down the road. Every interaction creates information that helps dealerships better understand their customers and build stronger relationships over time.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;Helping dealerships manage those relationships is at the core of the &lt;/SPAN&gt;&lt;/SPAN&gt;&lt;A href="https://www.cdkglobal.com/automotive-crm" target="_blank"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;CDK CRM platform&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;. &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;CDK&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; is a leading provider of cloud-based software to dealerships and OEMs across automotive and related industries in the US and Canada, helping &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;facilitate&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; more than &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;$540 billion&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; in annual automotive commerce. It gives dealership teams visibility into customer &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;interactions with their products and services &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;and creates continuity across the conversations that shape the buying journey. As customer expectations evolve, CDK&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;continu&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;es to look for new ways to&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; help dealerships work more efficiently&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;, &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;make better use of information&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;,&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; and &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;adapt quick&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;ly to market changes.&amp;nbsp;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;Supporting that next phase of innovation required a technology foundation that could &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;grow&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; alongside the business.&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;Working with Microsoft, CDK &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;evol&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;ved&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; its CRM platform on &lt;/SPAN&gt;&lt;/SPAN&gt;&lt;A href="https://azure.microsoft.com/en-us" target="_blank"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;Microsoft Azure&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; and &lt;/SPAN&gt;&lt;/SPAN&gt;&lt;A href="https://azure.microsoft.com/en-us/products/azure-sql/managed-instance" target="_blank"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;Azure SQL Managed Instance&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;.&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;The project included &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;a&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; large migration &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;to &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;Azure SQL Managed Instance and &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;marked &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;an important step&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; in the &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;future&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;of the CRM experience&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;Transforming&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;a business-critical system at scale&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;The size of the project reflected how central the CRM platform is to CDK’s business. The environment supports a broad set of applications for dealership operations and relies on a robust data foundation to keep information flowing.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;Microsoft and CDK worked together to implement a cloud architecture built on &lt;/SPAN&gt;&lt;/SPAN&gt;&lt;A href="https://azure.microsoft.com/en-us/products/app-service" target="_blank"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;Azure App Service&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="auto"&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;and&lt;/SPAN&gt; &lt;/SPAN&gt;&lt;A href="https://azure.microsoft.com/en-us/products/azure-sql/managed-instance" target="_blank"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;Azure SQL Managed Instance&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;.&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; Additional Azure services support application delivery, networking, and data movement across the &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;architecture&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;.&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; The teams&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;developed a Terraform-based infrastructure-as-code framework that standardizes how environments are deployed and managed.&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;That foundation brings greater consistency to the development process. Engineering teams can work within environments that are configured in a predictable way, reducing the likelihood of unexpected differences between testing and production.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;The &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;project&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; also provided an opportunity to strengthen security and governance practices&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;.&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; CDK updated applications to use managed identities, &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;reducing&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; reliance on static credentials&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;, and&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; implemented &lt;/SPAN&gt;&lt;/SPAN&gt;&lt;A href="https://www.microsoft.com/en-us/security/business/identity-access/microsoft-entra-id" target="_blank"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;Microsoft Entra ID&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; authentication and &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;private connectivity patterns that help secure access to critical resources.&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; While d&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;ealership users may &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;not &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;see these changes directly, they help &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;deliver&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; the stability and security that customers expect from the platform&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;Preserving continuity while moving to the cloud&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;As CDK evaluated its&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;cloud&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;strategy, the database layer became one of the most important decisions in the project.&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;The CRM platform supports business-critical dealership operations throughout the day. Customer interactions, sales activity, and service records all depend on information moving quickly and reliably between systems&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;. &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;Any&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;technology transformation&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;would need to preserve that experience while creating a path to future growth.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;Azure SQL Managed Instance stood out because it offered a familiar SQL Server environment while reducing much of the operational overhead associated with managing database infrastructure. &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;It&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; also aligned well with CDK&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;’&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;s long-term goals around resiliency, scalability, and &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;continuous innovation.&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;Equally important, Azure SQL Managed Instance provided a migration path that worked with the realities of the CRM environment.&lt;/SPAN&gt; &lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;“&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;One of the reasons Azure SQL Managed Instance &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;appealed to us&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; was that it allowed us to &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;innovate&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; without redesigning the CRM platform from the ground up,&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;”&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; says &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;Stan &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;Leong&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;, &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;Vice President of Modern Retailing Engineering at CDK&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;. &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;“&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;We could preserve compatibility with the applications our dealerships depend on while taking advantage of a fully managed cloud service.&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;”&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;Unlike some projects that can move applications gradually, &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;the &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;CDK CRM &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;platform required a coordinated transition. The databases that support the platform are highly interconnected, which meant the company needed an approach that would allow the environment to move together while minimizing disruption for customers.&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;To prepare for that transition, CDK and Microsoft used &lt;/SPAN&gt;&lt;/SPAN&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/azure-sql/managed-instance/managed-instance-link-feature-overview" target="_blank"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;Azure SQL Managed Instance link&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; to establish near real-time replication between environments. This allowed teams to begin &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;validating&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;the migration&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; long before the production cutover. Engineers could confirm that data was flowing correctly, &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;identify&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; potential issues, and gain confidence in the process before dealerships were ever affected.&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;The approach also gave CDK an added layer of flexibility during the transition period. Rather than making a one-way move, the company could &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;maintain&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; a rollback option while teams &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;validated&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; production operations in Azure.&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;Because&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;Azure SQL Managed Instance link&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;kept&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;environments&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; synchronized&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;, CDK &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;retained the ability &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;to&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;fail back&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; to its &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;on-premises&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; environment &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;if needed&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;while preserving continuity for dealership operations. After failing over to Azure, the company continued running with synchronized environments for more than four weeks, allowing teams to &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;validate&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; production workloads before completing the final cutover&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;Executing&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; a migration measured in terabytes&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;That preparation became especially important because the migration would take place during a &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;single&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;maintenance &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;window&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;.&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;By &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;establishing&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;synchronization&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; ahead of time, CDK was able to &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;keep &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;data &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;aligned &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;between environments before the&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; failover to &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;Azure&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;.&lt;/SPAN&gt; &lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;W&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;hen migration&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; time&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; arrived, &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;production data was already synchronized in Azure&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;, &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;al&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;l&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;owing&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; the &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;maintenance&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;event &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;to &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;focus&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; on transitioning operations rather than moving large volumes of data for the first time.&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;Planning&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; and architecture work that led to &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;the migration&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; spanned several months. The final migration preparation and validation effort was completed in just six &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;weeks&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;, &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;with e&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;ngineers work&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;ing&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; together to test migration scenarios, &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;optimize&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; replication performance, and &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;validate&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; data consistency ahead of production.&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;CDK migrated more than 1,000 databases and hundreds of terabytes of data to Azure SQL Managed Instance. The environment now processes billions of &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;database&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; queries every day. The &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;failover to Azure&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; was completed during a single weekend with minimal failover time and no data loss. In many cases, database failovers &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;completed&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; in seconds, with most finishing within minutes. Microsoft engineering, product, support, and field teams remained engaged throughout the event, working alongside CDK to &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;monitor&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; the transition and address issues in real time.&amp;nbsp;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;“This was one of the most significant technology initiatives we’ve undertaken for our CRM platform,” says Leong. “Working closely with Microsoft, we migrated more than 1,000 databases to Azure SQL Managed Instance while supporting dealerships throughout the process. The collaboration between our teams helped us execute the transition with minimal disruption to &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;customers.”&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;Strengthening reliabil&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;ity for dealership operations&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;Delivering more than a new cloud environment, the migration also gave CDK an opportunity to &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;evolve &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;the CRM platform operat&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;ions&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;. As part of this effort, CDK implemented resilience architecture built on Azure. The design incorporates &lt;/SPAN&gt;&lt;/SPAN&gt;&lt;A href="https://azure.microsoft.com/en-us/products/frontdoor" target="_blank"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;Azure Front Door&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;, geo-redundant storage, and Azure SQL Managed Instance failover capabilities, allowing the company to maintain continuity when unexpected disruptions occur.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;These improvements rarely attract attention when everything is working as expected, but they help the platform remain available when it matters most. Since the migration, CDK has heard positive feedback from dealerships that report faster and more responsive application experiences.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;Preparing dealerships for &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;what&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;’&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;s&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; next&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;With the CRM platform now running on Azure, CDK is focused on the next phase of its&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; CRM strategy&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;.&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;Beyond supporting &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;today’s&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; dealership operations, the cloud-based foundation gives the company greater flexibility to introduce new &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;functionality&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;and continue &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;advancing&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;the platform over time.&lt;/SPAN&gt; &lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;AI ca&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;pabilities are &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;already&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; help&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;ing&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; dealership employees&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;surface&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; relevant information at the right moment &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;while&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;reducing the effort required to complete routine tasks.&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;A&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;nd A&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;I &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;is&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;accelerat&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;ing&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; software development and deployment, &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;enabling&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; engineering teams to deliver new capabilities more efficiently.&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;CDK is&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;also &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;investing in new reporting experiences designed to make insights easier to access. Dealerships generate enormous amounts of data, but &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;its&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; value depends on how quickly users can find relevant information and act on it&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;The CRM migration has become an important reference point for future &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;platform &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;initiatives&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; across the organization. By successfully moving a business-critical platform at this scale, CDK established a blueprint &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;for&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; future cloud initiatives.&lt;/SPAN&gt; &lt;SPAN data-ccp-parastyle="No Spacing"&gt;“&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;The migration was an important milestone, but &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;it’s&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; really the starting point,&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;”&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; says Leong. &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;“&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;With Azure and Azure SQL Managed Instance, our teams can focus more energy on delivering new capabilities for dealerships&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; while using&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; AI&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; to &lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;power&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt; experiences that help customers operate more efficiently.&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="No Spacing"&gt;”&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 21 Jul 2026 23:21:36 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/customer-innovation-blog/cdk-global-modernizes-automotive-crm-on-azure-sql-managed/ba-p/4539479</guid>
      <dc:creator>RicardoDuncan</dc:creator>
      <dc:date>2026-07-21T23:21:36Z</dc:date>
    </item>
    <item>
      <title>Connecting Microsoft Discovery App to Azure HPC with Azure NetApp Files and CycleCloud</title>
      <link>https://techcommunity.microsoft.com/t5/azure-high-performance-computing/connecting-microsoft-discovery-app-to-azure-hpc-with-azure/ba-p/4539224</link>
      <description>&lt;H1&gt;Introduction&lt;/H1&gt;
&lt;P&gt;Traditional High-Performance Computing (HPC) is central to most scientific computing and engineering applications. These systems are based on well tested, repeatable, and scalable compute processes that enable the large-scale generation and validation that the semiconductor and other industries depend on.&lt;/P&gt;
&lt;P&gt;Integrating AI and agentic flows into this traditional model remains a challenge.&amp;nbsp; Where many agentic systems rely on “modern” compute architecture utilizing REST based communications and object storage, traditional HPC relies on POSIX based file systems and scheduler-based orchestration.&amp;nbsp; Combining these two different disciplines presents a challenge to many engineering and large-scale research enterprises.&lt;/P&gt;
&lt;P&gt;This blog describes how Microsoft Discovery app can directly interoperate with existing traditional HPC deployments utilizing Azure HPC. Microsoft Discovery is an enterprise agentic AI platform for research and development, designed to help specialized agents reason, plan, execute, and learn in a continuous loop across data, tools, and workflows. It is built on Azure and designed to integrate with capabilities such as Azure HPC and Microsoft Foundry, while leveraging industry proven tools and customer flows, making it a natural bridge between AI-native reasoning and compute-intensive engineering execution.&lt;/P&gt;
&lt;H1&gt;Agentic AI in HPC&lt;/H1&gt;
&lt;P&gt;The value of agentic AI in scientific and engineering workloads is not just as a chatbot or code generator; it is the ability to participate in an existing workflow: write parameter files, launch simulations, inspect logs, summarize failures, generate follow-on jobs, and preserve artifacts for traceability. To do this effectively, agents must be able also read data from POSIX-based file systems and dispatch jobs to the scheduler.&amp;nbsp;&lt;/P&gt;
&lt;P&gt;To maintain IP security, they must access data while respecting the access limits of the user that dispatched them and write data under the same RBAC rules as that same user. That way organizations can be assured that agents operate strictly under the dispatching user's access rights, so they cannot reach data the user isn't authorized to see and the data they write retains the permissions that the user who dispatches the agent enjoys.&lt;/P&gt;
&lt;P&gt;Discovery app is currently only available on Windows.&amp;nbsp; Linux and Mac versions will soon be available.&amp;nbsp; In the meantime, enabling the Windows based Discovery app to leverage the capabilities on Azure HPC today with little to no change to the current HPC flow.&lt;/P&gt;
&lt;H1&gt;Reference architecture&lt;/H1&gt;
&lt;P&gt;The architecture has five primary components: a Windows VM running the Discovery app on Azure, an Azure NetApp Files volume that is mounted by both Windows and Linux clients, an Azure CycleCloud-managed HPC cluster with a Linux login node used for job submission and monitoring, and a scheduler that dispatches work to compute nodes.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Discovery app on Windows:&lt;/STRONG&gt; Runs on a Windows VM placed inside the same VNet, or reachable through controlled private networking, so it can access storage and the login node without exposing the HPC environment publicly.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure NetApp Files:&lt;/STRONG&gt; Provides the shared file namespace for input decks, scripts, logs, generated data, and results. Linux nodes mount the volume with native NFS. Windows can mount the same namespace through the Windows NFS client or an equivalent enterprise file-access solution.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure CycleCloud:&lt;/STRONG&gt; Provisions and manages the HPC cluster, including scheduler, login node, execute nodes, autoscaling, and cluster configuration.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Login node:&lt;/STRONG&gt; Acts as the controlled command endpoint. Discovery should submit jobs and query scheduler state from here, not run heavy tools directly on the login node. This protects the actual Cyclecloud scheduler and cluster manager from being overwhelmed if the dispatch volume is high. This also allows for scalability if there are many clients scheduling jobs at once.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Scheduler:&lt;/STRONG&gt; Incumbent schedulers that customers are already using with Azure Cyclecloud handles execution. Discovery creates scripts and submits them through the scheduler so compute runs on the correct nodes with the correct environment, licenses, and policies. This remains the same as how traditional HPC schedules jobs today.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H1&gt;Connectivity model&lt;/H1&gt;
&lt;P&gt;Running on a Windows VM in the Azure VNet aligns with many enterprises’ requirement that none of their data or IP can leave their network. Doing this avoids the need to run jobs through a user laptop or unmanaged public endpoint. This also allows the user to start an agentic workload then disconnect from the VM while the job runs and reconnect when they come back. Maintaining the same working model most engineers have with their HPC environment today.&lt;/P&gt;
&lt;P&gt;For Discovery to access Azure HPC, it needs to map the Azure NetApp Files volumes and command access to the Linux login node. File access lets agents create and read files. Command access lets agents run Linux commands, submit jobs, and monitor execution. Discovery writes execution scripts to the ANF volume, then uses SSH to dispatch the job to the schedule with the execution script.&lt;/P&gt;
&lt;P&gt;The advantage of this model is that the user can directly provide Unix-style paths to Discovery and tell it to execute tools without the user worrying about being on a Windows platform.&amp;nbsp; For example, CAD teams often provide setup scripts with paths to tools that users are supposed to reference in creating their own tool execution scripts.&amp;nbsp; Using this model, I just tell Discovery to use the setup script /mount/setup/tools/vendora.csh as a template and it will happily find the file and do so.&amp;nbsp; I can then tell Discovery to dispatch the job to the hpc queue and it will schedule to the appropriate SLURM queue where a system is provisioned and assigned to run using the new execution script Discovery creates.&lt;/P&gt;
&lt;P&gt;This model also gets around the temporary limitation of Discovery app, unlike the enterprise Discovery service, only being available on Windows.&amp;nbsp; In this model, Windows is only acting as a terminal interface so it doesn’t matter that it’s not a platform that normally works with most EDA tools. In fact, users can deploy Discovery app on either x86 or ARM Windows since all the actual tools run on the Linux-based HPC deployment in Azure anyway.&lt;/P&gt;
&lt;H1&gt;Mapping Windows drive to NFS volumes&lt;/H1&gt;
&lt;P&gt;Discovery agents must understand how to access files on the mapped NFS volume from the Windows VM. The same file may appear as &lt;STRONG&gt;Z:\project\run1\exec.csh&lt;/STRONG&gt; on Windows and &lt;STRONG&gt;/mount/project/run1/exec.csh&lt;/STRONG&gt; on Linux. The skill or instruction file used by Discovery should explicitly describe that mapping and instruct agents to use Linux paths inside batch scripts and scheduler commands.&lt;/P&gt;
&lt;P&gt;A typical instruction might say: the Windows drive &lt;STRONG&gt;Z:&lt;/STRONG&gt; maps to the Linux mount point &lt;STRONG&gt;/mount&lt;/STRONG&gt;. If Discovery sees &lt;STRONG&gt;/mount/designs/test.sv&lt;/STRONG&gt;, it can read or write the corresponding Windows file at &lt;STRONG&gt;Z:\designs\test.sv&lt;/STRONG&gt;. When generating job scripts, it should always use the Linux path, because those scripts run on the cluster.&lt;/P&gt;
&lt;H1&gt;Implementation steps&lt;/H1&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;1. Deploy the HPC foundation&lt;/H2&gt;
&lt;P&gt;Start with the standard Azure HPC foundation: a virtual network, Azure CycleCloud, a scheduler-backed cluster, and shared storage. CycleCloud can deploy and manage Slurm clusters and supports external NFS mounts, including Azure NetApp Files. Configure the cluster so the scheduler and compute nodes mount the shared file system at a stable Linux path such as &lt;STRONG&gt;/mount&lt;/STRONG&gt;, &lt;STRONG&gt;/shared&lt;/STRONG&gt;, or a customer-specific project path.&lt;/P&gt;
&lt;P&gt;Make sure that systems can ssh to the login node using a public/private key pair.&amp;nbsp; The login node should be accessible using a command like:&lt;/P&gt;
&lt;P&gt;% ssh -i .ssh/keyfile user@login-node&lt;/P&gt;
&lt;H2&gt;2. Place the Windows VM in the same network boundary&lt;/H2&gt;
&lt;P&gt;Create a Windows VM that can reach both the Azure NetApp Files endpoint and the CycleCloud login node over private IP. The required Windows NFS client is available on Windows 11 Pro, Enterprise, and Education. It is not available on Windows 11 Home. Install the Discovery app on that VM. This VM becomes the user’s Discovery workstation for the HPC-connected workflow. Keep the VM inside the same security boundary as the HPC deployment, and use standard enterprise controls such as private networking, restricted inbound access, managed identity where appropriate, and least-privilege SSH keys.&lt;/P&gt;
&lt;H2&gt;3. Install Microsoft Discovery app&lt;/H2&gt;
&lt;OL&gt;
&lt;LI&gt;Download and install Microsoft Discovery app from:&lt;BR /&gt;&lt;A href="https://github.com/microsoft/discovery" target="_blank" rel="noopener"&gt;https://github.com/microsoft/discovery&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Once installed, start Microsoft Discovery app and connect to your Github Copilot account in order to configure access to the LLM AI models.&lt;/LI&gt;
&lt;LI&gt;Create a new project&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2&gt;4. Mount the Azure NetApp Files volume into Windows&lt;/H2&gt;
&lt;P&gt;Mount the same shared storage namespace into Windows. The Windows NFS client can map the NFS export to a drive letter such as &lt;STRONG&gt;Z:&lt;/STRONG&gt;. The Windows NFS is sufficient for reading logs, results, and data files, and writing execution scripts, parameters, and configuration files. Heavy simulation I/O remain on Linux compute nodes running HPC tools.&lt;/P&gt;
&lt;P&gt;To configure the Windows NFS Client, you will need to have administrator privileges on your Windows VM.&lt;/P&gt;
&lt;P&gt;To mount the ANF volume with the Windows NFS client.&amp;nbsp;&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Press Win+R&lt;/LI&gt;
&lt;LI&gt;Run “optional features”&lt;/LI&gt;
&lt;LI&gt;Select “More Windows features”&lt;/LI&gt;
&lt;LI&gt;Expand “Services for NFS”&lt;/LI&gt;
&lt;LI&gt;Enable “Administrative Tools” and “Client for NFS”&lt;BR /&gt;&lt;img /&gt;&lt;/LI&gt;
&lt;LI&gt;Click “OK” to install the components.&lt;/LI&gt;
&lt;LI&gt;By default, Windows will mount the NFS volume as an anonymous user. To set the correct user, you can define the correct Linux User ID and Group ID in the registry.&lt;BR /&gt;&lt;BR /&gt;Note: This is an example of a functional baseline to mount a Linux volume with a user identity as an example.&amp;nbsp; Production environments should employ proper identity mapping for security. &lt;BR /&gt;&lt;BR /&gt;&lt;/LI&gt;
&lt;LI&gt;Press Win+R&lt;/LI&gt;
&lt;LI&gt;Run “regedit” to start the Registry Editor&lt;/LI&gt;
&lt;LI&gt;Navigate to HKEY_LOCAL_MACHINE\SOFTWARE\Microsoft\ClientForNFS\CurrentVersion\Default&lt;BR /&gt;&lt;img /&gt;&lt;/LI&gt;
&lt;LI&gt;In Default, right click and add two new DWORD entires: AnonymousUID and AnonymousGID&lt;BR /&gt;&lt;img /&gt;&lt;/LI&gt;
&lt;LI&gt;Set the values to the decimal value of your Linux UID and GID. This will ensure that you connect to the ANF volume as the correct user.&lt;BR /&gt;&lt;img /&gt;&lt;/LI&gt;
&lt;LI&gt;Select OK and then launch the Command Prompt as Administrator.&lt;/LI&gt;
&lt;LI&gt;To allow the NFS Client to accept the new registry entries, restart the NFS Client with the following commands:&lt;BR /&gt;nfsadmin client stop&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;nfsadmin client start&lt;BR /&gt;&lt;BR /&gt;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;OL start="15"&gt;
&lt;LI&gt;Mount the ANF volume to a Windows drive:&lt;BR /&gt;&lt;img /&gt;&lt;/LI&gt;
&lt;LI&gt;You should now be able to access your NFS volume by going through the Z: drive.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2&gt;5. Configure SSH access to the login node&lt;/H2&gt;
&lt;P&gt;Configure the SSH client on the Windows VM with a dedicated key for Discovery-driven access. The key should connect to the CycleCloud login node as the user (same as with NFS above) with permissions to submit jobs, read scheduler state, and access the shared project directory. Verify connectivity by connecting to the Cyclecloud login node using ssh from the Command Prompt.&amp;nbsp;&lt;/P&gt;
&lt;P&gt;In this example, the username is ai4semi, the IP address of the login node is 10.16.16.25, and the cc_rsa is the keyfile located under the .ssh directory of the user’s home directory on Windows. Since this is the first time for this system to log into the Linux machine, you must accept the connection registration.&amp;nbsp; This will not be necessary for subsequent ssh connects by Discovery using this method.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;5. Teach Discovery how to use Azure HPC&lt;/H2&gt;
&lt;P&gt;Now that you have the NFS and SSH connections established in Windows, you simply need to tell Discovery how to use these connections properly to run HPC workloads. In the Discovery chat window, tell Discovery to map the Z: drive on windows to the /mount volume on NFS and how to interact with the files.&lt;/P&gt;
&lt;P&gt;The Windows drive `Z:` is mapped to the cluster NFS volume `/mount`.&lt;/P&gt;
&lt;P&gt;- `/mount/&amp;lt;path&amp;gt;` on the cluster corresponds to `z:\&amp;lt;path&amp;gt;` locally.&lt;/P&gt;
&lt;P&gt;- Example: when a file is referenced as `/mount/test.txt`, look at `z:\test.txt`.&lt;/P&gt;
&lt;P&gt;Test to make sure Discovery understand by asking it to read a file on the NFS volume using POSIX nomenclature.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; read the contents of /mount/newfile.txt&lt;/P&gt;
&lt;P&gt;Discovery should show you the file contents&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Next, instruct Discovery on how to access and interact with Cyclecloud&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; Connect to the CycleCloud login node over SSH using "ssh -i .ssh\cc_rsa ai4semi@10.16.16.25"&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;Then tell Discovery which scheduler Cyclecloud uses and tell it to check the queue.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; Cyclecloud is using SLURM as the scheduler. Check the status of the hpc queue.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;Since the Windows VM is only used to run Discovery and the agentic AI part of the workload and the login node is not meant to run actual HPC workloads, be explicit with Discovery on how it should treat HPC job runs.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;NEVER run EDA/build/simulation tools on the local system or directly on the login node. ALWAYS submit work to the Slurm scheduler via sbatch.&lt;/P&gt;
&lt;P&gt;- The login node is for job submission and monitoring only — not for compute.&lt;/P&gt;
&lt;P&gt;- Wrap every tool invocation in a batch script and submit it with sbatch.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;6. Validate the end-to-end workflow&lt;/H2&gt;
&lt;P&gt;Validate that everything works by dispatching a simple test script to Cyclecloud&lt;/P&gt;
&lt;P&gt;Create a batch script that references the Linux path, submit it to Cyclecloud.&lt;/P&gt;
&lt;P&gt;dispatch /mount/proj/ai4semi/exec.csh to the hpc queue&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;7. Save the knowledge for future reference&lt;/H2&gt;
&lt;P&gt;Now that Discovery understands how to access Linux-based files on NFS and how to dispatch jobs, tell it to save the information.&lt;/P&gt;
&lt;P&gt;Save this knowledge to a skill and use it for all subsequent job runs and file access for the project&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;You see that Discovery now saved the information on how to access files and run jobs to a SKILL.md for general knowledge and updated the copilot-instructions.md to ensure that all subsequent agents and engines understand how to access Azure HPC.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Note: These instructions are only scoped to a single project.&amp;nbsp; When creating a new project, you can simply tell Discovery to reach over to your older project to copy the knowledge:&lt;/P&gt;
&lt;P&gt;copy over instructions on how to access NFS files and paths, and how to dispatch and execute hpc jobs from workspace1&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;If you want to make this a portable instruction that you can just load for each new project, tell Discovery to create a portable file.&lt;/P&gt;
&lt;P&gt;package the instructions on how to access NFS files and paths, and how to dispatch and execute hpc jobs into a a single instruction I can provide to new workspaces&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;You can then copy the file to a central location like: C:\Users\ai4semi\project&lt;/P&gt;
&lt;P&gt;Then for each new project you can tell Discovery to use this file.&lt;/P&gt;
&lt;P&gt;read and understand the instructions in C:\Users\ai4semi\project\cyclecloud-hpc.instructions.md. Use for all agents and engines in this project and save this information as skills to use in the project&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H1&gt;Operational considerations&lt;/H1&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Identity and permissions:&lt;/STRONG&gt; Keep SSH access scoped to the project and avoid broad administrative privileges. For NFS, verify UID/GID mapping and file ownership behavior before production use.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Performance:&lt;/STRONG&gt; Use the Windows mount for orchestration artifacts, scripts, and logs. Let Linux compute nodes perform heavy I/O through native ANF mounts.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Security:&lt;/STRONG&gt; Keep Discovery, storage, and the login node on private networking. Avoid public exposure of scheduler or storage endpoints.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Scheduler hygiene:&lt;/STRONG&gt; Enforce a rule that Discovery submits jobs through the scheduler and does not run heavy tools interactively on the login node.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H1&gt;Conclusion&lt;/H1&gt;
&lt;P&gt;Connecting Microsoft Discovery app to Azure HPC allows users to seamlessly use their existing HPC environment. The most practical approach today is to place Discovery on a Windows VM inside the Azure network, map the shared Azure NetApp Files namespace into both Windows and Linux, and use SSH to submit scheduler-managed jobs through CycleCloud.&lt;/P&gt;
&lt;P&gt;Using this method, customers can quickly and easily use Discovery to integrate Agentic AI into their current HPC flows. Users can take advantage of AI agents to read and understand specification files, generate code and testbenches based on those specifications, run existing HPC tools on their Azure HPC environment. They can direct Discovery to use the scripts and resources which their CAD teams have developed. Users can create bookshelves by pointing to the documentation directories many tool vendors include in their installations, allowing Discovery to take advantage of documentation which most users don’t have time to fully digest.&lt;/P&gt;
&lt;P&gt;Over time, native Linux support can reduce the need for Windows-hosted bridging. Until then, this architecture gives teams a secure, and repeatable way to let Discovery agents create, launch, monitor, and learn from real HPC workloads on Azure.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 22 Jul 2026 16:47:17 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-high-performance-computing/connecting-microsoft-discovery-app-to-azure-hpc-with-azure/ba-p/4539224</guid>
      <dc:creator>richpaw</dc:creator>
      <dc:date>2026-07-22T16:47:17Z</dc:date>
    </item>
    <item>
      <title>Microsoft Discovery: Where HPC meets agentic AI for the next era of EDA</title>
      <link>https://techcommunity.microsoft.com/t5/azure-high-performance-computing/microsoft-discovery-where-hpc-meets-agentic-ai-for-the-next-era/ba-p/4539212</link>
      <description>&lt;H1&gt;Introduction&lt;/H1&gt;
&lt;P&gt;High-performance computing is the engine behind the most demanding engineering breakthroughs. In electronic design automation, HPC enables massive simulation farms, verification regressions, place-and-route exploration, timing closure, power analysis, and signoff workloads that would be impractical at scale. However, as chip complexity grows, the limiting factor is no longer only compute capacity. It is the ability to reason across enormous design spaces, coordinate specialized tools, preserve engineering context, and decide what to run next.&lt;/P&gt;
&lt;P&gt;That is where &lt;A href="https://azure.microsoft.com/en-us/solutions/discovery/" target="_blank" rel="noopener"&gt;Microsoft Discovery&lt;/A&gt; becomes especially interesting. Microsoft Discovery is an enterprise agentic AI platform for research and development, designed to help specialized agents reason, plan, execute, and learn in a continuous loop across data, tools, and workflows. It is built on Azure and designed to interoperate with capabilities such as Azure HPC and Microsoft Foundry, while leveraging industry proven tools and customer flows, making it a natural bridge between AI-native reasoning and compute-intensive engineering execution.&lt;/P&gt;
&lt;H1&gt;The Inflection Point: HPC is required, but not enough&lt;/H1&gt;
&lt;P&gt;Engineering teams already know how to scale compute. They run large regressions, distribute workloads across clusters, burst into cloud capacity, and optimize storage and schedulers around fast-moving design cycles. Yet the harder challenge is often deciding which simulations matter, which failures are related, which tool settings deserve another experiment, and which signals should trigger the next branch of exploration.&lt;/P&gt;
&lt;P&gt;Traditional HPC gives engineering teams scale. Agentic AI adds accelerated analysis, orchestration, memory, reasoning, and adaptation. The synthesis of the two is not about replacing engineers or replacing EDA tools. It is about turbocharging the engineers by creating an intelligent execution fabric where agents can understand goals, inspect results, choose tools, launch jobs, summarize outcomes, and recommend the next best action. Engineers remain in control of the design intent and final decisions. In this model, the most valuable resource is not compute capacity; it is engineering time.&lt;/P&gt;
&lt;H1&gt;Microsoft Discovery as the agentic layer for engineering R&amp;amp;D&lt;/H1&gt;
&lt;P&gt;Microsoft Discovery focuses on the full R&amp;amp;D lifecycle: knowledge reasoning, code and plan generation, simulation, analysis, and iteration. For semiconductor and EDA workflows, that maps naturally to the way design teams already operate. Timing closure, regression failure, achieving performance targets all require that engineers pour over massive amounts of information from tool run logs, interpret the results based on engineering goals and specifications, and determine how to resolve the gaps.&amp;nbsp; Discovery leverages the same scientific process used in scientific research to help engineers more quickly understand the resulting data, pick out critical learnings, and help compose the next steps and mitigations.&lt;/P&gt;
&lt;P&gt;Unlike chat interfaces, agents can act and execute specific functions. Agents can be assigned roles: a verification triage agent, a log-analysis agent, a simulation-planning agent, a physical-design exploration agent, a cost-and-capacity agent, or a documentation agent that preserves evidence and rationale. Together, these agents can coordinate around a shared objective and use HPC as the execution and validation vehicle.&lt;/P&gt;
&lt;P&gt;For semiconductor programs, engineers are the most expensive and constrained part of the process. Every hour spent searching logs, reconciling reports, relaunching routine jobs, copying context between systems, or documenting repetitive findings is an hour not spent on architecture, debug strategy, design tradeoffs, or design sign-off. Discovery’s value is therefore not simply that it can automate tasks; it can help shift engineering effort away from menial coordination work and toward the decisions that improve product outcomes and silicon spins.&lt;/P&gt;
&lt;H1&gt;What this looks like for EDA&lt;/H1&gt;
&lt;P&gt;In a complex silicon project, design teams may run thousands of tests across multiple scenarios, configurations, simulators, and coverage targets. Today, engineers spend significant time reading logs, correlating failures, rerunning jobs, and deciding whether a failure is new, known, flaky, or blocking. Much of that work is necessary, but repetitive. With an agentic workflow, a verification agent could monitor regression output, cluster related failures, inspect logs, identify likely root causes, and propose targeted reruns rather than brute-force rerunning everything. All this helps to provide engineers more time to focus on the work that requires their unique engineering skill and judgement.&lt;/P&gt;
&lt;P&gt;In physical design, an agentic workflow could help explore constraints, placement strategies, timing violations, congestion hot spots, and power tradeoffs. HPC provides the parallel capacity to run experiments. Discovery-style agent orchestration can help determine which experiments to run, capture why they were run, compare outcomes, and refine the next set of candidates. This provides a guided design-space exploration loop.&lt;/P&gt;
&lt;P&gt;For signoff and analysis, agents could help connect results across timing, power, reliability, manufacturability, and cost. Instead of treating each report as an isolated artifact, an agentic system can reason across reports, prior runs, known design patterns, and engineering guidance helping to converge faster.&lt;/P&gt;
&lt;P&gt;Discovery leverages tried-and-true, industry-proven tools to validate the output of AI using the same methods semiconductor teams have relied on for decades. AI can propose code and plans, prioritize validation and simulation, summarize analysis and debug, or recommend next steps, but the validation path still runs through established EDA flows that engineers understand and trust.&lt;/P&gt;
&lt;H1&gt;The Architecture: Agents, Tools, Data, and Compute&lt;/H1&gt;
&lt;P&gt;Discovery’s flow architecture leverages a closed-loop system. Knowledge sources provide context: design specs, prior bugs, regression history, tool documentation, scripts, and engineering notes. Agents reason over that context and formulate next steps. Tool integrations connect those agents to EDA applications, schedulers, storage systems, model endpoints, and analysis pipelines. Azure HPC supplies scalable execution for simulations and compute-heavy analysis. Results flow back into the system so the next decision is better informed than the last.&lt;/P&gt;
&lt;P&gt;This demonstrates the power of joining agentic AI and HPC. Discovery provides a platform model for coordinating that loop with enterprise expectations around security, governance, transparency, and human oversight. HPC tools provide assurance that the results from AI are valid.&lt;/P&gt;
&lt;H1&gt;Integrating Discovery with Azure HPC for EDA&lt;/H1&gt;
&lt;P&gt;Discovery becomes even more powerful when it is connected to Azure HPC infrastructure purpose-built for compute- and memory-intensive engineering workloads. For EDA, that includes AMD-based Azure HPC virtual machines such as HX-series instances, which are optimized for silicon design and memory-intensive workloads, with large memory capacity, AMD EPYC processors with 3D V-Cache, and high-performance InfiniBand networking. These platforms give agentic workflows the execution fabric needed to run large simulations, regressions, analysis jobs, and design-space exploration at cloud scale for both EDA and other HPC workloads.&lt;/P&gt;
&lt;P&gt;Storage is equally important. EDA workloads generate and consume enormous numbers of files, and performance often depends on low-latency shared file access as much as raw CPU capacity. Azure NetApp Files provides enterprise-grade file storage for mission-critical HPC and EDA workloads in Azure, giving teams a managed, high-performance storage layer that can support simulation farms, tool installations, design libraries, scratch spaces, and shared project data.&lt;/P&gt;
&lt;P&gt;Azure NetApp Files also enables customers to use Discovery in a hybrid environment, enabling customers to keep critical data on-prem while using Discovery to optimize cloud-based flows. Many semiconductor teams keep authoritative design data, source repositories, IP libraries, and sign-off environments in on-premises infrastructure. Azure NetApp Files cache volumes can help bridge that workflow by creating a cloud-based cache of active data from an external origin volume. Frequently accessed data can be served close to Azure compute, while the authoritative dataset remains in the existing on-premises design environment.&lt;/P&gt;
&lt;P&gt;Cluster automation is another important part of enabling the hybrid flow. Azure CycleCloud can help automate the creation, scaling, and lifecycle management of HPC clusters in Azure so teams can stand up cloud capacity around specific EDA workloads rather than manually managing static infrastructure. That is important for agentic workflows because Discovery can reason about what needs to run, determine the proper resources needed, while CycleCloud-backed automation helps ensure the right compute environment is available when those jobs are ready to execute.&lt;/P&gt;
&lt;P&gt;Just as importantly, hybrid adoption does not need to force customers to replace the scheduling systems their engineering organizations already use. Many semiconductor teams have deep operational investments in incumbent schedulers which CycleCloud already works with, established queues, policies, licenses, scripts, and user workflows. A practical Discovery-plus-Azure-HPC architecture helps bring agentic AI to their existing flows: integrating cloud capacity with existing scheduler patterns, extending familiar job-submission models, and allowing teams to burst selected workloads into Azure while preserving the operational controls that have governed EDA execution for years.&lt;/P&gt;
&lt;P&gt;In that architecture, Discovery can act as the intelligent orchestration layer across the hybrid estate as well as a productivity accelerator. Agents can reason over design intent, prior results, and engineering context; submit jobs into Azure HPC environments; use AMD-powered compute for demanding EDA runs; access hot data through Azure NetApp Files; and work through cluster automation and scheduler integrations that align with customer operating models. Combined with Azure HPC, Azure NetApp Files helps&amp;nbsp;eliminate&amp;nbsp;storage bottlenecks that can otherwise limit EDA job throughput, allowing compute resources, EDA licenses, and engineering teams&amp;nbsp;to remain&amp;nbsp;productive at scale. The result is a practical synthesis: AI-guided exploration, proven EDA validation, elastic cloud compute, hybrid data access, and scheduler-aware execution that respects the way semiconductor teams already work.&lt;/P&gt;
&lt;H1&gt;Why it matters for semiconductor teams&lt;/H1&gt;
&lt;P&gt;Semiconductor design is a systems problem. Teams must manage more IP, more verification complexity, more software interaction, more foundry requirements, and more pressure to maintain schedules while design sizes and complexity grows. The industry has invested heavily in automation, but much of that automation remains fragmented across scripts, dashboards, job schedulers, and individual expert workflows. The simple fact is that compute is expensive, but engineering time is even more expensive. Improving engineer productivity has an outsized impact on schedule, quality, and cost.&lt;/P&gt;
&lt;P&gt;Agentic AI offers a way to automation more of the design and verification process.&amp;nbsp; Automating much of the more mechanical parts of the flow. Teams can create workflows that observe, reason, and iterate while agents coordinate the results to achieve the engineering objective. Instead of relying only on tribal knowledge, teams can preserve decisions, evidence, and rationale in a repeatable workflow.&lt;/P&gt;
&lt;H1&gt;Human-in-the-loop by design&lt;/H1&gt;
&lt;P&gt;The goal is not autonomous chip design without engineers. The goal is amplified engineering outcomes. In the most valuable scenarios, agents handle the repetitive, high-volume, evidence-gathering work while engineers define objectives, validate assumptions, approve major decisions, and interpret tradeoffs. Discovery can help reduce the menial burden of searching, summarizing, comparing, rerunning, and documenting so engineers can spend more time on the creative and analytical work only they can do. This is especially important in EDA, where correctness, traceability, and sign-off confidence matter as much as speed.&lt;/P&gt;
&lt;P&gt;Microsoft Discovery’s emphasis on enterprise governance and transparency is therefore central to the story. In regulated or high-stakes engineering environments, teams need to know what data was used, which tools were invoked, what assumptions were made, and where human approval occurred. Just as important, they need confidence that AI-generated recommendations are validated through established engineering practices, not accepted on faith. Agentic workflows must be auditable, explainable, and grounded in the same EDA verification and signoff discipline the industry already depends on.&lt;/P&gt;
&lt;H1&gt;Conclusion: Going beyond just compute and AI&lt;/H1&gt;
&lt;P&gt;The next phase of EDA acceleration will not come from compute alone, and it will not come from AI alone. It will come from the synthesis of both: HPC infrastructure that can execute at scale, and agentic AI systems that can reason, coordinate, and learn across complex engineering workflows.&lt;/P&gt;
&lt;P&gt;Microsoft Discovery steps towards that future. For HPC and EDA teams, the opportunity is to move from faster batch execution to intelligent convergence: workflows that know what has been tried, understand what changed, recommend what to try next, and use scalable compute to validate ideas quickly through proven EDA tools and methodologies. In a world where design complexity keeps rising, the ability to combine elastic compute with agentic systems can help to greatly improve productivity, turn repetitive process work into guided automation, and validate AI-assisted decisions through the same trusted engineering practices the semiconductor industry has used for decades.&lt;/P&gt;
&lt;P&gt;The Microsoft Discovery team will be at &lt;A href="https://dac.com/2026" target="_blank" rel="noopener"&gt;Design Automation Conference&lt;/A&gt;, July 27-29, 2026, in Long Beach, California. The Microsoft Discovery team will be showcasing how engineers can use agentic AI to enable greater productivity for semiconductor design workloads. Drop by the Microsoft booth (651) to learn more about Microsoft Discovery for semiconductor design.&amp;nbsp; Learn how technologies from AMD and NetApp come together with Azure HPC to provide the high-performance environment for running EDA tools on the cloud. &amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 22 Jul 2026 15:25:57 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-high-performance-computing/microsoft-discovery-where-hpc-meets-agentic-ai-for-the-next-era/ba-p/4539212</guid>
      <dc:creator>richpaw</dc:creator>
      <dc:date>2026-07-22T15:25:57Z</dc:date>
    </item>
    <item>
      <title>Take command of your Microsoft Foundry AI agents with the Azure Copilot Observability Agent</title>
      <link>https://techcommunity.microsoft.com/t5/azure-observability-blog/take-command-of-your-microsoft-foundry-ai-agents-with-the-azure/ba-p/4538681</link>
      <description>&lt;P&gt;Generative AI (GenAI) agents are moving from prototypes to production, running as part of real distributed applications. When a Microsoft Foundry agent slows down, returns errors, or spikes in token usage, teams need to know why - fast, and with evidence. &lt;STRONG&gt;The Azure Copilot Observability Agent,&lt;/STRONG&gt; generally available and powered by Azure Monitor, &lt;STRONG&gt;makes Foundry and GenAI agent investigations a first-class scenario. &lt;/STRONG&gt;You can investigate agent failures, latency, tool-call errors, and backend saturation through the same signals that run your Azure estate.&lt;/P&gt;
&lt;H2&gt;Why AI observability matters in production&lt;/H2&gt;
&lt;P&gt;Building an agent is one thing. Running a fleet of them at scale, in front of real users, is another. AI applications also introduce behaviors that traditional monitoring wasn't built to explain:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Non-deterministic outputs&lt;/STRONG&gt; - the same prompt can take a different path through the agent's logic.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;New failure modes&lt;/STRONG&gt; - hallucinations, tool-call failures, and reasoning that silently goes off track.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Cost and latency that shift with behavior&lt;/STRONG&gt; - tokens and response time depend on what the agent decides to do, which models are used and so on.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Fleet-scale complexity&lt;/STRONG&gt; - many agents, across frameworks and hosting, each part of a larger system.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;AI observability closes this gap. You need to observe your AI workflows, understand their health, and investigate when an agent is failing, slow, or behaving in ways users notice.&lt;/P&gt;
&lt;H2&gt;From telemetry to analysis and investigation&lt;/H2&gt;
&lt;P&gt;&lt;STRONG&gt;Microsoft Foundry&lt;/STRONG&gt; is your home for building and evaluating agents, &lt;STRONG&gt;while Azure Monitor makes those agents a first-class item you can observe through traces, tool calls, dependencies, token usage, and evaluation signals - across frameworks and hosting environments. &lt;/STRONG&gt;Once you instrument your agents, their telemetry flows into Azure Monitor through Application Insights, and the Observability Agent reasons over it there. &lt;SPAN style="color: rgb(30, 30, 30);"&gt;Setting this up is a straightforward process: &lt;/SPAN&gt;instrument the agent, confirm telemetry is landing, then work with it in Azure Monitor. The &lt;A href="https://learn.microsoft.com/en-us/azure/azure-monitor/app/agents-view" target="_blank"&gt;Monitor AI Agents with Application Insights &lt;/A&gt;walks you through the experience and lets you jump directly from an agent's Monitoring tab in Foundry to Azure Monitor, and&amp;nbsp;&lt;A href="https://techcommunity.microsoft.com/blog/azureobservabilityblog/new-capabilities-to-observe-agents-in-azure-monitor/4524896" target="_blank" rel="noopener"&gt;New capabilities to observe agents in Azure Monitor&lt;/A&gt;&amp;nbsp;covers the wider telemetry pipeline behind it.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Telemetry tells you&amp;nbsp;&lt;EM&gt;what&lt;/EM&gt; happened&lt;/STRONG&gt;,&amp;nbsp;&lt;STRONG&gt;Analysis helps you make sense of it&lt;/STRONG&gt;, and &lt;STRONG&gt;investigations tell you&amp;nbsp;&lt;EM&gt;why&lt;/EM&gt;&lt;/STRONG&gt;. Most days start with analytics: you can&amp;nbsp;&lt;A href="https://learn.microsoft.com/azure/azure-monitor/aiops/observability-agent-chat" target="_blank" rel="noopener"&gt;chat with your observability data&lt;/A&gt; in natural language - ask for an error summary, a latency overview, or dependency health, then follow up to explore a pattern and decide whether something warrants a deeper look. So, the Observability Agent isn't only for investigations; it's also how you analyze and explore your agent telemetry day to day. When that exploration surfaces a real problem and you need the &lt;EM&gt;why&lt;/EM&gt;, it escalates into &lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/azure-monitor/aiops/observability-agent-deep-investigations" target="_blank"&gt;deep investigations&lt;/A&gt; and root cause analysis for GenAI scenarios. &lt;STRONG&gt;The Azure Copilot Observability Agent reasons across your telemetry - logs, metrics, traces, alerts, dependencies, Foundry telemetry, and changes - together with Azure resource context, discovered topology, and relevant instructions.&lt;/STRONG&gt; It frames hypotheses, gathers evidence, rules out weak explanations, and shows the reasoning behind its findings. Because it runs inside Azure Monitor, it inherits the access controls you already use - chat and investigations run with the signed-in user's identity and Azure role-based access control (RBAC), and your prompts and responses aren't used to train foundation models.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;For Foundry and GenAI agents, you can investigate issues including agent failures and errors, latency, tool-call failures and failing Model Context Protocol (MCP) or backend dependencies, token spikes, and regressions after deployments - prioritizing the failure, latency, and throttling patterns that matter most.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;It's powered by the same investigation model the Observability Agent applies to your applications, Azure Kubernetes Service (AKS), and virtual machines (VMs). Autonomous operations, in public preview, can analyze alerts in the background, correlate related alerts, and create Azure Monitor issues with findings and next steps - preparing context and reducing triage work, with humans in control.&lt;/P&gt;
&lt;H3&gt;Deep investigation through the Azure stack&lt;/H3&gt;
&lt;P&gt;One of the Observability Agent's key strengths is that it doesn't stop at the agent boundary. &lt;STRONG&gt;It investigates&amp;nbsp;&lt;EM&gt;through&lt;/EM&gt; the Azure stack&lt;/STRONG&gt; - from the agent to its dependencies and tool calls, down to the infrastructure behind them.&lt;/P&gt;
&lt;P&gt;Take a common &lt;STRONG&gt;Foundry &lt;/STRONG&gt;pattern:&amp;nbsp;&lt;STRONG&gt;agent failures caused by backend SQL saturation.&lt;/STRONG&gt; When a Foundry GenAI agent starts failing, the root cause can go beyond the model or the prompt. The agent scopes to the affected agent, its dependencies and tool calls, and the SQL backend; correlates the failures with SQL timeouts and a spike in failed dependency calls; rules out alternatives like cache issues or a broader outage; and identifies a short SQL saturation window often driven by a high-cost query as the cause. For more patterns, read the &lt;A href="https://learn.microsoft.com/azure/azure-monitor/aiops/observability-agent-deep-investigation-examples#agent-failures-caused-by-backend-sql-saturation" target="_blank" rel="noopener"&gt;deep investigation examples for the Azure Copilot Observability Agent&lt;/A&gt;.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Watch it in action:&lt;/STRONG&gt;&amp;nbsp;&lt;/P&gt;
&lt;div data-video-id="https://www.youtube.com/watch?v=yDklikC7zwI/1784465244626" data-video-remote-vid="https://www.youtube.com/watch?v=yDklikC7zwI/1784465244626" class="lia-video-container lia-media-is-center lia-media-size-large"&gt;&lt;iframe src="https://cdn.embedly.com/widgets/media.html?src=https%3A%2F%2Fwww.youtube.com%2Fembed%2FyDklikC7zwI%3Ffeature%3Doembed&amp;amp;display_name=YouTube&amp;amp;url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3DyDklikC7zwI&amp;amp;image=https%3A%2F%2Fi.ytimg.com%2Fvi%2FyDklikC7zwI%2Fhqdefault.jpg&amp;amp;type=text%2Fhtml&amp;amp;schema=youtube" allowfullscreen="" style="max-width: 100%"&gt;&lt;/iframe&gt;&lt;/div&gt;
&lt;H3&gt;A closer look: diagnosing an agent over a time range&lt;/H3&gt;
&lt;P&gt;To see how this plays out, point the Observability Agent at a specific agent and a time range say, the last 24 hours, and ask it to diagnose operation and performance. Instead of a single answer, it works through the problem the way an experienced operator would, across three lines of reasoning.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Failure patterns.&lt;/STRONG&gt; It first establishes whether failures trace back to one dominant cause or several unrelated ones, and identifies which &lt;STRONG&gt;tools&lt;/STRONG&gt; and dependencies are failing most and with what error and &lt;STRONG&gt;GenAI error &lt;/STRONG&gt;types. It separates&amp;nbsp;&lt;EM&gt;expected, app-level&lt;/EM&gt; failures - validation errors, 400s and 404s - from&amp;nbsp;&lt;EM&gt;platform and reliability&lt;/EM&gt; failures like 5xx responses, timeouts, and dropped connections, so you don't spend time on noise. For each top failure type, it points to the likely locus: Main agent, sub-agent, tool, or a downstream service shows the supporting evidence, and proposes the most actionable fix.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Latency and slow runs.&lt;/STRONG&gt; When a run is slow, it traces the end-to-end flow - entry point, agent steps, model and tool calls, and downstream dependencies - and shows where the time actually goes. If you give it a specific run, it walks that transaction step by step; if you don't, it looks across runs to &lt;STRONG&gt;find where latency concentrates: The large language model (LLM) call, a tool call, orchestration, or a downstream dependency,&lt;/STRONG&gt; and tells you whether the slowness is systemic or isolated.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Throttling and quota pressure.&lt;/STRONG&gt; It confirms whether you're seeing throttling signals - 429s, 503s, timeouts, and backoff patterns - and attributes them to the most likely source: the model or provider, the Foundry runtime, or another dependency. It also checks whether retries and backoff are helping or actually amplifying the problem.&lt;/P&gt;
&lt;P&gt;The result isn't a dashboard you have to interpret. It's a prioritized, evidence-backed read on what's failing, what's slow, and what's being throttled - with the reasoning shown at each step, so you can act on it or dig further.&lt;/P&gt;
&lt;img /&gt;
&lt;H2&gt;Get started&lt;/H2&gt;
&lt;P&gt;The Azure Copilot Observability Agent is generally available in Azure Monitor, and autonomous operations are available in public preview.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Instrument your Foundry and GenAI agents, then explore them in the&amp;nbsp;&lt;A href="https://learn.microsoft.com/azure/azure-monitor/app/agents-view" target="_blank" rel="noopener"&gt;agents view in Azure Monitor&lt;/A&gt;.&lt;/LI&gt;
&lt;LI&gt;See what's possible once agents are instrumented in&amp;nbsp;&lt;A href="https://techcommunity.microsoft.com/blog/azureobservabilityblog/new-capabilities-to-observe-agents-in-azure-monitor/4524896" target="_blank" rel="noopener"&gt;New capabilities to observe agents in Azure Monitor&lt;/A&gt;.&lt;/LI&gt;
&lt;LI&gt;Learn how to use the&amp;nbsp;&lt;A href="https://learn.microsoft.com/azure/azure-monitor/aiops/observability-agent-overview" target="_blank" rel="noopener"&gt;Azure Copilot Observability Agent&lt;/A&gt;.&lt;/LI&gt;
&lt;LI&gt;Explore how&amp;nbsp;&lt;A href="https://learn.microsoft.com/azure/azure-monitor/aiops/observability-agent-deep-investigations" target="_blank" rel="noopener"&gt;deep investigations in the Azure Copilot Observability Agent&lt;/A&gt;&amp;nbsp;work.&lt;/LI&gt;
&lt;LI&gt;Read the announcement:&amp;nbsp;&lt;A href="https://techcommunity.microsoft.com/blog/azureobservabilityblog/azure-copilot-observability-agent-is-generally-available-with-autonomous-operati/4528213" target="_blank" rel="noopener"&gt;Azure Copilot Observability Agent is generally available, with autonomous operations in preview&lt;/A&gt;.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;What's next for Foundry and GenAI agents&lt;/H2&gt;
&lt;P&gt;The Observability Agent already reasons about the scenarios that matter most for GenAI agents today: It investigates agent failures, latency, and performance, and it also understands token economy and works with evaluation signals. Here's where we're investing next:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Advanced token analysis&lt;/STRONG&gt; - clearer visibility into where tokens go, which runs are most expensive, and how model and prompt choices affect cost.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Deeper failure attribution&lt;/STRONG&gt; - pinpoint whether an issue comes from agent logic, model behavior, or a backend dependency, so you troubleshoot the right layer.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Richer evaluation correlation&lt;/STRONG&gt; - easier ways to tie quality regressions to the specific runs, prompts, and changes behind them.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;These additions extend the Observability Agent's existing strengths while keeping the same governed, evidence-first, human-in-control model you rely on today.&lt;/P&gt;
&lt;H2&gt;Stay connected&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Follow this blog for ongoing deep dives, updates on current capabilities, and a preview of what's coming next.&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Live webinar -&lt;/STRONG&gt;&amp;nbsp;a walkthrough of real Observability Agent scenarios, best practices, and what's available today - along with a look at what's coming next, and live Q&amp;amp;A with the product team.&amp;nbsp;&lt;A href="https://forms.office.com/r/XYAarZvFte" target="_blank" rel="noopener"&gt;Register for the Observability Agent webinar&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;We'd love your feedback&lt;/H2&gt;
&lt;P&gt;The Observability Agent continues to evolve based on real-world usage and operator feedback. Share your thoughts directly through the&amp;nbsp;&lt;STRONG&gt;Give Feedback&lt;/STRONG&gt; option in the experience or reach us at&amp;nbsp;&lt;A href="mailto:azureobsagent@microsoft.com" target="_blank" rel="noopener"&gt;rofrenke@microsoft.com&lt;/A&gt;.&lt;/P&gt;</description>
      <pubDate>Tue, 21 Jul 2026 16:28:48 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-observability-blog/take-command-of-your-microsoft-foundry-ai-agents-with-the-azure/ba-p/4538681</guid>
      <dc:creator>Ron Frenkel</dc:creator>
      <dc:date>2026-07-21T16:28:48Z</dc:date>
    </item>
    <item>
      <title>Azure Architecture Best Practices for Enterprise Applications</title>
      <link>https://techcommunity.microsoft.com/t5/azure/azure-architecture-best-practices-for-enterprise-applications/m-p/4538994#M22755</link>
      <description>&lt;P data-slot-rendered-content="true"&gt;As businesses continue to modernize their IT infrastructure, Microsoft Azure has become one of the leading cloud platforms for building enterprise-grade applications. Organizations of all sizes rely on Azure to improve scalability, strengthen security, reduce operational costs, and accelerate digital transformation. However, simply moving an application to the cloud does not guarantee success. A well-designed Azure architecture is the foundation of a secure, reliable, and high-performing enterprise application.&lt;/P&gt;
&lt;P data-slot-rendered-content="true"&gt;&lt;A class="lia-external-url" href="https://dellenny.com/azure-architecture-best-practices-for-enterprise-applications/?utm_source=linkedin&amp;amp;utm_medium=social&amp;amp;utm_campaign=organic_post" target="_blank"&gt;https://dellenny.com/azure-architecture-best-practices-for-enterprise-applications/&lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Mon, 20 Jul 2026 14:28:26 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure/azure-architecture-best-practices-for-enterprise-applications/m-p/4538994#M22755</guid>
      <dc:creator>JohnNaguib</dc:creator>
      <dc:date>2026-07-20T14:28:26Z</dc:date>
    </item>
    <item>
      <title>Join us for our MCP Live! — A free livestream covering all things MCP</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-developer-community/join-us-for-our-mcp-live-a-free-livestream-covering-all-things/ba-p/4537980</link>
      <description>&lt;P&gt;The &lt;A class="lia-external-url" href="https://modelcontextprotocol.io/" target="_blank"&gt;Model Context Protocol (MCP)&lt;/A&gt; is the open standard for connecting AI models to tools and data. It was first introduced to the world by Anthropic in November 2024, and is now the most widely adopted standard in the world of Generative AI. You can now use MCP servers to connect agents to your data sources across multiple clients: VS Code, Claude Desktop, Codex, Copilot App, Goose, and many more. Plus, MCP is the most common way to build your own agents that connect to internal enterprise tools, like when using Microsoft agent-framework or Langchain.&lt;/P&gt;
&lt;P&gt;To celebrate the success of MCP and its rich ecosystem, we are hosting MCP Live! on September 9th, from 9AM to 1PM PT. We'll hear both from the teams at Microsoft that are implementing MCP, but also from core MCP maintainers and the global community of MCP developers.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Register here:&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-teams="true"&gt;&lt;A href="https://aka.ms/MCPLive/99/b" aria-label="Link https://aka.ms/MCPLive/99/b" target="_blank"&gt;https://aka.ms/MCPLive/99/b&lt;/A&gt;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Here's our current planned agenda:&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 100%; height: 350px; border-width: 1px;"&gt;&lt;colgroup&gt;&lt;col style="width: 9.36052%" /&gt;&lt;col style="width: 31.0436%" /&gt;&lt;col style="width: 59.5959%" /&gt;&lt;/colgroup&gt;&lt;tbody&gt;&lt;tr style="height: 35px;"&gt;&lt;td style="height: 35px;"&gt;&lt;EM&gt;Time&lt;/EM&gt;&lt;/td&gt;&lt;td style="height: 35px;"&gt;&lt;EM&gt;Topic&lt;/EM&gt;&lt;/td&gt;&lt;td style="height: 35px;"&gt;&lt;EM&gt;Speakers&lt;/EM&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 35px;"&gt;&lt;td style="height: 35px;"&gt;9 AM PT&lt;/td&gt;&lt;td style="height: 35px;"&gt;&lt;STRONG&gt;Welcome to MCP Live!&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 35px;"&gt;Liam Hampton and Pamela Fox (Microsoft/GitHub), MCP advocates&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 35px;"&gt;&lt;td style="height: 35px;"&gt;9:15 AM&lt;/td&gt;&lt;td style="height: 35px;"&gt;&lt;STRONG&gt;State of MCP&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 35px;"&gt;Caitie McCaffrey (Microsoft), MCP core maintainer&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 35px;"&gt;&lt;td style="height: 35px;"&gt;9:45 AM&amp;nbsp;&lt;/td&gt;&lt;td style="height: 35px;"&gt;&lt;STRONG&gt;MCP at GitHub&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 35px;"&gt;Sam Morrow (GitHub), GitHub MCP server maintainer&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 35px;"&gt;&lt;td style="height: 35px;"&gt;10:05 AM&amp;nbsp;&lt;/td&gt;&lt;td style="height: 35px;"&gt;&lt;STRONG&gt;MCP at Microsoft&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 35px;"&gt;Linda Li (Microsoft), Foundry Toolbox PM&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 35px;"&gt;&lt;td style="height: 35px;"&gt;10:45 AM&lt;/td&gt;&lt;td style="height: 35px;"&gt;&lt;STRONG&gt;Building MCP servers in VS Code&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 35px;"&gt;VS Code Engineering Team&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 35px;"&gt;&lt;td style="height: 35px;"&gt;11:30 AM&lt;/td&gt;&lt;td style="height: 35px;"&gt;&lt;STRONG&gt;MCP Authorization: DCR, EMA, and more!&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 35px;"&gt;MCP authorization maintainers (To be announced!)&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 35px;"&gt;&lt;td style="height: 35px;"&gt;12:15 AM&lt;/td&gt;&lt;td style="height: 35px;"&gt;&lt;STRONG&gt;MCP Apps&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 35px;"&gt;Jeremiah Lowin (Prefect), FastMCP maintainer&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 35px;"&gt;&lt;td style="height: 35px;"&gt;12:30 AM&lt;/td&gt;&lt;td style="height: 35px;"&gt;&lt;STRONG&gt;MCP Tasks, Events, Triggers&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 35px;"&gt;Clare Liguori (Amazon), MCP core maintainer&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 35px;"&gt;&lt;td style="height: 35px;"&gt;12:45 AM&lt;/td&gt;&lt;td style="height: 35px;"&gt;&lt;STRONG&gt;Closing&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 35px;"&gt;Liam Hampton and Pamela Fox (Microsoft/GitHub), MCP advocates&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;We're very excited to bring together so many folks working on both the MCP specification and on production MCP servers, so we can all learn about the latest MCP features and best practices together.&lt;/P&gt;
&lt;H3&gt;Learn more MCP&lt;/H3&gt;
&lt;P&gt;Brand new to MCP? We still want you to join! If you want, you can learn MCP fundamentals from our free resources:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://aka.ms/mcp-for-beginners" target="_blank"&gt;MCP-for-beginners&lt;/A&gt; : A step-by-step written tutorial covering all the MCP features&lt;/LI&gt;
&lt;LI&gt;
&lt;P&gt;&lt;A class="lia-external-url" href="https://aka.ms/pythonagents/rewatch" target="_blank"&gt;Python + MCP series&lt;/A&gt; : Recordings of a 3-part livestream series focusing on building MCP servers with Python&lt;/P&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;Meet the MCP community&lt;/H3&gt;
&lt;P&gt;Want to meet the MCP community in person? We're also planning several IRL events:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://globalai.community/e/bay9vh24" target="_blank"&gt;San Francisco (14 Sep 2026)&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://globalai.community/e/bd1o37ln" target="_blank"&gt;Bengaluru (26 Sep 2026)&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Hope to see you in the live chat on September 9th!&lt;/P&gt;</description>
      <pubDate>Mon, 20 Jul 2026 07:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-developer-community/join-us-for-our-mcp-live-a-free-livestream-covering-all-things/ba-p/4537980</guid>
      <dc:creator>Pamela_Fox</dc:creator>
      <dc:date>2026-07-20T07:00:00Z</dc:date>
    </item>
    <item>
      <title>Round Table: Building Browser-Capable Agents with the Browser Automation Tool</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-developer-community/round-table-building-browser-capable-agents-with-the-browser/ba-p/4538581</link>
      <description>&lt;P&gt;Some of the most valuable work still lives inside a browser: booking a class, pulling a figure off a dashboard, filling in a portal form, gathering research across a dozen tabs. These are exactly the tasks people wish an agent could just&amp;nbsp;&lt;EM&gt;do&lt;/EM&gt;. On &lt;STRONG&gt;22 July 2026 at 2:30 PM BST (7:00 PM IST)&lt;/STRONG&gt;, the Microsoft Foundry community is running a 40‑minute &lt;STRONG&gt;Discord round table&lt;/STRONG&gt; on the &lt;STRONG&gt;Browser Automation tool&lt;/STRONG&gt; : how how it helps agents complete real browser workflows, where you see risk or friction, and what samples, docs, and product improvements would help you adopt it.&lt;/P&gt;
&lt;P&gt;This is a discussion, not a slideshow. Bring your real projects : the the web workflows you'd love to hand off, and the guardrails you'd want first. Join us in the &lt;A href="https://aka.ms/foundry/discord" target="_blank" rel="noopener"&gt;Microsoft Foundry Discord community&lt;/A&gt;. Please arrive at the scheduled time for a quick tech check.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;Event at a glance&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;What:&lt;/STRONG&gt; Microsoft Foundry Discord Community Round Table : &lt;EM&gt;Building Browser-Capable Agents with the Browser Automation Tool&lt;/EM&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;When:&lt;/STRONG&gt; 22 July 2026, 2:30 PM BST / 7:00 PM IST (40 minutes)&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Where:&lt;/STRONG&gt; &lt;A href="https://aka.ms/foundry/discord" target="_blank" rel="noopener"&gt;https://aka.ms/foundry/discord&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Event link&amp;nbsp;
&lt;P&gt;&lt;A class="lia-external-url" href="https://discord.gg/Z8JZsrP5P5?event=1527676149264679013" target="_blank" rel="noopener"&gt;https://discord.gg/Z8JZsrP5P5?event=1527676149264679013&lt;/A&gt;&amp;nbsp;&lt;/P&gt;
&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Format:&lt;/STRONG&gt; Interactive discussion : voice and chat, live polls, and a short prioritisation exercise voice and chat, live polls, and a short prioritisation exercise&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Who it's for:&lt;/STRONG&gt; AI engineers and developers building agents that need to act on the web&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Opening question we'll start with:&lt;/STRONG&gt; "What browser-based task would you love an AI agent to automate for you today?"&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;The problem: the last mile of automation still runs in a browser&lt;/H2&gt;
&lt;P&gt;Most real-world workflows eventually hit a website with no clean API , such as a supplier portal, an internal admin console, a booking page, or a legacy dashboard a supplier portal, an internal admin console, a booking page, a legacy dashboard. Traditional scripting can automate these, but selectors break, pages change, and every new site means another brittle script to maintain. What developers actually want is an agent that can &lt;EM&gt;look at a page, decide what to do, and do it&lt;/EM&gt; : navigate, read, click, type, and hand back a structured result navigate, read, click, type, and hand back a structured result.&lt;/P&gt;
&lt;P&gt;That's the gap the &lt;A href="https://learn.microsoft.com/azure/foundry/agents/how-to/tools/browser-automation" target="_blank" rel="noopener"&gt;Browser Automation tool&lt;/A&gt; in Microsoft Foundry is built to close , and doing it responsibly and doing it responsibly, with the right safeguards, is a big part of why we want your feedback.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;What is the Browser Automation tool?&lt;/H2&gt;
&lt;P&gt;The &lt;STRONG&gt;Browser Automation Tool (BAT)&lt;/STRONG&gt; gives Foundry agents the ability to drive a real browser to complete web workflows. It's available as an &lt;STRONG&gt;MCP tool&lt;/STRONG&gt;, and it uses &lt;A href="https://aka.ms/pww/docs" target="_blank" rel="noopener"&gt;Playwright Workspaces&lt;/A&gt; , a generally available, cloud-scale service, a generally available, cloud-scale service : navigating, clicking at coordinates, typing, and applying filters as its headless browser infrastructure. When an agent gets a request, Foundry spins up an &lt;STRONG&gt;isolated, sandboxed browser session&lt;/STRONG&gt; per interaction, so each run is private and segregated.&lt;/P&gt;
&lt;H3&gt;How agents actually interact with a page&lt;/H3&gt;
&lt;P&gt;BAT runs a perception–action loop. The model receives the current state of the page (including screenshots), decides the next action, and BAT executes it in the sandbox using Playwright &amp;nbsp;and real oversight navigating, clicking at coordinates, typing, applying filters. After each action, BAT captures the updated state and sends it back to the model, repeating until the goal is met or the user stops. Because the model can parse HTML into a DOM, it can reason about the page rather than follow a fixed script. It also supports &lt;STRONG&gt;multi-turn conversations&lt;/STRONG&gt;, so you can refine a request mid-flow to complete form-filling or scraping scenarios.&lt;/P&gt;
&lt;H3&gt;Built for real use : watch the automation happen in real time for debugging. and real oversight&lt;/H3&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Live View&lt;/STRONG&gt; : a human-in-the-loop override for ambiguous or sensitive steps. watch the automation happen in real time for debugging.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Take Control&lt;/STRONG&gt; : each interaction gets its own sandboxed browser. a human-in-the-loop override for ambiguous or sensitive steps.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Isolated sessions&lt;/STRONG&gt; : for reliability, optimisation, and audit. each interaction gets its own sandboxed browser.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Built-in observability&lt;/STRONG&gt; : for internal systems (private preview). for reliability, optimisation, and audit.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Private website browsing&lt;/STRONG&gt; : Python, C#, JavaScript, Java, and the REST API. for internal systems (private preview).&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Broad SDK support&lt;/STRONG&gt; , and the agent can make mistakes or be misled by malicious page content Python, C#, JavaScript, Java, and the REST API.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;A word on responsible use.&lt;/STRONG&gt; BAT is powerful precisely because an AI can use credentials you share with it to reach email, financial, enterprise, or social accounts , watching for &amp;nbsp;and the agent can make mistakes or be misled by malicious page content. You're responsible for reviewing your applications, scoping which credentials you provide, and adding your own mitigations. See the &lt;A href="https://learn.microsoft.com/azure/foundry/responsible-ai/agents/transparency-note" target="_blank" rel="noopener"&gt;Foundry Agent Service transparency note&lt;/A&gt;. This is exactly the kind of trade-off we want to talk through together.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;Example scenario&lt;/H2&gt;
&lt;P&gt;A user asks: &lt;EM&gt;"Report the year-to-date percent change of Microsoft's stock price."&lt;/EM&gt; The agent navigates to a finance site, enters MSFT in the search bar, opens the stock page, clicks the &lt;STRONG&gt;YTD&lt;/STRONG&gt; view on the chart, reads the value, and returns a clean, structured answer ; that's the &amp;nbsp;no bespoke scraper, no hard-coded selectors, and a full trace of what it did.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Discussion prompt:&lt;/STRONG&gt; "Where would browser automation fit into your current projects or workflows?"&lt;/P&gt;
&lt;H2&gt;How setup works (the short version)&lt;/H2&gt;
&lt;P&gt;You'll want to understand the wiring before you scale, so it's worth a look ahead of the session. There are two moving parts:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Create a Playwright Workspace&lt;/STRONG&gt; in the Azure portal, enable the access token auth method, and grab the wss:// browser endpoint. Give your project identity a Contributor (or custom) role on the workspace.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Connect the tool in Foundry&lt;/STRONG&gt; under &lt;STRONG&gt;Build &amp;gt; Tools&lt;/STRONG&gt;: create a toolbox, add &lt;STRONG&gt;Browser Automation&lt;/STRONG&gt;, point it at your Playwright workspace and auth type, and publish. Copy the &lt;STRONG&gt;Project connection ID&lt;/STRONG&gt; from the tool's details page : how agents interact with the web, example use cases, and responsible use. that's the BROWSER_CONNECTION_ID in your code.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;What we'll cover in the 40 minutes&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Welcome &amp;amp; opening question (0:00–0:03)&lt;/STRONG&gt; : navigate, gather, interact, and return a structured result, end to end. the browser task you'd most love to automate.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;What is Browser Automation (0:03–0:07)&lt;/STRONG&gt; : the workflows you're building, public vs. internal targets, and where you'd pick automation over scripting. how agents interact with the web, example use cases, and responsible use.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Scenario walkthrough (0:07–0:12)&lt;/STRONG&gt; : what agents may do autonomously, what needs approval, and the observability and enterprise safeguards you'd require. navigate, gather, interact, return a structured result : your biggest adoption blockers, missing docs, and the SDK samples and demos you'd prioritise. end to end.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Use cases &amp;amp; opportunities (0:12–0:22)&lt;/STRONG&gt; : vote live on top use cases, challenges, and feature requests. the workflows you're building, public vs. internal targets, and where you'd pick automation over scripting.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Trust, security &amp;amp; governance (0:22–0:31)&lt;/STRONG&gt; , and which would benefit most from a capable agent. what agents may do autonomously, what needs approval, and the observability and enterprise safeguards you'd require.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Developer experience feedback (0:31–0:36)&lt;/STRONG&gt; ; which should always require approval. your biggest adoption blockers, missing docs, and the SDK samples and demos you'd prioritise.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Prioritisation &amp;amp; next steps (0:36–0:40)&lt;/STRONG&gt; , sometimes with credentials, vote live on top use cases, challenges, and feature requests.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;Come prepared to talk about&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;The browser-based workflows you're building today : navigation, data gathering, form filling, and research, via an MCP tool powered by Playwright Workspaces. and which would benefit most from a capable agent.&lt;/LI&gt;
&lt;LI&gt;Whether your scenarios target public websites, internal systems, or both.&lt;/LI&gt;
&lt;LI&gt;Why you'd choose browser automation over traditional scripting.&lt;/LI&gt;
&lt;LI&gt;Which actions you'd let an agent perform autonomously : a perception-action loop with screenshots and DOM parsing handles pages that break brittle scripts. and which should always require approval.&lt;/LI&gt;
&lt;LI&gt;The observability, audit, and enterprise safeguards you'd expect before running this in production.&lt;/LI&gt;
&lt;LI&gt;The examples, samples, and tutorials that would help you get started fastest.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;Responsible and secure by design&lt;/H2&gt;
&lt;P&gt;Because BAT lets an agent take real actions on live websites : isolated sessions, Live View, Take Control, and observability for reliability and audit. sometimes with credentials : scope credentials carefully and add your own mitigations; the tool is powerful and in preview. governance is a first-class part of the conversation, not a footnote. Isolated per-session sandboxes, Live View, Take Control human-in-the-loop, and built-in observability are there so you can see, pause, and audit what an agent does. Bring your trust concerns, required guardrails, and governance requirements: they directly shape the roadmap.&lt;/P&gt;
&lt;P&gt;&lt;EM&gt;Note:&lt;/EM&gt; the Browser Automation tool is in &lt;STRONG&gt;preview&lt;/STRONG&gt;; APIs and capabilities may change, and it isn't recommended for production workloads yet.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;Key takeaways&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Browser Automation&lt;/STRONG&gt; lets Foundry agents complete real web workflows : this round table feeds directly into the engineering and product teams. navigation, data gathering, form filling, research , and arrive on time for the tech check. via an MCP tool powered by Playwright Workspaces.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Agents reason, not just replay&lt;/STRONG&gt; , and browser-capable agents are how we cross it. a perception–action loop with screenshots and DOM parsing handles pages that break brittle scripts.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Oversight is built in: isolated sessions&lt;/STRONG&gt;, Live View, Take Control, and observability for reliability and audit.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Responsibility is shared: scope credentials&lt;/STRONG&gt; carefully and add your own mitigations; the tool is powerful and in preview.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Your feedback shapes the product: this round table&lt;/STRONG&gt; feeds directly into the engineering and product teams.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;Save your spot&lt;/H2&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;Add it to your calendar: 22 July 2026, 2:30 PM BST / 7:00 PM IST, and arrive&lt;/STRONG&gt; on time for the tech check.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Join the community:&lt;/STRONG&gt; &lt;A href="https://aka.ms/foundry/discord" target="_blank" rel="noopener"&gt;https://aka.ms/foundry/discord&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Prep with the sample:&lt;/STRONG&gt; explore the &lt;A href="https://github.com/microsoft-foundry/foundry-samples/tree/main/samples/csharp/hosted-agents/agent-framework/browser-automation" target="_blank" rel="noopener"&gt;browser automation sample&lt;/A&gt; in foundry-samples.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Read the docs:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/azure/foundry/agents/how-to/tools/browser-automation" target="_blank" rel="noopener"&gt;Automate browser tasks with Foundry agents&lt;/A&gt; and the &lt;A href="https://learn.microsoft.com/azure/foundry/agents/how-to/tools/browser-automation-hosted-agent-quickstart" target="_blank" rel="noopener"&gt;hosted-agent quickstart&lt;/A&gt;.&lt;/LI&gt;
&lt;LI&gt;Event registration link&amp;nbsp;
&lt;P&gt;&lt;A class="lia-external-url" href="https://discord.gg/Z8JZsrP5P5?event=1527676149264679013" target="_blank" rel="noopener"&gt;https://discord.gg/Z8JZsrP5P5?event=1527676149264679013&lt;/A&gt;&amp;nbsp;&lt;/P&gt;
&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;The last mile of automation still runs in a browser, and browser-capable agents are how we cross it. Come tell us what you'd automate, where you'd draw the line, and what you'd need to trust it in production. See you on &lt;STRONG&gt;22 July&lt;/STRONG&gt;.&lt;/P&gt;</description>
      <pubDate>Sat, 18 Jul 2026 22:12:25 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-developer-community/round-table-building-browser-capable-agents-with-the-browser/ba-p/4538581</guid>
      <dc:creator>Lee_Stott</dc:creator>
      <dc:date>2026-07-18T22:12:25Z</dc:date>
    </item>
    <item>
      <title>Microsoft Foundry Now Has an AI Gateway Control Plane — What Changes for App Service</title>
      <link>https://techcommunity.microsoft.com/t5/apps-on-azure-blog/microsoft-foundry-now-has-an-ai-gateway-control-plane-what/ba-p/4538320</link>
      <description>&lt;P&gt;In May, I published a &lt;A class="lia-internal-link lia-internal-url lia-internal-url-content-type-blog" href="https://techcommunity.microsoft.com/blog/appsonazureblog/you-can-build-a-framework-agnostic-ai-gateway-on-azure-app-service-%E2%80%94-heres-how/4522004" target="_blank" rel="noopener" data-lia-auto-title="runnable sample that put Azure API Management in front of an AI agent on Azure App Service" data-lia-auto-title-active="0"&gt;runnable sample that put Azure API Management in front of an AI agent on Azure App Service&lt;/A&gt;. APIM handled token limits, semantic caching, token metrics, and access to the model. The point was simple: keep the agent framework interchangeable and put the production controls at the gateway.&lt;/P&gt;
&lt;P&gt;A recent Apps on Azure post, &lt;A class="lia-internal-link lia-internal-url lia-internal-url-content-type-blog" href="https://techcommunity.microsoft.com/blog/appsonazureblog/from-ai-adoption-to-ai-governance---using-apim-as-the-gateway-for-azure-ai-found/4536247" target="_blank" rel="noopener" data-lia-auto-title="From AI Adoption to AI Governance" data-lia-auto-title-active="0"&gt;From AI Adoption to AI Governance&lt;/A&gt;, arrived at the same architectural conclusion from a platform-governance angle. It uses APIM to turn model identity and token consumption into a per-model chargeback signal for Microsoft Foundry workloads.&lt;/P&gt;
&lt;P&gt;That is useful validation, but the more interesting update is in the product itself: &lt;STRONG&gt;Microsoft Foundry can now create or associate an APIM-based AI Gateway directly from its Admin console.&lt;/STRONG&gt; The data path is still APIM. What changed is who can configure and govern it.&lt;/P&gt;
&lt;P&gt;This post explains what that means for an App Service-hosted agent, what Foundry now manages, what still belongs in APIM, and what I would change in the sample we built earlier.&lt;/P&gt;
&lt;H2&gt;The short version: same gateway, new control plane&lt;/H2&gt;
&lt;P&gt;The architecture is still familiar:&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Client -&amp;gt; App Service agent -&amp;gt; APIM AI Gateway -&amp;gt; Foundry model&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;APIM remains the traffic gateway. It authenticates callers, applies policies, forwards requests, and emits telemetry. App Service remains the application runtime for the agent and its APIs.&lt;/P&gt;
&lt;P&gt;The new piece is the Foundry control plane. From &lt;STRONG&gt;Operate &amp;gt; Admin console &amp;gt; AI Gateway&lt;/STRONG&gt;, a platform team can now:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Create a new APIM instance or associate a supported existing instance.&lt;/LI&gt;
&lt;LI&gt;Enable the gateway for individual Foundry projects.&lt;/LI&gt;
&lt;LI&gt;Apply independent token limits to projects sharing the gateway.&lt;/LI&gt;
&lt;LI&gt;Verify gateway traffic and inspect logs.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Once the gateway is enabled, separate Foundry experiences can also register custom agents and govern supported MCP tools. Those actions do not happen in the AI Gateway tab itself, but the gateway provides their governed traffic path.&lt;/P&gt;
&lt;P&gt;That is a meaningful shift. In the original sample, the infrastructure owner created APIM, imported the model API, deployed policy XML, and wired the agent to the gateway. Foundry does not remove APIM or replace those advanced capabilities. It gives Foundry administrators a first-party path into the same architecture.&lt;/P&gt;
&lt;H2&gt;What Foundry now owns&lt;/H2&gt;
&lt;H3&gt;Project onboarding&lt;/H3&gt;
&lt;P&gt;An AI Gateway is associated with a Foundry resource, then enabled for projects inside it. New projects can inherit the gateway automatically. Existing projects must be added explicitly.&lt;/P&gt;
&lt;P&gt;This creates a useful governance boundary: multiple teams can share the same APIM instance without sharing one undifferentiated token pool. Each project can receive its own token ceiling. If a project exceeds its configured limit, the gateway returns &lt;CODE&gt;429 Too Many Requests&lt;/CODE&gt; while other projects continue using their allocations.&lt;/P&gt;
&lt;H3&gt;Foundry-level limits&lt;/H3&gt;
&lt;P&gt;Our earlier sample used an APIM token-limit policy keyed by an APIM subscription. That is still a valid pattern, especially when the consumer boundary is an application, business unit, or external customer.&lt;/P&gt;
&lt;P&gt;The Foundry integration adds another natural counter key: the Foundry project. For organizations already organizing models and agents by project, limits can now follow that structure instead of being recreated independently in every client application.&lt;/P&gt;
&lt;H3&gt;Inventory for agents and tools&lt;/H3&gt;
&lt;P&gt;The AI Gateway experience is expanding beyond model endpoints in two distinct ways. Foundry can register custom agents running outside Foundry, including agents that use A2A as their communication protocol. Separately, AI Gateway can govern supported MCP tools. The current tools-governance preview is scoped to MCP; it does not cover every Foundry or OpenAPI tool type.&lt;/P&gt;
&lt;P&gt;That matters for App Service because the runtime does not need to move. An agent can continue running on App Service while participating in Foundry's inventory, and its supported MCP traffic can flow through the gateway.&lt;/P&gt;
&lt;P&gt;These experiences are still preview features, so I would treat them as a control-plane integration to evaluate, not a reason to redesign a working production agent.&lt;/P&gt;
&lt;H2&gt;What still belongs in APIM&lt;/H2&gt;
&lt;P&gt;Foundry makes the common path easier. It does not make the full APIM surface unnecessary.&lt;/P&gt;
&lt;P&gt;I would still use APIM directly for:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Custom policy expressions and organization-specific headers or claims.&lt;/LI&gt;
&lt;LI&gt;Semantic caching backed by Azure Managed Redis.&lt;/LI&gt;
&lt;LI&gt;Backend pools, priority routing, and circuit breakers.&lt;/LI&gt;
&lt;LI&gt;Private networking, multi-region gateways, and advanced topology decisions.&lt;/LI&gt;
&lt;LI&gt;Custom metrics and dimensions needed for internal chargeback.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The clean mental model is: &lt;STRONG&gt;Foundry manages the AI assets and common project guardrails; APIM manages the traffic contract and advanced gateway behavior.&lt;/STRONG&gt;&lt;/P&gt;
&lt;H2&gt;Applying this to the App Service sample&lt;/H2&gt;
&lt;P&gt;The existing &lt;A class="lia-external-url" href="https://github.com/seligj95/app-service-ai-gateway-mcp-apim-python" target="_blank" rel="noopener"&gt;framework-agnostic AI Gateway sample&lt;/A&gt; hosts a Microsoft Agent Framework agent and an MCP server on one Linux App Service. The agent reaches its model through APIM, which applies token limits, semantic caching, token metrics, and managed-identity authentication.&lt;/P&gt;
&lt;P&gt;Most of that architecture should stay exactly as it is.&lt;/P&gt;
&lt;H3&gt;1. Keep the App Service application unchanged&lt;/H3&gt;
&lt;P&gt;The agent should continue calling a stable gateway endpoint. It should not need to understand whether APIM was provisioned by Bicep, associated in Foundry, or managed by a central platform team. That separation was the reason to introduce the gateway in the first place.&lt;/P&gt;
&lt;H3&gt;2. Choose the APIM ownership model deliberately&lt;/H3&gt;
&lt;P&gt;For a new proof of concept, Foundry can create a Basic v2 APIM instance. For production, Microsoft recommends evaluating Standard v2 or Premium v2 based on throughput and networking requirements.&lt;/P&gt;
&lt;P&gt;For an existing enterprise gateway, use the Foundry option to associate an APIM instance only after confirming that it meets the eligibility requirements: it must be in the same Microsoft Entra tenant and Azure subscription as the Foundry resource, you need the &lt;STRONG&gt;API Management Service Contributor&lt;/STRONG&gt; role (or Owner), and the instance must use a supported v2 tier.&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Important migration detail:&lt;/STRONG&gt; the original sample defaults to the APIM Developer tier. That instance will not appear in Foundry's &lt;EM&gt;Use existing APIM&lt;/EM&gt; list because the direct integration currently requires a v2 tier. Do not assume an existing classic-tier gateway can simply be attached.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3&gt;3. Add existing projects explicitly&lt;/H3&gt;
&lt;P&gt;Creating or associating the gateway at the Foundry resource level does not automatically enable every existing project. Add each existing project to the gateway, then verify its status is &lt;STRONG&gt;Enabled&lt;/STRONG&gt;.&lt;/P&gt;
&lt;H3&gt;4. Decide where each limit belongs&lt;/H3&gt;
&lt;P&gt;Avoid layering limits without a clear ownership model. A Foundry project limit, an APIM subscription limit, and a model deployment TPM limit answer different questions:&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Control&lt;/th&gt;&lt;th&gt;Best boundary&lt;/th&gt;&lt;th&gt;Question it answers&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Foundry project limit&lt;/td&gt;&lt;td&gt;Team or workload project&lt;/td&gt;&lt;td&gt;How much shared capacity can this project consume?&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;APIM policy limit&lt;/td&gt;&lt;td&gt;Subscription, user, tenant, or application&lt;/td&gt;&lt;td&gt;How much can this specific consumer use?&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Model deployment quota&lt;/td&gt;&lt;td&gt;Backend deployment&lt;/td&gt;&lt;td&gt;What capacity exists at the model endpoint?&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;Use all three when the boundaries are intentional. Otherwise, operators end up debugging a &lt;CODE&gt;429&lt;/CODE&gt; without knowing which layer produced it.&lt;/P&gt;
&lt;H3&gt;5. Verify the data path&lt;/H3&gt;
&lt;P&gt;After enabling a project, make a model request and confirm that the APIM request metric increments. Microsoft also recommends checking the &lt;CODE&gt;GatewayLogs&lt;/CODE&gt; table for a successful response and an API name matching the AI Gateway.&lt;/P&gt;
&lt;P&gt;Then test the failure path. Set a deliberately small project limit, exceed it, and confirm that the gateway returns &lt;CODE&gt;429&lt;/CODE&gt;. A governance feature is not finished until the team has observed both the allowed and denied behavior.&lt;/P&gt;
&lt;H2&gt;Which path should you choose?&lt;/H2&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Scenario&lt;/th&gt;&lt;th&gt;Recommended approach&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;New Foundry proof of concept&lt;/td&gt;&lt;td&gt;Create the AI Gateway from Foundry and start with project-level limits.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Existing v2 enterprise APIM&lt;/td&gt;&lt;td&gt;Associate the existing instance and preserve central networking and policy ownership.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Existing classic-tier APIM&lt;/td&gt;&lt;td&gt;Keep the working gateway or plan a deliberate v2 migration; it cannot be selected directly today.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Advanced routing or custom policy needs&lt;/td&gt;&lt;td&gt;Use Foundry for inventory and common guardrails, then manage the advanced behavior in APIM.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Strict isolation between project groups&lt;/td&gt;&lt;td&gt;Use separate Foundry resources and separate gateways rather than assuming one gateway per project.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H2&gt;The architecture was right; the experience caught up&lt;/H2&gt;
&lt;P&gt;The most important takeaway is not that everyone should rebuild an AI Gateway in a new portal. It is that the composable architecture now has a more accessible control plane.&lt;/P&gt;
&lt;P&gt;App Service still runs the agent. APIM still governs the traffic. Foundry now gives model and platform owners a direct way to connect projects, allocate capacity, and bring agents and tools into the same governance view.&lt;/P&gt;
&lt;P&gt;If you already built the earlier sample, the application boundary does not need to change. Evaluate the Foundry integration, decide whether its project model matches your organization, and check the APIM tier before planning any migration. The gateway remains the contribution; Foundry now makes it easier to operate.&lt;/P&gt;
&lt;H2&gt;Resources&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A class="lia-internal-link lia-internal-url lia-internal-url-content-type-blog" href="https://techcommunity.microsoft.com/blog/appsonazureblog/you-can-build-a-framework-agnostic-ai-gateway-on-azure-app-service-%E2%80%94-heres-how/4522004" target="_blank" rel="noopener" data-lia-auto-title="You can build a framework-agnostic AI Gateway on Azure App Service" data-lia-auto-title-active="0"&gt;You can build a framework-agnostic AI Gateway on Azure App Service&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-internal-link lia-internal-url lia-internal-url-content-type-blog" href="https://techcommunity.microsoft.com/blog/appsonazureblog/from-ai-adoption-to-ai-governance---using-apim-as-the-gateway-for-azure-ai-found/4536247" target="_blank" rel="noopener" data-lia-auto-title="From AI Adoption to AI Governance" data-lia-auto-title-active="0"&gt;From AI Adoption to AI Governance&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/foundry/configuration/enable-ai-api-management-gateway-portal" target="_blank" rel="noopener"&gt;Configure AI Gateway in your Foundry resources&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/api-management/genai-gateway-capabilities" target="_blank" rel="noopener"&gt;AI Gateway capabilities in Azure API Management&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://github.com/seligj95/app-service-ai-gateway-mcp-apim-python" target="_blank" rel="noopener"&gt;App Service AI Gateway sample&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;</description>
      <pubDate>Fri, 17 Jul 2026 17:04:16 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/apps-on-azure-blog/microsoft-foundry-now-has-an-ai-gateway-control-plane-what/ba-p/4538320</guid>
      <dc:creator>jordanselig</dc:creator>
      <dc:date>2026-07-17T17:04:16Z</dc:date>
    </item>
  </channel>
</rss>

