<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>New blog articles in Microsoft Community Hub</title>
    <link>https://techcommunity.microsoft.com/t5/</link>
    <description>Microsoft Community Hub</description>
    <pubDate>Wed, 26 Aug 2026 23:32:00 GMT</pubDate>
    <dc:creator>Community</dc:creator>
    <dc:date>2026-08-26T23:32:00Z</dc:date>
    <item>
      <title>Partner Blog | Building the foundation for AI: Cloud, data, security, and AI skills for partners</title>
      <link>https://techcommunity.microsoft.com/t5/partner-news/partner-blog-building-the-foundation-for-ai-cloud-data-security/ba-p/4550578</link>
      <description>&lt;P&gt;Customers are moving beyond AI experimentation. They are looking for partners who can connect AI ambition to the cloud, data, security, governance, and business application capabilities required to put AI to work. That makes skilling across the Microsoft stack increasingly important.&lt;/P&gt;
&lt;P&gt;It can also make the question of where to start more difficult. This month, there is a simpler starting point. The&amp;nbsp;&lt;A href="https://www.youtube.com/watch?v=xZ0guJQ20wM" target="_blank" rel="noopener"&gt;Microsoft Partner Skilling Hub Agent&amp;nbsp;&lt;/A&gt;can recommend training, answer skilling-related questions, and generate personalized technical skilling plans based on your role and goals. From there, you can build a focused path across Frontier Transformation, agents, Microsoft 365 Copilot, certifications, hands-on learning, and co-sell execution.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&amp;nbsp;&lt;/STRONG&gt;&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;The foundation for AI is broader than AI skills&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;Frontier Transformation is the shift from targeted AI pilots to repeatable, governed AI capabilities embedded into the flow of work, business processes, and customer engagement. For partners, delivering that transformation requires more than expertise in a single AI product. It requires teams that understand how cloud infrastructure, data, security, agents, and business applications work together.&lt;/P&gt;
&lt;P&gt;That foundation matters across customer segments. For partners serving small and medium-sized businesses (SMBs), it can support you in guiding customers toward practical AI adoption while addressing security, governance, productivity, and business process needs together.&lt;/P&gt;
&lt;P&gt;This month, focus your skilling plan on five areas: validating your technical expertise, building agent platform capabilities, developing AI business application skills, earning industry-recognized certifications, and applying those skills through hands-on learning.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&amp;nbsp;&lt;/STRONG&gt;&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;Validate your Frontier Transformation expertise&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;The&amp;nbsp;&lt;A href="https://skillupwithlevelup.com/frontier" target="_blank" rel="noopener"&gt;Frontier Transformation Engineer Badge&lt;/A&gt;&amp;nbsp;gives technical professionals a path to validate their ability to design, build, and deliver AI solutions.&lt;/P&gt;
&lt;P&gt;The journey progresses from Microsoft certifications through project-ready skilling and advanced Frontier engineering expertise. It is designed to build capability across agents, Microsoft 365 Copilot, Microsoft Foundry, Security, and the broader Frontier stack, giving your technical teams a structured way to move from foundational knowledge toward advanced implementation skills.&lt;/P&gt;
&lt;P&gt;If you are developing a deeper AI engineering practice, consider identifying technical professionals who are ready to progress through the Frontier Transformation Engineer journey.&lt;/P&gt;
&lt;P&gt;&lt;A class="lia-external-url" href="https://aka.ms/partnerblog-skillingAug26" target="_blank"&gt;Continue reading here&lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Wed, 26 Aug 2026 21:08:26 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/partner-news/partner-blog-building-the-foundation-for-ai-cloud-data-security/ba-p/4550578</guid>
      <dc:creator>JillArmourMicrosoft</dc:creator>
      <dc:date>2026-08-26T21:08:26Z</dc:date>
    </item>
    <item>
      <title>Prepare for the launch of growth margins with API readiness resources</title>
      <link>https://techcommunity.microsoft.com/t5/partner-news/prepare-for-the-launch-of-growth-margins-with-api-readiness/ba-p/4550566</link>
      <description>&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Starting October 1, 2026, growth margins will provide eligible partners with&amp;nbsp;incremental&amp;nbsp;margin on qualifying Microsoft 365 growth opportunities. Built on partner feedback, growth margins give you more flexibility to structure deals, compete for new business, reward reseller performance, and reinvest in capabilities that support long-term growth. Review the available readiness resources now so you can maximize growth opportunities,&amp;nbsp;maintain&amp;nbsp;business continuity, and prepare for upcoming technical changes.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:0,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:0,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Technical guidance is available to help you navigate API changes and implement best practices for data export and billing automation.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:0,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="3" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;A href="https://partner.microsoft.com/asset/collection/growth-margins-partner-resources?wt.mc_id=9aj8mwtit3" target="_blank"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;Review the partner resource collection&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;469777462&amp;quot;:[142],&amp;quot;469777927&amp;quot;:[0],&amp;quot;469777928&amp;quot;:[1]}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="3" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="2" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;Evaluate opportunities to&amp;nbsp;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;A href="https://learn.microsoft.com/partner-center/pricing/growth-margins?wt.mc_id=u3norzm814" target="_blank"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;incorporate growth margins into your Microsoft 365 sales and growth strategies&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;469777462&amp;quot;:[142],&amp;quot;469777927&amp;quot;:[0],&amp;quot;469777928&amp;quot;:[1]}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="3" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="3" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;Eligible partners may&amp;nbsp;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;A href="https://learn.microsoft.com/partner-center/benefits/technical-benefits?wt.mc_id=wgkswgas5s" target="_blank"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;open an&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;a&lt;/SPAN&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;dvisory&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;c&lt;/SPAN&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;ase with Microsoft Partner Technical Consulting&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="List Bullet"&gt;&amp;nbsp;for API transition guidance.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;469777462&amp;quot;:[142],&amp;quot;469777927&amp;quot;:[0],&amp;quot;469777928&amp;quot;:[1]}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 26 Aug 2026 19:36:50 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/partner-news/prepare-for-the-launch-of-growth-margins-with-api-readiness/ba-p/4550566</guid>
      <dc:creator>JillArmourMicrosoft</dc:creator>
      <dc:date>2026-08-26T19:36:50Z</dc:date>
    </item>
    <item>
      <title>Understanding and Reclaiming Storage After DELETE in Azure Database for PostgreSQL</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-blog-for-postgresql/understanding-and-reclaiming-storage-after-delete-in-azure/ba-p/4548797</link>
      <description>&lt;H2&gt;&lt;STRONG&gt;Why DELETE Does Not Immediately Shrink Storage and Azure Monitor Storage Graphs Continue to Show High Storage Consumption&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;PostgreSQL uses Multi-Version Concurrency Control, commonly called MVCC. MVCC allows readers and writers to work at the same time without blocking each other unnecessarily. This is great for concurrency, but it changes how storage behaves after DELETE and UPDATE operations.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;When you execute DELETE:&lt;/STRONG&gt;&lt;/P&gt;
&lt;PRE&gt;&amp;nbsp;&lt;/PRE&gt;
&lt;LI-CODE lang="sql"&gt;DELETE FROM orders
WHERE order_date &amp;lt; '2026-08-20';&lt;/LI-CODE&gt;
&lt;P&gt;PostgreSQL does &lt;STRONG&gt;not&lt;/STRONG&gt; physically remove the rows from disk immediately.&lt;/P&gt;
&lt;P&gt;Instead, it:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Marks the rows as &lt;STRONG&gt;dead tuples&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI&gt;Keeps the row versions on disk so that transactions that started earlier can still see a consistent view of the data.&lt;/LI&gt;
&lt;LI&gt;Leaves the space occupied by those rows in the table file.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;As a result:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Logical row count decreases.&lt;/LI&gt;
&lt;LI&gt;Physical table size typically remains unchanged.&lt;/LI&gt;
&lt;LI&gt;Azure Monitor storage graphs continue to show similar storage consumption.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The deleted pages remain allocated until PostgreSQL rewrites or reorganizes them.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;When you execute UPDATE:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="sql"&gt;UPDATE customers 
SET city = 'London' 
WHERE customer_id = 100; &lt;/LI-CODE&gt;
&lt;P&gt;PostgreSQL:&amp;nbsp;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Marks the existing row version as obsolete.&amp;nbsp;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI&gt;Creates a new row version with the updated value.&amp;nbsp;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI&gt;Retains the old version until no active transaction needs it.&amp;nbsp;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;This behavior is a direct result of PostgreSQL's MVCC (Multi-Version Concurrency Control) architecture, which allows concurrent transactions without blocking readers.&lt;/P&gt;
&lt;P&gt;This allows readers and writers to work concurrently, reduces locking, and gives long-running queries a consistent view of the data. The trade-off is that old row versions continue to occupy space until they can be cleaned up.&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Learn more about MVCC&amp;nbsp; &lt;A href="https://www.postgresql.org/docs/7.1/mvcc.html" target="_blank" rel="noopener"&gt;PostgreSQL: Documentation: 7.1: Multi-Version Concurrency Control&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&amp;nbsp;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Key Azure Metrics used for Storage analysis for Azure Database for PostgreSQL Flexible server:&lt;BR /&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;The following metrics are usually the most useful during storage analysis:&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;Please refer&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/azure/postgresql/monitor/concepts-monitoring" target="_blank" rel="noopener"&gt;Monitor Using Metrics and Logs in Azure Database for PostgreSQL Flexible Server - Azure Database for PostgreSQL | Microsoft Learn&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;After large DELETE operations, it's important to understand the difference between reclaiming database space for reuse and reducing Azure storage allocation and billing.&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;&amp;nbsp;&lt;/STRONG&gt;&lt;STRONG&gt;Storage That Can Be Reclaimed&lt;/STRONG&gt;&lt;/H5&gt;
&lt;P&gt;The following types of space can often be reclaimed or reused:&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Dead Tuple Space&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Obsolete row versions created by UPDATE and DELETE are called dead tuples.&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Dead tuples are no longer visible to normal queries, but they remain in the table files until vacuuming makes their space reusable. Workloads with frequent updates, deletes, batch processing, or ETL operations can accumulate large numbers of dead tuples. Autovacuum or manual VACUUM operations can clean up these tuples and make the space available for future database activity&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;How to Diagnose Table Bloat&amp;nbsp;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Start by checking tables with high dead tuple counts:&amp;nbsp;&lt;BR /&gt;&lt;BR /&gt;&lt;/P&gt;
&lt;LI-CODE lang="sql"&gt;SELECT 
    relname, 
    n_live_tup, 
    n_dead_tup 
FROM pg_stat_user_tables 
ORDER BY n_dead_tup DESC;&lt;/LI-CODE&gt;
&lt;P&gt;Below query identifies the largest tables:&amp;nbsp;&lt;/P&gt;
&lt;LI-CODE lang="sql"&gt;SELECT 
    relname, 
    pg_size_pretty(pg_total_relation_size(relid)) AS total_size 
FROM pg_catalog.pg_statio_user_tables 
ORDER BY pg_total_relation_size(relid) DESC; &lt;/LI-CODE&gt;
&lt;P&gt;Tables that are both large and have high dead tuple counts are strong candidates for maintenance review.&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Table and Index Bloat&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Over time, frequent INSERT, UPDATE, and DELETE operations can create bloat. This can be addressed using:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;VACUUM FULL&lt;/LI&gt;
&lt;LI&gt;pg_repack (preferred for production environments due to minimal locking)&lt;/LI&gt;
&lt;LI&gt;Partition maintenance strategies&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;These operations can physically shrink tables and indexes, returning unused space to the operating system.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Temporary Storage Consumption&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Storage consumed by temporary files generated from sorts, hash joins, and large query operations can be reclaimed automatically after the query completes. Identifying and tuning such workloads can prevent recurring storage spikes.&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;What Autovacuum Actually Does&lt;/STRONG&gt;&lt;/H5&gt;
&lt;H6&gt;PostgreSQL uses autovacuum to clean up obsolete row versions in the background. It also updates planner statistics and helps prevent transaction ID wraparound.&amp;nbsp;&lt;/H6&gt;
&lt;H6&gt;A regular vacuum operation:&amp;nbsp;&lt;/H6&gt;
&lt;UL&gt;
&lt;LI&gt;&amp;nbsp; &amp;nbsp; Removes dead tuples when they are no longer visible to any transaction.&amp;nbsp;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;&amp;nbsp; &amp;nbsp; Marks their space as reusable.&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;&amp;nbsp; &amp;nbsp; Updates statistics when configured and also the maintenance information.&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;&amp;nbsp; &amp;nbsp; Does not normally shrink the table file.&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H6&gt;Check whether Autovacuum is enabled:&amp;nbsp;&lt;/H6&gt;
&lt;LI-CODE lang="sql"&gt;SHOW Autovacuum;&lt;/LI-CODE&gt;
&lt;P&gt;Below SQL query reviews tables with high dead-tuple counts:&lt;/P&gt;
&lt;LI-CODE lang="sql"&gt;SELECT 
    schemaname, 
    relname, 
    n_live_tup, 
    n_dead_tup, 
    last_autovacuum, 
    last_vacuum 
FROM pg_stat_user_tables 
ORDER BY n_dead_tup DESC; &lt;/LI-CODE&gt;
&lt;H6&gt;A high count may mean that autovacuum is not keeping up, large batch operations are producing dead tuples quickly, or old transactions are preventing cleanup.&lt;/H6&gt;
&lt;H6&gt;Learn more about Autovacuum &lt;A href="https://www.postgresql.org/docs/current/runtime-config-vacuum.html" target="_blank" rel="noopener"&gt;PostgreSQL: Documentation: 18: 19.10.&amp;nbsp;Vacuuming&lt;/A&gt;&amp;nbsp;&lt;/H6&gt;
&lt;H5&gt;&lt;STRONG&gt;&amp;nbsp;F&lt;/STRONG&gt;&lt;STRONG&gt;ollowing options helps us to reclaim/reuse the storage &lt;/STRONG&gt;&lt;/H5&gt;
&lt;H6&gt;&lt;STRONG&gt;Option 1: Standard VACUUM&lt;/STRONG&gt;&lt;/H6&gt;
&lt;LI-CODE lang="sql"&gt;VACUUM audit_logs;&lt;/LI-CODE&gt;
&lt;P&gt;Use standard VACUUM as routine maintenance. It makes deleted space reusable and helps improve database health, but it does not physically reduce table files.&lt;/P&gt;
&lt;H6&gt;&lt;STRONG&gt;Option 2: VACUUM FULL&lt;/STRONG&gt;&lt;/H6&gt;
&lt;H6&gt;To return unused table space to the operating system, PostgreSQL provides:&amp;nbsp;&lt;/H6&gt;
&lt;LI-CODE lang="sql"&gt;VACUUM FULL schema_name.table_name; &lt;/LI-CODE&gt;
&lt;P&gt;&amp;nbsp;VACUUM FULL rewrites the table into a compact physical file containing only the active rows. It can substantially reduce the table's file size.&amp;nbsp;However, it has several operational costs:&amp;nbsp;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;&amp;nbsp; It takes an ACCESS EXCLUSIVE lock.&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;&amp;nbsp; Reads and writes to the table are blocked.&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;&amp;nbsp; It may run for a long time on large tables.&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;&amp;nbsp; &amp;nbsp;It requires additional temporary storage during the rewrite.&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;&amp;nbsp; &amp;nbsp;Associated indexes may also need time and space to be rebuilt.&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H6&gt;&amp;nbsp;Please refer &lt;A href="https://wiki.postgresql.org/wiki/VACUUM_FULL" target="_blank" rel="noopener"&gt;VACUUM FULL - PostgreSQL wiki&lt;/A&gt;&amp;nbsp;&lt;/H6&gt;
&lt;H6&gt;&lt;STRONG&gt;Option 3: pg_repack&lt;/STRONG&gt;&lt;/H6&gt;
&lt;P&gt;For production tables that cannot be unavailable for an extended period, pg_repack is often a better option.&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Example:&amp;nbsp;&lt;/P&gt;
&lt;LI-CODE lang="shell"&gt;pg_repack -d mydb -t schema_name.large_table &lt;/LI-CODE&gt;
&lt;P&gt;It rebuilds tables and indexes with much less blocking than VACUUM FULL. A short lock may still be required during setup or the final object swap, so it is not entirely lock-free.&amp;nbsp;&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;img /&gt;&lt;/DIV&gt;
&lt;P&gt;Before using pg_repack on Azure Database for PostgreSQL, confirm that the required extension is supported and enabled for the server configuration.&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Please refer &lt;A href="https://learn.microsoft.com/azure/postgresql/troubleshoot/how-to-perform-fullvacuum-pg-repack" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/azure/postgresql/troubleshoot/how-to-perform-fullvacuum-pg-repack&lt;/A&gt;&lt;/P&gt;
&lt;H6&gt;&lt;STRONG&gt;Option 4: Partitioning for Future Prevention&lt;/STRONG&gt;&lt;/H6&gt;
&lt;H6&gt;For historical data such as audit logs, telemetry, or event tables, partitioning is often the best long-term solution. Instead of deleting millions of old rows, you can remove an old partition.&amp;nbsp;&lt;/H6&gt;
&lt;LI-CODE lang="sql"&gt;DROP TABLE audit_logs_2026; &lt;/LI-CODE&gt;
&lt;H6&gt;This approach is faster, produces less bloat, and reduces the pressure on Autovacuum.&lt;/H6&gt;
&lt;H6&gt;Azure storage metrics may remain high after rows are deleted and vacuumed.&amp;nbsp;&lt;/H6&gt;
&lt;H6&gt;For example:&amp;nbsp;&lt;/H6&gt;
&lt;img /&gt;
&lt;H6&gt;This can happen for two separate reasons:&amp;nbsp;&lt;/H6&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;&amp;nbsp;PostgreSQL may have converted the deleted space into reusable space without shrinking its files.&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;&amp;nbsp;Azure may retain the server's provisioned storage allocation even after PostgreSQL files are compacted.&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H6&gt;As a result, a reduction in table size does not necessarily reduce the storage provisioned for the Flexible Server. Once Azure storage has grown, it generally should not be assumed that it will automatically shrink.&lt;/H6&gt;
&lt;img /&gt;
&lt;P&gt;&lt;STRONG&gt;Provisioned Storage Cannot Be Reduced&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Azure Database for PostgreSQL Flexible Server allows storage to be increased, but currently does not support reducing the configured storage size afterward.&lt;/P&gt;
&lt;P&gt;For example:&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;img /&gt;&lt;/DIV&gt;
&lt;P&gt;Even if PostgreSQL reclaims internal space through VACUUM FULL or pg_repack, the server remains provisioned at 512 GB.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Storage Costs Continue Based on Provisioned Capacity&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Because billing is tied to the provisioned storage allocation, deleting data does not automatically lower storage costs.&lt;/P&gt;
&lt;P&gt;The following actions do &lt;STRONG&gt;not&lt;/STRONG&gt; reduce provisioned storage billing:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;DELETE operations&lt;/LI&gt;
&lt;LI&gt;VACUUM&lt;/LI&gt;
&lt;LI&gt;VACUUM FULL&lt;/LI&gt;
&lt;LI&gt;Autovacuum cleanup&lt;/LI&gt;
&lt;LI&gt;pg_repack&lt;/LI&gt;
&lt;/UL&gt;
&lt;H4&gt;&lt;STRONG&gt;Best Practices to Avoid Storage Surprises&lt;/STRONG&gt;&lt;/H4&gt;
&lt;UL&gt;
&lt;LI&gt;Do not run massive DELETE operations during peak workload hours.&lt;/LI&gt;
&lt;LI&gt;Delete in controlled batches when a large purge is unavoidable.&lt;/LI&gt;
&lt;LI&gt;Monitor dead tuples, table size, index size, WAL generation, and storage percent.&lt;/LI&gt;
&lt;LI&gt;Investigate long-running transactions before assuming Autovacuum is broken.&lt;/LI&gt;
&lt;LI&gt;Use partitioning for retention-based cleanup.&lt;/LI&gt;
&lt;LI&gt;Use pg_repack or planned VACUUM FULL when physical compaction is required.&lt;/LI&gt;
&lt;LI&gt;Set Azure Monitor alerts for storage utilization and plan capacity before full-disk scenarios occur.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H4&gt;&lt;STRONG&gt;Conclusion&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P&gt;When Azure PostgreSQL storage stays high after deleting, the database is usually behaving as designed. DELETE creates dead tuples and reusable internal space, but it does not automatically shrink physical files or reduce Azure allocated storage.&lt;/P&gt;
&lt;P&gt;The right strategy depends on your goal. Use VACUUM for routine health, pg_repack for minimal-downtime compaction, VACUUM FULL when a maintenance window is acceptable, and partitioning to avoid large deletes in the first place.&lt;BR /&gt;Storage charges are tied to the amount of storage provisioned for the server. As a result, deleting data alone does not lower storage-related costs.&lt;/P&gt;</description>
      <pubDate>Wed, 26 Aug 2026 19:36:07 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-blog-for-postgresql/understanding-and-reclaiming-storage-after-delete-in-azure/ba-p/4548797</guid>
      <dc:creator>Shweta_Anjankar</dc:creator>
      <dc:date>2026-08-26T19:36:07Z</dc:date>
    </item>
    <item>
      <title>Your on-call rotation has a new member: 10 production incidents, end to end, with Azure SRE Agent</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-mission-critical-blog/your-on-call-rotation-has-a-new-member-10-production-incidents/ba-p/4545187</link>
      <description>&lt;DIV class="mce-toc"&gt;
&lt;H2&gt;Table of Contents&lt;/H2&gt;
&lt;UL&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_1" target="_self"&gt;TL;DR&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_2" target="_self"&gt;1. Why the "investigation" half is the real prize&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_3" target="_self"&gt;2. What Azure SRE Agent actually does&lt;/A&gt;&lt;UL&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_4" target="_self"&gt;The three primary use cases&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_5" target="_self"&gt;The five extension primitives&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_6" target="_self"&gt;Integrations you can assume&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_7" target="_self"&gt;Skills vs. custom agents vs. knowledge files&lt;/A&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_8" target="_self"&gt;3. Anatomy of an agent-run incident&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_9" target="_self"&gt;4. Build the demo lab&lt;/A&gt;&lt;UL&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_10" target="_self"&gt;4.1 Resource group and agent&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_11" target="_self"&gt;4.2 Choose a permission level deliberately&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_12" target="_self"&gt;4.3 Connect ServiceNow&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_13" target="_self"&gt;4.4 Connect deployment correlation&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_14" target="_self"&gt;4.5 Create the demo resources&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_15" target="_self"&gt;4.6 One custom agent per domain&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_16" target="_self"&gt;4.7 Response plans&lt;/A&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_17" target="_self"&gt;5. Set your guardrails before your first incident&lt;/A&gt;&lt;UL&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_18" target="_self"&gt;5.1 Run modes&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_19" target="_self"&gt;5.2 What the product blocks for you&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_20" target="_self"&gt;5.3 Hooks: the guardrail you write yourself&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_21" target="_self"&gt;5.4 The recommended production policy&lt;/A&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_22" target="_self"&gt;6. The ten use cases&lt;/A&gt;&lt;UL&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_23" target="_self"&gt;Use case #1 — App Service: HTTP 500 after deployment&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_24" target="_self"&gt;Use case #2 — AKS: pods in CrashLoopBackOff&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_25" target="_self"&gt;Use case #3 — Azure SQL Database: CPU saturation and timeouts&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_26" target="_self"&gt;Use case #4 — Azure Cosmos DB: HTTP 429 throttling&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_27" target="_self"&gt;Use case #5 — Azure VM: OS/root disk full&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_28" target="_self"&gt;Use case #6 — Linux VM: anomalous CPU saturation&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_29" target="_self"&gt;Use case #7 — Windows VM with IIS: memory leak&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_30" target="_self"&gt;Use case #8 — Virtual Machine Scale Set: unhealthy instance&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_31" target="_self"&gt;Use case #9 — Application Gateway: HTTP 502 from unhealthy backends&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_32" target="_self"&gt;Use case #10 — Azure Service Bus: queue and dead-letter backlog&lt;/A&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_33" target="_self"&gt;7. The ITSM integration model&lt;/A&gt;&lt;UL&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_34" target="_self"&gt;Recommended incident fields&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_35" target="_self"&gt;Recommended agent-generated timeline&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_36" target="_self"&gt;Deduplication strategy&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_37" target="_self"&gt;Record responsibilities&lt;/A&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_38" target="_self"&gt;8. Reality check: where you have to build&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_39" target="_self"&gt;9. Approval and autonomy policy&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_40" target="_self"&gt;10. Cross-cutting security controls&lt;/A&gt;&lt;UL&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_41" target="_self"&gt;What the platform gives you for free&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_42" target="_self"&gt;The audit trail&lt;/A&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_43" target="_self"&gt;11. A 30/60/90 pilot that survives contact with your CAB&lt;/A&gt;&lt;UL&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_44" target="_self"&gt;The best first five candidates&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_45" target="_self"&gt;What to measure in week one&lt;/A&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_46" target="_self"&gt;12. Measuring whether it's actually working&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_47" target="_self"&gt;13. Resources&lt;/A&gt;&lt;/LI&gt;&lt;LI&gt;&lt;A href="#community--1-mcetoc_blog_48" target="_self"&gt;Closing&lt;/A&gt;&lt;/LI&gt;&lt;/UL&gt;
&lt;/DIV&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;What this post is.&lt;/STRONG&gt; A hands-on, reproducible walkthrough of ten real production failure modes — App Service, AKS, Azure SQL, Cosmos DB, VMs, VM Scale Sets, Application Gateway, and Service Bus — each one driven end to end by Azure SRE Agent: detection, hypothesis-driven investigation, ITSM ticketing, bounded remediation, recovery validation, and follow-up records.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;What this post is not.&lt;/STRONG&gt; A claim that you install SRE Agent and all of this happens on day one. Every workflow below needs telemetry, scoped RBAC, a response plan, an approved action surface, and an ITSM integration. I'll be explicit about which parts are documented product behavior and which parts you have to wire yourself — because that distinction is the difference between a demo that works on stage and one that works at 3 AM.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;HR /&gt;
&lt;H2 id="mcetoc_blog_1"&gt;TL;DR&lt;/H2&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;The pattern&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Alert → agent investigates read-only → agent opens/updates the ticket with evidence → agent proposes a &lt;STRONG&gt;bounded&lt;/STRONG&gt; action → human approves → agent executes → agent validates recovery → agent files the follow-up record&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;The unlock&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Not "AI fixes prod." The unlock is that &lt;EM&gt;investigation&lt;/EM&gt; — the 20 minutes of tab-switching between Azure Monitor, App Insights, deployment history, and Activity Log — is fully automated and consistent, and the fix arrives pre-justified with an evidence chain&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;The guardrail&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read-only automatic. Ticketing automatic. Production writes in &lt;STRONG&gt;Review&lt;/STRONG&gt; mode. Guest-OS work through fixed-purpose runbooks, never a shell prompt&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;The reality check&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;SRE Agent hard-blocks &lt;CODE&gt;delete&lt;/CODE&gt;/&lt;CODE&gt;remove&lt;/CODE&gt; and all &lt;CODE&gt;az keyvault&lt;/CODE&gt; commands, respects Azure management locks, and only one incident platform can be active at a time. Several patterns in this post need a custom tool or MCP server to complete the loop&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Time to first value&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;One resource group, one alert rule, one response plan. You can reproduce use case #1 in an afternoon&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;HR /&gt;
&lt;H2 id="mcetoc_blog_2"&gt;1. Why the "investigation" half is the real prize&lt;/H2&gt;
&lt;P&gt;Every conversation about AI in operations goes straight to remediation. "Will it restart my app?" That's the least interesting question, and it's the one with the most downside risk.&lt;/P&gt;
&lt;P&gt;Think about what actually consumes the minutes during a Sev1. The alert fires. Someone acknowledges. Then:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Open Azure Monitor, confirm the metric is real and not a probe artifact&lt;/LI&gt;
&lt;LI&gt;Open Application Insights, find the dominant exception&lt;/LI&gt;
&lt;LI&gt;Open the deployment pipeline, find what shipped and when&lt;/LI&gt;
&lt;LI&gt;Open Activity Log, check whether someone changed configuration&lt;/LI&gt;
&lt;LI&gt;Open Resource Health, rule out a platform incident&lt;/LI&gt;
&lt;LI&gt;Open the &lt;EM&gt;other&lt;/EM&gt; environment, confirm the previous version is healthy&lt;/LI&gt;
&lt;LI&gt;Assemble all of that into a sentence a human can act on&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;That's twenty minutes of context assembly performed by a tired human, differently every time, with quality that depends entirely on who happens to be on call. It is the single most automatable part of incident response and the part nobody automates, because scripts can't reason about which of six hypotheses fits the evidence.&lt;/P&gt;
&lt;P&gt;This is exactly what &lt;A href="https://learn.microsoft.com/azure/sre-agent/root-cause-analysis" target="_blank"&gt;root cause analysis in SRE Agent&lt;/A&gt; is designed for. The agent doesn't grep logs — it forms hypotheses and invalidates them:&lt;/P&gt;
&lt;LI-CODE lang="text"&gt;HYPOTHESIS 1: Recent deployment broke something
├─ Checked: Last deployment was 3 days ago
├─ Evidence: Error rate stable until 30 minutes ago
└─ Result: INVALIDATED

HYPOTHESIS 2: Database overloaded
├─ Checked: Azure SQL metrics (CPU, DTU, connections)
├─ Evidence: DTU at 98%, query duration 4x normal
├─ Traced: SELECT * FROM orders WHERE... taking 8.2s
└─ Result: VALIDATED

ROOT CAUSE: Orders table missing index on customer_id column.
Query plan shows full table scan on 2.1M rows.

RECOMMENDED ACTION: Add index on orders.customer_id
Similar fix applied in INC-2341 (3 weeks ago)
&lt;/LI-CODE&gt;
&lt;P&gt;That last line — recalling a similar incident from three weeks ago — is the compounding part. Every thread produces a &lt;A href="https://learn.microsoft.com/azure/sre-agent/memory" target="_blank"&gt;session insight&lt;/A&gt;: symptoms observed, steps that worked, root cause, and pitfalls to avoid. Thirty minutes after a thread goes quiet, the agent indexes those learnings. Next time the same resource misbehaves, that history surfaces first.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;So the framing for the rest of this post:&lt;/STRONG&gt; remediation is the punchline, but investigation is the product. Every one of the ten use cases below has a large read-only phase you can turn on tomorrow with zero write permissions, and a small write phase you should gate behind approval for a long time.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="mcetoc_blog_3"&gt;2. What Azure SRE Agent actually does&lt;/H2&gt;
&lt;P&gt;Before the use cases, here is the honest capability map, drawn from the product documentation rather than from a keynote.&lt;/P&gt;
&lt;H3 id="mcetoc_blog_4"&gt;The three primary use cases&lt;/H3&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Use case&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;What it means&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Automate incidents&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Alert fires → agent queries monitoring tools, correlates signals across systems, identifies probable root cause, proposes mitigations&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Automate scheduled workflows&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Proactive health checks, compliance sweeps, and routine tasks on a schedule, with results routed to your incident platform or notification channel&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Investigate and advise&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Natural-language questions — "what changed in the last hour?" — answered with grounded citations&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;H3 id="mcetoc_blog_5"&gt;The five extension primitives&lt;/H3&gt;
&lt;P&gt;Everything you customize sits in one of five buckets:&lt;/P&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Primitive&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;What it is&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;When you reach for it&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Skills&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Procedural guidance (&lt;CODE&gt;SKILL.md&lt;/CODE&gt;) plus optional attached tools; auto-loaded when relevant&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Team troubleshooting runbooks that should also &lt;EM&gt;execute&lt;/EM&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Subagents / custom agents&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Purpose-built specialists invoked via &lt;CODE&gt;/agent&lt;/CODE&gt; or routed to by a response plan&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;A &lt;CODE&gt;DatabaseExpert&lt;/CODE&gt; that owns every SQL incident&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Python tools&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Custom logic, transformations, API calls&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Anything that needs code, e.g. writing to the ServiceNow Table API&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;MCP servers&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;40+ managed connectors (Datadog, New Relic, Splunk, Elastic, Dynatrace…) plus any custom MCP tool&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Bringing your non-Azure telemetry and your ITSM write surface into the loop&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Agent hooks&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Event-triggered automations at &lt;CODE&gt;Stop&lt;/CODE&gt; and &lt;CODE&gt;PostToolUse&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Policy enforcement, audit emission, blocking risky commands&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;Six generic subagents ship built in — &lt;STRONG&gt;Explore, Plan, CodeReview, Bash, Verification, GeneralPurpose&lt;/STRONG&gt; — and the agent can parallelize investigation, planning, review, shell, and verification work across them.&lt;/P&gt;
&lt;P&gt;A &lt;STRONG&gt;permission gate&lt;/STRONG&gt; sits in front of all five primitives and evaluates every proposed tool call &lt;EM&gt;before&lt;/EM&gt; it runs.&lt;/P&gt;
&lt;H3 id="mcetoc_blog_6"&gt;Integrations you can assume&lt;/H3&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Category&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Supported&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Monitoring&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Azure Monitor (metrics, logs, alerts, workbooks), Application Insights, Log Analytics, Grafana&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Incident platforms&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Azure Monitor Alerts, PagerDuty, ServiceNow — &lt;STRONG&gt;only one active at a time&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Source control / CI&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;GitHub (repos, issues), Azure DevOps (repos, work items)&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Data&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Azure Data Explorer (Kusto), MCP servers&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Comms&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Slack, Microsoft Teams&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;⚠️ &lt;STRONG&gt;Design constraint worth internalizing early.&lt;/STRONG&gt; Only one incident platform can be active at a time, and switching disconnects the current one. If your org runs PagerDuty for paging &lt;EM&gt;and&lt;/EM&gt; ServiceNow for records of truth, you must pick which one the agent is bound to and reach the other through a connector or custom tool. Every use case below assumes &lt;STRONG&gt;ServiceNow is the bound platform&lt;/STRONG&gt;.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3 id="mcetoc_blog_7"&gt;Skills vs. custom agents vs. knowledge files&lt;/H3&gt;
&lt;P&gt;The three are constantly confused. This table settles it:&lt;/P&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Skills&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Custom agents&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Knowledge files&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Access&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Automatic when relevant&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Explicit (&lt;CODE&gt;/agent&lt;/CODE&gt;) or routed by response plan&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Automatic search&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Tools&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Can attach&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Has its own&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;None&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Context&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Uses thread context&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Shares&lt;/STRONG&gt; thread context (no clean slate)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Reference only&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Best for&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Team procedures with execution&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Domain specialists&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Runbooks, architecture docs&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;Practical limits: a maximum of &lt;STRONG&gt;five concurrent active skills&lt;/STRONG&gt; (oldest auto-unloads), knowledge base uploads up to &lt;STRONG&gt;16 MB per file&lt;/STRONG&gt;, and custom agent knowledge bases up to &lt;STRONG&gt;1,000 files&lt;/STRONG&gt;.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="mcetoc_blog_8"&gt;3. Anatomy of an agent-run incident&lt;/H2&gt;
&lt;P&gt;Every use case in section 6 is an instance of this one shape. Learn it once.&lt;/P&gt;
&lt;P&gt;&lt;IMG src="https://techcommunity.microsoft.com/t5/s/gxcuf89792/uploaded_images/MzMzMzQzLVRrNnhUcA" alt="Four-phase flow of an agent-run incident: detect and route, investigate and document with no blast radius, decide and act behind an approval gate, then validate and close." width="2498" height="2685" style="max-width:100%;height:auto;display:block;margin:18px auto;" data-lia-image="" /&gt;&lt;/P&gt;
&lt;P&gt;Five properties make this shape safe, and they're worth stating as rules:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;The read phase has no blast radius.&lt;/STRONG&gt; Turn it on everywhere, immediately, with &lt;CODE&gt;Reader&lt;/CODE&gt;.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;One action per incident.&lt;/STRONG&gt; Not "roll back and scale and restart." One bounded, reversible step, then re-measure.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;The proposed action is always the smallest reversible one.&lt;/STRONG&gt; Swap a slot, don't redeploy. Bump one service tier, don't resize the cluster.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Validation is a first-class phase, not a vibe.&lt;/STRONG&gt; Define the metric, the threshold, and the duration before you approve.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Technical recovery ≠ business recovery.&lt;/STRONG&gt; Use case #10 makes this painfully clear.&lt;/LI&gt;
&lt;/OL&gt;
&lt;HR /&gt;
&lt;H2 id="mcetoc_blog_9"&gt;4. Build the demo lab&lt;/H2&gt;
&lt;P&gt;Everything below runs in a single throwaway resource group. Nothing here should touch a subscription you care about.&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;🧪 &lt;STRONG&gt;Lab hygiene.&lt;/STRONG&gt; Create it, demo it, delete it. &lt;CODE&gt;az group delete -n rg-sre-agent-demo --yes --no-wait&lt;/CODE&gt; when you're done. Several of these use cases deliberately break things.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3 id="mcetoc_blog_10"&gt;4.1 Resource group and agent&lt;/H3&gt;
&lt;LI-CODE lang="bash"&gt;LOC=eastus2
RG=rg-sre-agent-demo

az group create -n $RG -l $LOC
&lt;/LI-CODE&gt;
&lt;P&gt;Create the SRE Agent from the Azure portal and point it at &lt;CODE&gt;$RG&lt;/CODE&gt;. Three things are created for you automatically:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;an &lt;STRONG&gt;Application Insights&lt;/STRONG&gt; instance (this is where your audit trail lands),&lt;/LI&gt;
&lt;LI&gt;a &lt;STRONG&gt;Log Analytics workspace&lt;/STRONG&gt;,&lt;/LI&gt;
&lt;LI&gt;a &lt;STRONG&gt;user-assigned managed identity (UAMI)&lt;/STRONG&gt; — the identity every action runs as.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3 id="mcetoc_blog_11"&gt;4.2 Choose a permission level deliberately&lt;/H3&gt;
&lt;P&gt;At creation you pick one of two levels:&lt;/P&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Level&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Grants&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Use it when&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Reader&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Core monitoring roles + resource-type reader roles. Prompts for temporary elevation via on-behalf-of when it needs to act&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Start here.&lt;/STRONG&gt; Production. Weeks 1–4 of any pilot&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Privileged&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Core monitoring roles + resource-type &lt;STRONG&gt;contributor&lt;/STRONG&gt; roles&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Non-production, or after a proven pilot&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;Regardless of level, these are always assigned:&lt;/P&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Role&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Scope&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Why&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;Reader&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Resource group&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;See resources and properties&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;Log Analytics Reader&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Resource group&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Query logs and workspaces&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;Monitoring Reader&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Resource group&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read metrics&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;Monitoring Contributor&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Subscription&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Acknowledge and close Azure Monitor alerts&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;Grant additional access explicitly and narrowly:&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;SUB=$(az account show --query id -o tsv)
AGENT_MI=&amp;lt;agent-managed-identity-principal-id&amp;gt;

# Read everywhere you want visibility
az role assignment create \
  --assignee $AGENT_MI \
  --role "Reader" \
  --scope "/subscriptions/$SUB"

# Write ONLY where you intend the agent to act
az role assignment create \
  --assignee $AGENT_MI \
  --role "Website Contributor" \
  --scope "/subscriptions/$SUB/resourceGroups/$RG/providers/Microsoft.Web/sites/app-checkout-demo"
&lt;/LI-CODE&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;📌 &lt;STRONG&gt;A sharp edge.&lt;/STRONG&gt; You can't remove individual permissions from an agent — only entire resource groups. Removing a resource group from the agent's scope revokes all access to it. Plan your resource group boundaries as your &lt;EM&gt;blast-radius&lt;/EM&gt; boundaries, because that's exactly what they are.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3 id="mcetoc_blog_12"&gt;4.3 Connect ServiceNow&lt;/H3&gt;
&lt;P&gt;Use a &lt;STRONG&gt;dedicated, least-privileged integration user&lt;/STRONG&gt; — not a shared admin account.&lt;/P&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Method&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;When&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;What you need&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Basic auth&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Quick setup, PDI, testing&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Username + password, &lt;CODE&gt;itil&lt;/CODE&gt; or &lt;CODE&gt;admin&lt;/CODE&gt; role&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;OAuth 2.0&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Production&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;ServiceNow OAuth app (client ID + secret); register redirect &lt;CODE&gt;https://logic-apis-{region}.consent.azure-apim.net/redirect&lt;/CODE&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;Scope the connection so you don't drown:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Assignment group&lt;/STRONG&gt; — essential on a shared enterprise instance&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Priority&lt;/STRONG&gt; — Critical through Planning&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Category&lt;/STRONG&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Scanner defaults worth knowing:&lt;/P&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Setting&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Value&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Scan interval&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;1 minute&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Incidents per page&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;20&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Max incidents per cycle&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;220 (11 pages)&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Initial lookback&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;30 days&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;Setup performs a &lt;STRONG&gt;real connectivity check&lt;/STRONG&gt; by fetching an actual incident, so credential and endpoint mistakes surface immediately instead of six hours later when nothing syncs.&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;🚨 &lt;STRONG&gt;Delete the quickstart plan.&lt;/STRONG&gt; Connecting an incident platform auto-creates a &lt;CODE&gt;quickstart_handler&lt;/CODE&gt; response plan that runs in &lt;STRONG&gt;fully autonomous&lt;/STRONG&gt; mode across all impacted services. If you then build your own plans, incidents get routed twice or to the wrong agent. Go to &lt;STRONG&gt;Builder → Incident response plans → Table view&lt;/STRONG&gt; and delete it before you do anything else.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3 id="mcetoc_blog_13"&gt;4.4 Connect deployment correlation&lt;/H3&gt;
&lt;P&gt;Half the use cases below hinge on "what shipped three minutes before the spike." That correlation requires a source control connector:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;GitHub&lt;/STRONG&gt; — repositories and issues&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure DevOps&lt;/STRONG&gt; — repos and work items&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Connect one. Without it, the agent can still read Azure Activity Log and deployment history, but it can't reach commits, PRs, or work items.&lt;/P&gt;
&lt;H3 id="mcetoc_blog_14"&gt;4.5 Create the demo resources&lt;/H3&gt;
&lt;LI-CODE lang="bash"&gt;# 1 · App Service with a staging slot (use cases 1)
az appservice plan create -g $RG -n plan-demo --sku P1V3 --is-linux
az webapp create -g $RG -p plan-demo -n app-checkout-demo --runtime "DOTNETCORE:8.0"
az webapp deployment slot create -g $RG -n app-checkout-demo --slot previous

# 2 · AKS (use case 2)
az aks create -g $RG -n aks-demo --node-count 3 --generate-ssh-keys --enable-addons monitoring

# 3 · Azure SQL (use case 3)
az sql server create -g $RG -n sqlsrv-sre-demo -u sqladmin -p "&amp;lt;use-a-generated-password&amp;gt;"
az sql db create -g $RG -s sqlsrv-sre-demo -n db-customer --service-objective S1

# 4 · Cosmos DB (use case 4)
az cosmosdb create -g $RG -n cosmos-sre-demo
az cosmosdb sql database create -g $RG -a cosmos-sre-demo -n catalog
az cosmosdb sql container create -g $RG -a cosmos-sre-demo -d catalog \
  -n products --partition-key-path /category --throughput 400

# 5–7 · VMs (use cases 5, 6, 7)
az vm create -g $RG -n vm-payments-linux --image Ubuntu2204 --size Standard_B2s --generate-ssh-keys
az vm create -g $RG -n vm-claims-win --image Win2022Datacenter --size Standard_B2ms \
  --admin-username azureadmin --admin-password "&amp;lt;use-a-generated-password&amp;gt;"

# 8 · VM Scale Set (use case 8)
az vmss create -g $RG -n vmss-api-demo --image Ubuntu2204 --instance-count 12 \
  --upgrade-policy-mode automatic --generate-ssh-keys

# 10 · Service Bus (use case 10)
az servicebus namespace create -g $RG -n sb-sre-demo --sku Standard
az servicebus queue create -g $RG --namespace-name sb-sre-demo -n order-events \
  --enable-dead-lettering-on-message-expiration true
&lt;/LI-CODE&gt;
&lt;P&gt;Install the &lt;STRONG&gt;Azure Monitor Agent&lt;/STRONG&gt; on the VMs and associate a data collection rule — without guest telemetry, use cases 5, 6, and 7 have nothing to detect.&lt;/P&gt;
&lt;H3 id="mcetoc_blog_15"&gt;4.6 One custom agent per domain&lt;/H3&gt;
&lt;P&gt;Don't build a single mega-agent. Build specialists and let response plans route to them:&lt;/P&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Custom agent&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Owns&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Attached tools&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;DeploymentAnalyzer&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;App Service, AKS, anything release-correlated&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;RunAzCliReadCommands&lt;/CODE&gt;, GitHub connector, Kusto&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;DatabaseExpert&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Azure SQL, Cosmos DB&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;RunAzCliReadCommands&lt;/CODE&gt;, Kusto, read-only SQL diagnostics tool&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;GuestOSResponder&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;VM / VMSS / IIS&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;RunAzCliReadCommands&lt;/CODE&gt;, &lt;STRONG&gt;fixed-purpose runbook tools only&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;NetworkPathExpert&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Application Gateway, NSG, DNS, TLS&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;RunAzCliReadCommands&lt;/CODE&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;IntegrationExpert&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Service Bus, Event Hubs&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;RunAzCliReadCommands&lt;/CODE&gt;, Kusto&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;A custom agent definition is small:&lt;/P&gt;
&lt;LI-CODE lang="yaml"&gt;name: database_expert
system_prompt: |
  You are a database specialist for this estate. Analyze query performance,
  diagnose connection and saturation issues, and recommend the smallest
  reversible mitigation. Never propose schema changes, index changes, plan
  forcing, or session termination — those are DBA-owned and require a change record.
handoff_description: Handles Azure SQL and Cosmos DB troubleshooting
tools:
  - execute_kusto_query
  - RunAzCliReadCommands
allowed_skills:
  - azure-sql-saturation-runbook
  - cosmos-throughput-runbook
&lt;/LI-CODE&gt;
&lt;P&gt;Note what that &lt;CODE&gt;system_prompt&lt;/CODE&gt; is really doing: it is &lt;STRONG&gt;narrowing the action space&lt;/STRONG&gt;. Half of production safety with an agent is telling it, in plain English, which categories of fix are off the table.&lt;/P&gt;
&lt;H3 id="mcetoc_blog_16"&gt;4.7 Response plans&lt;/H3&gt;
&lt;P&gt;Create one plan per domain, all starting in &lt;STRONG&gt;Review&lt;/STRONG&gt;:&lt;/P&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Plan&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Filter&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Custom agent&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Mode&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;appsvc-p1&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Priority 1 + 2, service &lt;CODE&gt;checkout&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;DeploymentAnalyzer&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Review&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;aks-p1&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Priority 1, service &lt;CODE&gt;orders-api&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;DeploymentAnalyzer&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Review&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;db-critical&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Priority 1 + 2, service &lt;CODE&gt;customer-api&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;DatabaseExpert&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Review&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;vm-guest&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Priority 1 + 2, title contains &lt;CODE&gt;disk&lt;/CODE&gt;, &lt;CODE&gt;memory&lt;/CODE&gt;, &lt;CODE&gt;CPU&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;GuestOSResponder&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Review&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;network-p1&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Priority 1, title contains &lt;CODE&gt;502&lt;/CODE&gt;, &lt;CODE&gt;backend&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;NetworkPathExpert&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Review&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;integration-p1&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Priority 1, service &lt;CODE&gt;fulfillment&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;IntegrationExpert&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Review&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;Filters available: &lt;STRONG&gt;severity/priority&lt;/STRONG&gt; (multiselect), &lt;STRONG&gt;impacted service&lt;/STRONG&gt;, &lt;STRONG&gt;incident type&lt;/STRONG&gt;, and &lt;STRONG&gt;title contains&lt;/STRONG&gt;. Plans can be turned &lt;STRONG&gt;off&lt;/STRONG&gt; without deleting them, which is exactly what you want during maintenance windows.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="mcetoc_blog_17"&gt;5. Set your guardrails before your first incident&lt;/H2&gt;
&lt;P&gt;If you read only one section of this post, read this one. The guardrails are not paperwork; they are the reason this is deployable.&lt;/P&gt;
&lt;H3 id="mcetoc_blog_18"&gt;5.1 Run modes&lt;/H3&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Mode&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Behavior&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Default&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Review&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Agent proposes; you approve or deny&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Agent-level default&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Autonomous&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Agent executes immediately and reports&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Per-plan and per-task default&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;Two subtleties that bite people:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;Run modes are set per response plan and per scheduled task&lt;/STRONG&gt;, not globally. The agent-level setting is only a fallback. And per-plan the default is &lt;EM&gt;Autonomous&lt;/EM&gt; — so if you don't set it, you get autonomy.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Review mode shows Approve/Deny only for Azure infrastructure operations&lt;/STRONG&gt; (Azure CLI, ARM writes). Sending an email, posting to Teams, or querying an external source proceeds based on the agent's reasoning. To gate &lt;EM&gt;those&lt;/EM&gt;, you need &lt;A href="https://learn.microsoft.com/azure/sre-agent/agent-hooks" target="_blank"&gt;hooks&lt;/A&gt; or tool access policies.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Only users holding the &lt;STRONG&gt;SRE Agent Administrator&lt;/STRONG&gt; role can approve. Standard users cannot, and personal Microsoft accounts can't authorize on-behalf-of at all — it requires a work or school (Entra ID) account.&lt;/P&gt;
&lt;H3 id="mcetoc_blog_19"&gt;5.2 What the product blocks for you&lt;/H3&gt;
&lt;P&gt;These are enforced at the command level, independent of your RBAC:&lt;/P&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Guardrail&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Behavior&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Delete operations&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;The agent never runs &lt;CODE&gt;delete&lt;/CODE&gt; or &lt;CODE&gt;remove&lt;/CODE&gt; commands. It returns an error pointing you at the portal&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Key Vault&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;All &lt;CODE&gt;az keyvault&lt;/CODE&gt; commands are blocked, to prevent credential exposure&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Management locks&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Resources with &lt;CODE&gt;ReadOnly&lt;/CODE&gt; locks can't be modified, regardless of permissions or run mode&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Subscription validation&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Subscription IDs are validated as well-formed GUIDs before execution&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;💡 &lt;STRONG&gt;Use the delete block architecturally.&lt;/STRONG&gt; Put a &lt;CODE&gt;ReadOnly&lt;/CODE&gt; management lock on anything that must never change during an incident — your Key Vaults, your production databases, your golden images. That lock is respected before every modification, which gives you a control that survives RBAC drift and misconfigured response plans.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3 id="mcetoc_blog_20"&gt;5.3 Hooks: the guardrail you write yourself&lt;/H3&gt;
&lt;P&gt;Two events are supported — &lt;CODE&gt;Stop&lt;/CODE&gt; (agent about to return a final response) and &lt;CODE&gt;PostToolUse&lt;/CODE&gt; (a tool finished). Hooks run as an LLM &lt;STRONG&gt;prompt&lt;/STRONG&gt; or a sandboxed &lt;STRONG&gt;command&lt;/STRONG&gt; script.&lt;/P&gt;
&lt;P&gt;Here is the single most useful hook for this entire post — a deterministic policy gate on shell execution:&lt;/P&gt;
&lt;LI-CODE lang="yaml"&gt;hooks:
  PostToolUse:
    - type: command
      matcher: "Bash|ExecuteShellCommand"
      timeout: 30
      failMode: block
      script: |
        #!/usr/bin/env python3
        import sys, json, re

        context = json.load(sys.stdin)
        command = context.get('tool_input', {}).get('command', '')

        dangerous = [
            r'\brm\s+-rf\b',
            r'\bsudo\b',
            r'\bchmod\s+777\b',
            r'\bmkfs\b',
            r'\bdd\s+if=',
            r'\btruncate\b',
            r'\bDROP\s+TABLE\b',
        ]

        for pattern in dangerous:
            if re.search(pattern, command, re.IGNORECASE):
                print(json.dumps({"decision": "block",
                                  "reason": f"Blocked by policy: {pattern}"}))
                sys.exit(0)

        print(json.dumps({"decision": "allow"}))
&lt;/LI-CODE&gt;
&lt;P&gt;And a &lt;CODE&gt;Stop&lt;/CODE&gt; hook that refuses to let the agent declare victory without evidence:&lt;/P&gt;
&lt;LI-CODE lang="yaml"&gt;hooks:
  Stop:
    - type: prompt
      model: ReasoningFast
      timeout: 30
      prompt: |
        Review the agent's final response.
        $ARGUMENTS

        It is only acceptable if it contains ALL of:
          1. The specific resource acted upon (full name or resource ID)
          2. Metric values BEFORE and AFTER the action
          3. The observation window over which recovery was confirmed
          4. Who approved the action and at what UTC timestamp

        Respond with:
          {"ok": true}
          {"ok": false, "reason": "&amp;lt;what is missing&amp;gt;"}
&lt;/LI-CODE&gt;
&lt;P&gt;Hook mechanics you'll want on hand:&lt;/P&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Setting&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Default&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Range / notes&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;timeout&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;30s&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;1–300s&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;failMode&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;allow&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;allow&lt;/CODE&gt; or &lt;CODE&gt;block&lt;/CODE&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;maxRejections&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;3&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;1–25; prompt-type &lt;CODE&gt;Stop&lt;/CODE&gt; hooks only&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;matcher&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;—&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Regex, anchored &lt;CODE&gt;^(pattern)$&lt;/CODE&gt;, case-sensitive; &lt;CODE&gt;*&lt;/CODE&gt; matches all&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Script size&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;—&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;64 KB max&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Shebangs&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;—&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;#!/bin/bash&lt;/CODE&gt;, &lt;CODE&gt;#!/usr/bin/env python3&lt;/CODE&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;⚠️ &lt;STRONG&gt;The footgun:&lt;/STRONG&gt; for &lt;CODE&gt;Stop&lt;/CODE&gt; hooks, a rejection &lt;STRONG&gt;without a &lt;CODE&gt;reason&lt;/CODE&gt; field is treated as approval&lt;/STRONG&gt;. Always populate &lt;CODE&gt;reason&lt;/CODE&gt;.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;Agent-level hooks and custom-agent-level hooks both run when both match; agent-level fires first.&lt;/P&gt;
&lt;H3 id="mcetoc_blog_21"&gt;5.4 The recommended production policy&lt;/H3&gt;
&lt;P&gt;This is the policy I'd put in front of a change board:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Allow &lt;STRONG&gt;read-only investigation&lt;/STRONG&gt; automatically, everywhere.&lt;/LI&gt;
&lt;LI&gt;Allow &lt;STRONG&gt;automatic incident creation and work-note updates&lt;/STRONG&gt; — after sanitization.&lt;/LI&gt;
&lt;LI&gt;Use &lt;STRONG&gt;Review mode for all production remediation&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI&gt;Require &lt;STRONG&gt;human approval for every production write&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI&gt;Use &lt;STRONG&gt;fixed-purpose runbooks&lt;/STRONG&gt; instead of unrestricted VM shell access.&lt;/LI&gt;
&lt;LI&gt;Require &lt;STRONG&gt;separate approval&lt;/STRONG&gt; for destructive or data-affecting actions.&lt;/LI&gt;
&lt;LI&gt;Initially require &lt;STRONG&gt;human confirmation before resolving&lt;/STRONG&gt; an incident.&lt;/LI&gt;
&lt;/UL&gt;
&lt;HR /&gt;
&lt;H2 id="mcetoc_blog_22"&gt;6. The ten use cases&lt;/H2&gt;
&lt;P&gt;Each use case follows the same seven-part structure so you can skim to the one you're firefighting:&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;What happened&lt;/STRONG&gt; → &lt;STRONG&gt;How Azure Monitor detected it&lt;/STRONG&gt; → &lt;STRONG&gt;How the agent found the cause&lt;/STRONG&gt; → &lt;STRONG&gt;Agent action table&lt;/STRONG&gt; → &lt;STRONG&gt;Recovery note&lt;/STRONG&gt; → &lt;STRONG&gt;Try it yourself&lt;/STRONG&gt; → &lt;STRONG&gt;Guardrails&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;In the action tables, &lt;CODE&gt;Type&lt;/CODE&gt; is one of:&lt;/P&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Type&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Meaning&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;Read&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;No blast radius. Safe to run automatically&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;Decision&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Agent reasoning or an approval gate. No resource change&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;Write (ITSM)&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Ticket create/update. Sanitized&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;Write (Azure)&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Azure control-plane change. &lt;STRONG&gt;Approval required in production&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;Write (Guest OS)&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Inside the VM, via a fixed-purpose runbook. &lt;STRONG&gt;Approval required&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;Write (K8s/DevOps)&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Cluster or pipeline change. &lt;STRONG&gt;Approval required&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;Validation&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Post-action measurement against defined thresholds&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;HR /&gt;
&lt;H3 id="mcetoc_blog_23"&gt;Use case #1 — App Service: HTTP 500 after deployment&lt;/H3&gt;
&lt;P&gt;&lt;EM&gt;The one you should build first. Clean trigger, bounded action, unambiguous validation.&lt;/EM&gt;&lt;/P&gt;
&lt;H4&gt;What happened&lt;/H4&gt;
&lt;P&gt;Release &lt;CODE&gt;2026.03.12.4&lt;/CODE&gt; was deployed to a production checkout App Service. The new code referenced an application setting that was never defined in the production slot. The application started fine — which is what makes this class of failure nasty — but every checkout operation that touched that setting returned HTTP 500.&lt;/P&gt;
&lt;H4&gt;How Azure Monitor detected it&lt;/H4&gt;
&lt;P&gt;Azure Monitor and Application Insights fired on a composite condition, not a single metric:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;HTTP 5xx rate above the operational threshold&lt;/LI&gt;
&lt;LI&gt;Failed availability tests&lt;/LI&gt;
&lt;LI&gt;Increased application exceptions&lt;/LI&gt;
&lt;LI&gt;Increased dependency failures&lt;/LI&gt;
&lt;LI&gt;Request latency above baseline&lt;/LI&gt;
&lt;/UL&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;P1 – Checkout App Service returning elevated HTTP 500 responses&lt;/STRONG&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;How the agent found the cause&lt;/H4&gt;
&lt;OL&gt;
&lt;LI&gt;Identified the affected App Service and slot.&lt;/LI&gt;
&lt;LI&gt;Queried HTTP response and latency metrics.&lt;/LI&gt;
&lt;LI&gt;Queried Application Insights failed requests and exceptions.&lt;/LI&gt;
&lt;LI&gt;Identified the missing-setting exception as the dominant failure.&lt;/LI&gt;
&lt;LI&gt;Reviewed deployment history and Azure Activity Log.&lt;/LI&gt;
&lt;LI&gt;Correlated the error increase with release &lt;CODE&gt;2026.03.12.4&lt;/CODE&gt;.&lt;/LI&gt;
&lt;LI&gt;Compared production against the previous deployment slot.&lt;/LI&gt;
&lt;LI&gt;Confirmed the previous slot remained healthy.&lt;/LI&gt;
&lt;LI&gt;Confirmed no matching Azure platform incident.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;STRONG&gt;Agent finding:&lt;/STRONG&gt;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;HTTP 500 responses increased from 0.4% to 18% within three minutes of release 2026.03.12.4. Most failures reference a missing application setting. The previous deployment slot passes availability and dependency checks.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;Notice the shape of that sentence: a &lt;STRONG&gt;delta&lt;/STRONG&gt;, a &lt;STRONG&gt;time correlation&lt;/STRONG&gt;, and a &lt;STRONG&gt;known-good comparison&lt;/STRONG&gt;. That's what makes it actionable rather than merely true.&lt;/P&gt;
&lt;H4&gt;Agent action table&lt;/H4&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;#&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Action&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Type&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Detail&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;1&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Investigate alert&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Reads HTTP 5xx rate, failed requests, latency, availability-test results&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;2&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Analyze exceptions&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Identifies missing-configuration exception as dominant failure&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;3&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Correlate deployment&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Finds release 2026.03.12.4 deployed three minutes before the spike&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;4&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Compare slots&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Confirms previous slot healthy while production is failing&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;5&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Rule out platform issue&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Checks Azure Resource Health and dependency health&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;6&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Assess blast radius&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Determines checkout affected, unrelated services healthy&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;7&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;P1 against the Checkout CI, assigned to Application Operations&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;8&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update INC evidence&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Adds exceptions, deployment correlation, resource links, affected operations&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;9&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Classify response&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Classifies as deployment-caused; selects rollback&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;10&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Pause release&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (DevOps)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Pauses the failed release pipeline&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;11&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Request approval&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Requests approval to swap the previous healthy slot into production&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;12&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Swap slot&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Azure)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Performs the approved swap &lt;STRONG&gt;only&lt;/STRONG&gt; on the named App Service&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;13&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validate recovery&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Confirms 5xx, latency, exceptions, dependencies, availability recover&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;14&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update/resolve INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Records approval, rollback, timestamps, recovery evidence&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;15&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW PRB&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Problem record for configuration-validation improvements&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;H4&gt;Recovery note&lt;/H4&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Production was reverted to the previous healthy deployment slot at 14:26 UTC. HTTP 500 responses declined from 18% to below 1%, and availability tests passed for 15 consecutive minutes. Preliminary cause: missing production configuration in release 2026.03.12.4.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;Try it yourself&lt;/H4&gt;
&lt;P&gt;&lt;STRONG&gt;Break it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;# Put a healthy build in the 'previous' slot first, then break production
# by deploying code that reads an app setting which only exists in staging.
az webapp config appsettings delete -g $RG -n app-checkout-demo \
  --setting-names Checkout__PaymentProviderKey

az webapp restart -g $RG -n app-checkout-demo
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Alert it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;az monitor metrics alert create -g $RG -n "alert-checkout-5xx" \
  --scopes "/subscriptions/$SUB/resourceGroups/$RG/providers/Microsoft.Web/sites/app-checkout-demo" \
  --condition "total Http5xx &amp;gt; 20" \
  --window-size 5m --evaluation-frequency 1m --severity 1 \
  --description "P1 – Checkout App Service returning elevated HTTP 500 responses"
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Ask it (before the alert, to see the read phase in isolation):&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="text"&gt;app-checkout-demo is returning HTTP 500s. Do not change anything.
Investigate and tell me:
  1. the dominant exception and its share of total failures
  2. what deployed in the 30 minutes before the error rate changed
  3. whether the 'previous' slot is healthy right now
  4. whether Azure Resource Health shows a platform issue
Give me the evidence chain and the smallest reversible mitigation.
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Expect it to propose:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="text"&gt;Proposed action: Swap slot 'previous' into production on app-checkout-demo
Risk: brief connection drain (~10s). Fully reversible by swapping back.
Validation: Http5xx &amp;lt; 1% and availability test passing for 15 minutes.

[Approve] [Deny]
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Verify it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="kusto"&gt;requests
| where timestamp &amp;gt; ago(1h)
| summarize failed = countif(success == false), total = count() by bin(timestamp, 1m)
| extend failureRate = 100.0 * failed / total
| render timechart
&lt;/LI-CODE&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;🎬 &lt;STRONG&gt;Best thing to record for a demo:&lt;/STRONG&gt; the moment the Approve button appears with the evidence already attached. That single screen is the whole value proposition.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;Guardrails&lt;/H4&gt;
&lt;UL&gt;
&lt;LI&gt;Scope write permissions to the &lt;STRONG&gt;named App Service&lt;/STRONG&gt;, not the resource group.&lt;/LI&gt;
&lt;LI&gt;Verify the previous slot is healthy &lt;STRONG&gt;before&lt;/STRONG&gt; swapping — a swap into a broken slot doubles the outage.&lt;/LI&gt;
&lt;LI&gt;Require human approval for production slot swaps.&lt;/LI&gt;
&lt;LI&gt;Preserve the failed deployment for analysis; don't let the pipeline overwrite it.&lt;/LI&gt;
&lt;LI&gt;Never grant subscription-level Contributor or Owner.&lt;/LI&gt;
&lt;/UL&gt;
&lt;HR /&gt;
&lt;H3 id="mcetoc_blog_24"&gt;Use case #2 — AKS: pods in CrashLoopBackOff&lt;/H3&gt;
&lt;H4&gt;What happened&lt;/H4&gt;
&lt;P&gt;Image &lt;CODE&gt;orders-api:4.18.0&lt;/CODE&gt; referenced a configuration key that didn't exist in the production namespace. Eight of ten pods entered &lt;CODE&gt;CrashLoopBackOff&lt;/CODE&gt;. The two survivors couldn't carry production traffic, producing latency and ingress 5xx.&lt;/P&gt;
&lt;H4&gt;How Azure Monitor detected it&lt;/H4&gt;
&lt;P&gt;Managed Prometheus / Container Insights detected:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Unavailable replicas&lt;/LI&gt;
&lt;LI&gt;Increasing container restart count&lt;/LI&gt;
&lt;LI&gt;&lt;CODE&gt;CrashLoopBackOff&lt;/CODE&gt; pod state&lt;/LI&gt;
&lt;LI&gt;Failed readiness and liveness checks&lt;/LI&gt;
&lt;LI&gt;Elevated ingress HTTP 5xx&lt;/LI&gt;
&lt;LI&gt;Reduced successful-request rate&lt;/LI&gt;
&lt;/UL&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;P1 – Orders API unavailable replicas in production AKS&lt;/STRONG&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;How the agent found the cause&lt;/H4&gt;
&lt;OL&gt;
&lt;LI&gt;Identified the cluster, namespace, and deployment.&lt;/LI&gt;
&lt;LI&gt;Reviewed deployment availability and pod states.&lt;/LI&gt;
&lt;LI&gt;Examined logs from failing containers.&lt;/LI&gt;
&lt;LI&gt;Reviewed Kubernetes warning events.&lt;/LI&gt;
&lt;LI&gt;Compared current and previous ReplicaSets.&lt;/LI&gt;
&lt;LI&gt;Correlated the failure with revision 42.&lt;/LI&gt;
&lt;LI&gt;Identified the missing configuration key.&lt;/LI&gt;
&lt;LI&gt;Confirmed revision 41 was previously healthy.&lt;/LI&gt;
&lt;LI&gt;Checked node CPU, memory, storage, networking, status.&lt;/LI&gt;
&lt;LI&gt;Determined AKS infrastructure was healthy.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;STRONG&gt;Agent finding:&lt;/STRONG&gt;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Eight of ten Orders API pods entered CrashLoopBackOff during rollout of revision 42. Container logs show a missing configuration key. Revision 41, using image 4.17.6, was healthy.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;Step 10 is the one people skip. "Rule out the infrastructure" is what stops you from spending an hour on a node pool that was never the problem.&lt;/P&gt;
&lt;H4&gt;Agent action table&lt;/H4&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;#&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Action&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Type&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Detail&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;1&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Investigate workload&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Deployment status, unavailable replicas, readiness, restart counts&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;2&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Analyze pod logs&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Startup failure from a missing configuration key&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;3&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Analyze events&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Image pulls, mounts, scheduling, probes, container events&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;4&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Correlate rollout&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Revision 42 deployed immediately before failure&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;5&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Compare ReplicaSets&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Revision 41 was the previous healthy workload&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;6&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Rule out infrastructure&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Nodes, memory, CPU, networking, storage healthy&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;7&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Assess blast radius&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Eight of ten replicas unavailable&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;8&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;P1 against the Orders API CI&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;9&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update INC evidence&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Namespace, image, revision, errors, rollout correlation&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;10&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Classify response&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Application/configuration-caused&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;11&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Request approval&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Approval to pause and roll back revision 42&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;12&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Pause failed rollout&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (K8s/DevOps)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Prevents the release progressing or being reapplied&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;13&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Roll back deployment&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (K8s)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Rolls back &lt;STRONG&gt;only&lt;/STRONG&gt; the named deployment to revision 41&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;14&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validate replicas&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;All expected replicas ready and stable&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;15&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validate application&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Ingress errors decline; synthetic orders succeed&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;16&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update/resolve INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Revision, approval, rollback, recovery evidence&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;17&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create defect/task&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Add configuration validation to CI/CD&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;H4&gt;Recovery note&lt;/H4&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Orders API was rolled back from revision 42 to revision 41. All ten replicas are ready, restart counts have stabilized, ingress HTTP 5xx responses are below threshold, and synthetic order transactions are succeeding.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;Try it yourself&lt;/H4&gt;
&lt;P&gt;&lt;STRONG&gt;Break it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;az aks get-credentials -g $RG -n aks-demo

kubectl create namespace orders
kubectl create deployment orders-api -n orders --image=nginx:1.25 --replicas=10
kubectl rollout status deployment/orders-api -n orders   # revision 41 equivalent, healthy

# Now break it: point at an image whose entrypoint requires a missing env var
kubectl set image deployment/orders-api -n orders orders-api=busybox:1.36 
kubectl patch deployment orders-api -n orders --type=json -p='[
  {"op":"add","path":"/spec/template/spec/containers/0/command",
   "value":["sh","-c","test -n \"$ORDERS_CONFIG_KEY\" || (echo \"FATAL: missing ORDERS_CONFIG_KEY\" &amp;gt;&amp;amp;2; exit 1); sleep 3600"]}
]'
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Watch it break:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;kubectl get pods -n orders -w
kubectl rollout history deployment/orders-api -n orders
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Ask it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="text"&gt;Pods in namespace 'orders' on aks-demo are crash looping. Read-only.
Tell me: which revision introduced it, the exact container error,
whether the node pool is healthy, and how many replicas are actually serving.
Then tell me the last known-good revision and why you believe it was healthy.
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Expect it to propose:&lt;/STRONG&gt; &lt;CODE&gt;kubectl rollout undo deployment/orders-api -n orders --to-revision=&amp;lt;n&amp;gt;&lt;/CODE&gt; scoped to that single deployment.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Verify it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;kubectl get deployment orders-api -n orders \
  -o jsonpath='{.status.readyReplicas}/{.status.replicas}{"\n"}'
kubectl get pods -n orders --no-headers | awk '{print $4}' | sort | uniq -c
&lt;/LI-CODE&gt;
&lt;H4&gt;Guardrails&lt;/H4&gt;
&lt;UL&gt;
&lt;LI&gt;Restrict access to the required &lt;STRONG&gt;cluster, namespace, and deployment&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI&gt;Do &lt;STRONG&gt;not&lt;/STRONG&gt; grant &lt;CODE&gt;cluster-admin&lt;/CODE&gt;.&lt;/LI&gt;
&lt;LI&gt;Require approval for rollback.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Coordinate with GitOps reconciliation.&lt;/STRONG&gt; If Flux or Argo owns that deployment, a &lt;CODE&gt;kubectl rollout undo&lt;/CODE&gt; gets reverted within minutes and you've created a flapping outage. Either pause reconciliation first or roll back through Git.&lt;/LI&gt;
&lt;LI&gt;Do not permit namespace, persistent-volume, or cluster deletion.&lt;/LI&gt;
&lt;/UL&gt;
&lt;HR /&gt;
&lt;H3 id="mcetoc_blog_25"&gt;Use case #3 — Azure SQL Database: CPU saturation and timeouts&lt;/H3&gt;
&lt;P&gt;&lt;EM&gt;The use case where the agent's job is to buy time, not to fix the problem.&lt;/EM&gt;&lt;/P&gt;
&lt;H4&gt;What happened&lt;/H4&gt;
&lt;P&gt;A query execution plan changed after an application release. The new plan consumed substantially more CPU and workers. The database saturated, producing SQL dependency timeouts and failed customer requests.&lt;/P&gt;
&lt;H4&gt;How Azure Monitor detected it&lt;/H4&gt;
&lt;UL&gt;
&lt;LI&gt;Sustained CPU saturation&lt;/LI&gt;
&lt;LI&gt;Worker or session pressure&lt;/LI&gt;
&lt;LI&gt;Increased connection failures&lt;/LI&gt;
&lt;LI&gt;SQL dependency timeouts&lt;/LI&gt;
&lt;LI&gt;Query-duration deviation from baseline&lt;/LI&gt;
&lt;LI&gt;Degraded customer API success rate&lt;/LI&gt;
&lt;/UL&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;P1 – Customer API database saturation causing request timeouts&lt;/STRONG&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;How the agent found the cause&lt;/H4&gt;
&lt;OL&gt;
&lt;LI&gt;Reviewed database CPU, workers, sessions, connections.&lt;/LI&gt;
&lt;LI&gt;Reviewed application SQL dependency failures.&lt;/LI&gt;
&lt;LI&gt;Checked for blocking and deadlocks.&lt;/LI&gt;
&lt;LI&gt;Ran approved &lt;STRONG&gt;read-only&lt;/STRONG&gt; Query Store diagnostics.&lt;/LI&gt;
&lt;LI&gt;Identified the primary CPU-consuming query.&lt;/LI&gt;
&lt;LI&gt;Detected a recent execution-plan change.&lt;/LI&gt;
&lt;LI&gt;Correlated with an application release.&lt;/LI&gt;
&lt;LI&gt;Ruled out Azure service health and storage issues.&lt;/LI&gt;
&lt;LI&gt;Determined a temporary scale operation could restore service.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Left permanent query remediation to the DBA team.&lt;/STRONG&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;STRONG&gt;Agent finding:&lt;/STRONG&gt;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Database CPU has remained saturated for 17 minutes. One query accounts for most recent CPU consumption and changed execution plan shortly before the incident. Application SQL dependency timeout rate is 23%.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;Step 10 is the design decision that makes this safe. The agent correctly diagnoses a plan regression and then &lt;EM&gt;deliberately does not fix it&lt;/EM&gt;, because forcing a plan or dropping an index is a permanent, DBA-owned, change-controlled action. It buys capacity instead.&lt;/P&gt;
&lt;H4&gt;Agent action table&lt;/H4&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;#&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Action&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Type&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Detail&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;1&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Investigate database&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;CPU, workers, sessions, connections, storage, availability&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;2&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Analyze app impact&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;SQL dependency latency, failures, affected API operations&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;3&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Check blocking&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Approved diagnostics for blocking, deadlocks, connection growth&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;4&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Analyze Query Store&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Highest-impact query and recent plan change&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;5&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Correlate changes&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Recent application and database deployments&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;6&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Rule out platform issue&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Service health, storage, database availability&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;7&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Assess blast radius&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Affected applications; other databases healthy&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;8&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;P1 against the production database CI&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;9&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update INC evidence&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Utilization, query ID, timeouts, change correlation&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;10&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Classify response&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Scaling is mitigation; query changes remain DBA-owned&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;11&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Calculate bounded scale&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Selects the &lt;STRONG&gt;smallest&lt;/STRONG&gt; pre-approved capacity increase&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;12&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Request approval&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Database owner or incident commander&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;13&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Scale database&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Azure)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Increases the affected database by one approved service step&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;14&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validate recovery&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;CPU, workers, timeouts, application success recover&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;15&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Previous/new capacity, approval, timing, results&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;16&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW PRB&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Permanent query-remediation work&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;17&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create scale-down task&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Task/change to restore normal capacity after stability&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;H4&gt;Recovery note&lt;/H4&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Azure SQL capacity was temporarily increased by one approved service step. CPU declined from sustained saturation to 54%, and SQL dependency timeouts returned to baseline. Query Store indicates a probable execution-plan regression requiring permanent DBA remediation.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;Try it yourself&lt;/H4&gt;
&lt;P&gt;&lt;STRONG&gt;Break it&lt;/STRONG&gt; — generate a saturating workload against &lt;CODE&gt;db-customer&lt;/CODE&gt; (an unindexed &lt;CODE&gt;LIKE '%...%'&lt;/CODE&gt; scan in a tight loop from a container in the same region works fine on an S1).&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Alert it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;az monitor metrics alert create -g $RG -n "alert-sql-cpu" \
  --scopes "/subscriptions/$SUB/resourceGroups/$RG/providers/Microsoft.Sql/servers/sqlsrv-sre-demo/databases/db-customer" \
  --condition "avg cpu_percent &amp;gt; 90" \
  --window-size 5m --evaluation-frequency 1m --severity 1 \
  --description "P1 – Customer API database saturation causing request timeouts"
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Ask it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="text"&gt;db-customer is saturated. Read-only investigation.
Identify the top CPU-consuming query, whether its plan changed recently,
and what application release correlates.
Do NOT propose index, schema, plan-forcing, or session-kill actions.
Propose only the smallest temporary capacity step that restores service,
and tell me what it costs per day and when we should scale back down.
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Read-only Query Store diagnostics the agent should run:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="sql"&gt;SELECT TOP 10
    qsq.query_id,
    qsp.plan_id,
    qsp.last_execution_time,
    SUM(qsrs.count_executions)                       AS executions,
    SUM(qsrs.avg_cpu_time * qsrs.count_executions)   AS total_cpu_us
FROM sys.query_store_query           AS qsq
JOIN sys.query_store_plan            AS qsp  ON qsp.query_id = qsq.query_id
JOIN sys.query_store_runtime_stats   AS qsrs ON qsrs.plan_id = qsp.plan_id
JOIN sys.query_store_runtime_stats_interval AS qsrsi
     ON qsrsi.runtime_stats_interval_id = qsrs.runtime_stats_interval_id
WHERE qsrsi.start_time &amp;gt; DATEADD(hour, -2, GETUTCDATE())
GROUP BY qsq.query_id, qsp.plan_id, qsp.last_execution_time
ORDER BY total_cpu_us DESC;
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Expect it to propose:&lt;/STRONG&gt; &lt;CODE&gt;az sql db update -g $RG -s sqlsrv-sre-demo -n db-customer --service-objective S2&lt;/CODE&gt; — exactly one step, on exactly that database.&lt;/P&gt;
&lt;H4&gt;Guardrails&lt;/H4&gt;
&lt;UL&gt;
&lt;LI&gt;Restrict scaling to the &lt;STRONG&gt;named database&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI&gt;Define minimum and maximum capacity.&lt;/LI&gt;
&lt;LI&gt;Require approval for scale-up &lt;STRONG&gt;and&lt;/STRONG&gt; scale-down.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Time-limit temporary capacity&lt;/STRONG&gt; — an un-reversed emergency scale-up is how a P1 becomes a budget incident.&lt;/LI&gt;
&lt;LI&gt;Do not autonomously force plans, terminate sessions, modify indexes, or change schema.&lt;/LI&gt;
&lt;LI&gt;Track the temporary cost impact with an Azure Cost Management alert.&lt;/LI&gt;
&lt;/UL&gt;
&lt;HR /&gt;
&lt;H3 id="mcetoc_blog_26"&gt;Use case #4 — Azure Cosmos DB: HTTP 429 throttling&lt;/H3&gt;
&lt;P&gt;&lt;EM&gt;The one where the correct root cause is "we're succeeding."&lt;/EM&gt;&lt;/P&gt;
&lt;H4&gt;What happened&lt;/H4&gt;
&lt;P&gt;A marketing campaign increased Product Catalog traffic by roughly 40%. The Cosmos DB container hit its provisioned throughput ceiling. HTTP 429s increased, and client retries amplified application latency.&lt;/P&gt;
&lt;H4&gt;How Azure Monitor detected it&lt;/H4&gt;
&lt;UL&gt;
&lt;LI&gt;High normalized RU consumption&lt;/LI&gt;
&lt;LI&gt;Increased HTTP 429 responses&lt;/LI&gt;
&lt;LI&gt;Elevated server-side latency&lt;/LI&gt;
&lt;LI&gt;Application dependency failures&lt;/LI&gt;
&lt;LI&gt;Sustained operation near the throughput limit&lt;/LI&gt;
&lt;/UL&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;P2 – Cosmos DB throttling affecting Product Catalog requests&lt;/STRONG&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;How the agent found the cause&lt;/H4&gt;
&lt;OL&gt;
&lt;LI&gt;Reviewed normalized RU consumption and throttled requests.&lt;/LI&gt;
&lt;LI&gt;Identified the affected database and container.&lt;/LI&gt;
&lt;LI&gt;Reviewed regional and partition behavior.&lt;/LI&gt;
&lt;LI&gt;Checked for hot-partition evidence.&lt;/LI&gt;
&lt;LI&gt;Reviewed application retry telemetry.&lt;/LI&gt;
&lt;LI&gt;Compared current traffic with the historical baseline.&lt;/LI&gt;
&lt;LI&gt;Correlated demand with the marketing campaign.&lt;/LI&gt;
&lt;LI&gt;Checked recent application deployments.&lt;/LI&gt;
&lt;LI&gt;Checked Azure service health.&lt;/LI&gt;
&lt;LI&gt;Determined the primary cause was &lt;STRONG&gt;legitimate demand&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;STRONG&gt;Agent finding:&lt;/STRONG&gt;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;The Product Catalog container is at its configured throughput ceiling, and 21% of requests are being throttled. Traffic increased by approximately 40% following a campaign launch. No deployment or regional platform issue correlates with the event.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;Step 4 is the fork in the road. If consumption is uneven across partitions, more RU/s is money set on fire — the correct answer is an architecture change, not a scale-up. The agent has to check before it recommends.&lt;/P&gt;
&lt;H4&gt;Agent action table&lt;/H4&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;#&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Action&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Type&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Detail&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;1&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Investigate throttling&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;RU consumption, 429 rate, latency, requests, availability&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;2&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Identify scope&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Account, database, container, operations, regions&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;3&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Analyze demand&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Compares traffic and RU consumption with historical patterns&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;4&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Check partitions&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Looks for uneven partition consumption where telemetry permits&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;5&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Analyze retries&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Determines whether client retries are amplifying the incident&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;6&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Correlate events&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Links demand to campaign traffic; excludes release/platform issues&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;7&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Assess blast radius&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Affected catalog operations; unaffected containers&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;8&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;P2 against the Product Catalog CI&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;9&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update INC evidence&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;RU, throttling, latency, traffic, partition evidence&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;10&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Classify response&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Demand-driven &lt;STRONG&gt;unless&lt;/STRONG&gt; hot-partition evidence exists&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;11&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Calculate throughput&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Smallest increase within the cost ceiling&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;12&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Request approval&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Approval for a temporary throughput increase&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;13&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Increase throughput&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Azure)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Raises throughput only to the approved maximum&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;14&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validate recovery&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;429 rate and latency recover&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;15&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update/resolve INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Throughput, approval, cost implication, recovery&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;16&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create capacity task&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Work to return throughput to normal&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;17&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW PRB&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Partition or retry improvements, if required&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;H4&gt;Recovery note&lt;/H4&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Provisioned throughput was increased within the approved production limit. HTTP 429 responses declined from 21% to below 1%, and Product Catalog latency returned to baseline. The increase is temporary and will be reviewed after campaign traffic subsides.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;Try it yourself&lt;/H4&gt;
&lt;P&gt;&lt;STRONG&gt;Break it:&lt;/STRONG&gt; the container was created at 400 RU/s. Drive a few hundred reads per second at it and you'll be throttled within seconds.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Alert it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;az monitor metrics alert create -g $RG -n "alert-cosmos-429" \
  --scopes "/subscriptions/$SUB/resourceGroups/$RG/providers/Microsoft.DocumentDB/databaseAccounts/cosmos-sre-demo" \
  --condition "total TotalRequests where StatusCode == 429 &amp;gt; 100" \
  --window-size 5m --evaluation-frequency 1m --severity 2 \
  --description "P2 – Cosmos DB throttling affecting Product Catalog requests"
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Ask it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="text"&gt;cosmos-sre-demo container 'products' is throttling. Read-only.
Before recommending anything, tell me whether RU consumption is EVEN across
physical partitions or concentrated. If it is concentrated, do not recommend
a throughput increase — recommend an architecture problem record instead.
If it is even, tell me the smallest RU/s that clears throttling and the
daily cost delta.
&lt;/LI-CODE&gt;
&lt;P&gt;That prompt is the whole use case. Getting the agent to &lt;EM&gt;refuse&lt;/EM&gt; the easy answer under a stated condition is the skill.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Verify it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="kusto"&gt;AzureDiagnostics
| where ResourceProvider == "MICROSOFT.DOCUMENTDB"
| where Category == "DataPlaneRequests"
| summarize throttled = countif(statusCode_s == "429"), total = count()
        by bin(TimeGenerated, 1m)
| extend throttleRate = 100.0 * throttled / total
| render timechart
&lt;/LI-CODE&gt;
&lt;H4&gt;Guardrails&lt;/H4&gt;
&lt;UL&gt;
&lt;LI&gt;Define a &lt;STRONG&gt;maximum throughput ceiling&lt;/STRONG&gt; the agent may not exceed.&lt;/LI&gt;
&lt;LI&gt;Restrict changes to the &lt;STRONG&gt;named container&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI&gt;Require approval for increases &lt;EM&gt;and&lt;/EM&gt; reductions.&lt;/LI&gt;
&lt;LI&gt;Do not permit deletion, consistency-level changes, or region changes.&lt;/LI&gt;
&lt;LI&gt;Create cost alerts for prolonged increased throughput.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Treat persistent hot partitions as an architecture issue&lt;/STRONG&gt;, never as a scaling issue.&lt;/LI&gt;
&lt;/UL&gt;
&lt;HR /&gt;
&lt;H3 id="mcetoc_blog_27"&gt;Use case #5 — Azure VM: OS/root disk full&lt;/H3&gt;
&lt;P&gt;&lt;EM&gt;The most operationally dangerous use case in this post, and the one with the most interesting guardrail design.&lt;/EM&gt;&lt;/P&gt;
&lt;H4&gt;What happened&lt;/H4&gt;
&lt;P&gt;A legacy payment application generated excessive trace logs. Log rotation stopped working and the root filesystem filled. The VM stayed available at the Azure platform layer — heartbeat green, Resource Health fine — but the application stopped, because it could no longer write to disk.&lt;/P&gt;
&lt;P&gt;This is the classic "green dashboard, dead service" failure. Platform-layer monitoring alone will never catch it.&lt;/P&gt;
&lt;H4&gt;How Azure Monitor detected it&lt;/H4&gt;
&lt;P&gt;Azure Monitor Agent and guest telemetry detected:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Critically low filesystem free space&lt;/LI&gt;
&lt;LI&gt;Rapid filesystem consumption&lt;/LI&gt;
&lt;LI&gt;Application process stopped&lt;/LI&gt;
&lt;LI&gt;Failed availability tests&lt;/LI&gt;
&lt;LI&gt;Elevated HTTP 5xx&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Healthy VM heartbeat but unhealthy application&lt;/STRONG&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;P1 – Legacy payment application unavailable due to full VM OS disk&lt;/STRONG&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;How the agent found the cause&lt;/H4&gt;
&lt;OL&gt;
&lt;LI&gt;Confirmed the VM was online.&lt;/LI&gt;
&lt;LI&gt;Confirmed Azure Monitor Agent heartbeat.&lt;/LI&gt;
&lt;LI&gt;Identified the affected root/OS filesystem.&lt;/LI&gt;
&lt;LI&gt;Reviewed the free-space trend.&lt;/LI&gt;
&lt;LI&gt;Correlated application failure with disk exhaustion.&lt;/LI&gt;
&lt;LI&gt;Found 86 GB of growth in the application trace directory.&lt;/LI&gt;
&lt;LI&gt;Reviewed recent deployments and logging changes.&lt;/LI&gt;
&lt;LI&gt;Checked log-rotation status.&lt;/LI&gt;
&lt;LI&gt;Confirmed the files matched the approved cleanup policy.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Excluded customer data, database files, audit logs, and system files.&lt;/STRONG&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;STRONG&gt;Agent finding:&lt;/STRONG&gt;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;The VM is healthy at the Azure platform layer, but the OS volume has less than 1% free space. The approved application trace directory grew by 86 GB in six hours. The payment service stopped when it could no longer write to disk.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;Agent action table&lt;/H4&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;#&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Action&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Type&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Detail&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;1&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Investigate VM&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;VM, agent heartbeat, Azure platform health&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;2&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Analyze filesystem&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Root volume and free-space trend&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;3&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Correlate app failure&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Service stopped after disk exhaustion&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;4&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Identify disk consumer&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Abnormal growth in an approved trace directory&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;5&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Check changes&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Deployments, logging changes, rotation, scheduled tasks&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;6&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validate cleanup scope&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Candidate files meet approved path, type, and age rules&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;7&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Protect data&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Excludes databases, customer data, security logs, unknown files&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;8&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;P1 against the payment VM/application CI&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;9&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update INC evidence&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Disk usage, growth timeline, service impact, cleanup scope&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;10&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Request approval&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Approval for the restricted recovery runbook&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;11&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Archive logs&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Guest OS)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Archives eligible files to protected Azure Storage&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;12&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Remove eligible files&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Guest OS)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Removes &lt;STRONG&gt;only&lt;/STRONG&gt; successfully archived, allowlisted files&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;13&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Run log rotation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Guest OS)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Executes the approved rotation operation&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;14&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Restart service&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Guest OS)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Restarts &lt;STRONG&gt;only&lt;/STRONG&gt; the named payment service if necessary&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;15&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validate recovery&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Disk, application, archive, availability, growth stabilization&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;16&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update/resolve INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Bytes processed, exclusions, approval, results&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;17&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW PRB&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Permanent logging and rotation remediation&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;H4&gt;Recovery note&lt;/H4&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;The approved recovery runbook archived and removed 82 GB of eligible application trace files. OS-volume free space is now 31%. The payment service was restarted and has passed health checks for 15 minutes. No database, customer, security, or system files were modified.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;That last sentence is not decoration. It is the sentence your auditor will read.&lt;/P&gt;
&lt;H4&gt;Try it yourself&lt;/H4&gt;
&lt;P&gt;&lt;STRONG&gt;Break it&lt;/STRONG&gt; &lt;EM&gt;(lab VM only — this fills the root disk)&lt;/EM&gt;:&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;az vm run-command invoke -g $RG -n vm-payments-linux \
  --command-id RunShellScript --scripts "
    mkdir -p /var/log/payments/trace
    fallocate -l 24G /var/log/payments/trace/trace-$(date +%s).log
    df -h /
  "
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Ask it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="text"&gt;vm-payments-linux root filesystem is nearly full and the payment service is down.
Read-only first. Tell me:
  - exactly which directory grew, by how much, over what window
  - whether log rotation is configured and when it last ran
  - what changed in the last 24 hours that would explain it
Then tell me which files are inside the approved cleanup allowlist
(/var/log/payments/trace/*.log, older than 2h) and which are NOT,
and confirm no database, audit, or customer data files are in scope.
Do not delete anything.
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;The critical design point.&lt;/STRONG&gt; SRE Agent &lt;STRONG&gt;blocks &lt;CODE&gt;delete&lt;/CODE&gt; and &lt;CODE&gt;remove&lt;/CODE&gt; commands outright&lt;/STRONG&gt;. You cannot have it &lt;CODE&gt;rm&lt;/CODE&gt; those files through its Azure CLI surface, and you should be glad. The correct implementation is a &lt;STRONG&gt;fixed-purpose, version-controlled runbook&lt;/STRONG&gt; that the agent &lt;EM&gt;invokes&lt;/EM&gt; with tightly bounded parameters:&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;# What the agent is allowed to call — one runbook, allowlisted parameters, nothing else
az automation runbook start \
  -g $RG --automation-account-name aa-sre-runbooks \
  -n "Reclaim-TraceDiskSpace" \
  --parameters vmName=vm-payments-linux \
               allowedPath=/var/log/payments/trace \
               pattern='*.log' \
               minAgeHours=2 \
               maxBytes=90000000000 \
               archiveToContainer=payments-trace-archive \
               requireArchiveBeforeDelete=true \
               dryRun=false
&lt;/LI-CODE&gt;
&lt;P&gt;The runbook — not the agent — owns the destructive logic, and it enforces:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;path allowlist (refuse anything outside &lt;CODE&gt;allowedPath&lt;/CODE&gt;)&lt;/LI&gt;
&lt;LI&gt;filename pattern allowlist&lt;/LI&gt;
&lt;LI&gt;minimum file age&lt;/LI&gt;
&lt;LI&gt;maximum total bytes per execution&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;successful archive verified before any deletion&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;hard-refuse if any candidate file is unknown, or matches a protected pattern (&lt;CODE&gt;*.mdf&lt;/CODE&gt;, &lt;CODE&gt;*.bak&lt;/CODE&gt;, &lt;CODE&gt;/var/log/audit/*&lt;/CODE&gt;, &lt;CODE&gt;*.key&lt;/CODE&gt;, &lt;CODE&gt;*.pem&lt;/CODE&gt;)&lt;/LI&gt;
&lt;LI&gt;a cooldown that prevents re-execution within N hours&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Verify it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;az vm run-command invoke -g $RG -n vm-payments-linux \
  --command-id RunShellScript \
  --scripts "df -h /; systemctl is-active payments.service; ls -la /var/log/payments/trace | head"
&lt;/LI-CODE&gt;
&lt;H4&gt;Guardrails&lt;/H4&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Do not give SRE Agent unrestricted SSH or shell access.&lt;/STRONG&gt; Ever. This is the single highest-leverage rule in this post.&lt;/LI&gt;
&lt;LI&gt;Use a version-controlled, fixed-purpose runbook.&lt;/LI&gt;
&lt;LI&gt;Allowlist paths, patterns, file ages, and maximum cleanup size.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Stop if the responsible files are unknown.&lt;/STRONG&gt; A disk filled by something you can't identify is a security event until proven otherwise.&lt;/LI&gt;
&lt;LI&gt;Require successful archival before deletion.&lt;/LI&gt;
&lt;LI&gt;Prevent repeated execution with a cooldown.&lt;/LI&gt;
&lt;LI&gt;Treat cleanup as &lt;STRONG&gt;temporary mitigation&lt;/STRONG&gt; — the problem record is the fix.&lt;/LI&gt;
&lt;/UL&gt;
&lt;HR /&gt;
&lt;H3 id="mcetoc_blog_28"&gt;Use case #6 — Linux VM: anomalous CPU saturation&lt;/H3&gt;
&lt;P&gt;&lt;EM&gt;Where "anomalous" is doing all the work.&lt;/EM&gt;&lt;/P&gt;
&lt;H4&gt;What happened&lt;/H4&gt;
&lt;P&gt;Release 7.3.1 introduced an immediate retry loop when an inventory dependency failed. The application retried continuously without backoff, consuming nearly all VM CPU and causing request timeouts.&lt;/P&gt;
&lt;H4&gt;How Azure Monitor detected it&lt;/H4&gt;
&lt;P&gt;The alert deliberately combined conditions rather than firing on a threshold:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;CPU significantly outside the &lt;STRONG&gt;historical baseline&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;Persistent saturation over multiple evaluations&lt;/LI&gt;
&lt;LI&gt;Increased request latency&lt;/LI&gt;
&lt;LI&gt;Increased dependency failures&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;No approved maintenance or batch workload active&lt;/STRONG&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;P1 – Production order-processing Linux VM experiencing anomalous CPU saturation&lt;/STRONG&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;A static "CPU &amp;gt; 90%" rule on a batch-processing VM is a pager that everyone learns to ignore. The composite condition is what makes the alert worth waking someone — or an agent — for.&lt;/P&gt;
&lt;H4&gt;How the agent found the cause&lt;/H4&gt;
&lt;OL&gt;
&lt;LI&gt;Confirmed the VM was online.&lt;/LI&gt;
&lt;LI&gt;Compared current CPU with historical behavior.&lt;/LI&gt;
&lt;LI&gt;Identified the order-processing service as the main CPU consumer.&lt;/LI&gt;
&lt;LI&gt;Reviewed application request latency.&lt;/LI&gt;
&lt;LI&gt;Reviewed downstream dependency failures.&lt;/LI&gt;
&lt;LI&gt;Analyzed application retry logs.&lt;/LI&gt;
&lt;LI&gt;Correlated CPU growth with release 7.3.1.&lt;/LI&gt;
&lt;LI&gt;Excluded expected batch jobs.&lt;/LI&gt;
&lt;LI&gt;Excluded Azure maintenance or platform issues.&lt;/LI&gt;
&lt;LI&gt;Confirmed other application instances had sufficient capacity.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;STRONG&gt;Agent finding:&lt;/STRONG&gt;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;CPU increased from a normal range of 35–50% to 98% four minutes after deployment 7.3.1. The order-processing service is repeatedly calling a failed dependency without backoff. No expected batch job or platform maintenance is active.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;Agent action table&lt;/H4&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;#&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Action&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Type&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Detail&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;1&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Confirm anomaly&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Compares current CPU with the historical baseline&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;2&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Analyze impact&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Latency, timeouts, availability, dependencies&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;3&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Identify process&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Named order service is the primary CPU consumer&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;4&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Correlate deployment&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Release 7.3.1, four minutes before saturation&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;5&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Analyze logs&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Dependency retry loop without backoff&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;6&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Rule out expected work&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Excludes batch, backup, maintenance, scheduled processing&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;7&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Assess blast radius&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Traffic and available capacity on other instances&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;8&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;P1 against the Order Processing CI&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;9&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update INC evidence&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;CPU, process, dependency, release, customer impact&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;10&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Classify response&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Rollback, &lt;STRONG&gt;not&lt;/STRONG&gt; VM resize or arbitrary process termination&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;11&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Request approval&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Approval to drain, roll back, restart the service&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;12&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Drain VM&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Azure)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Removes the VM from load-balancer rotation&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;13&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Restore prior release&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Guest OS/DevOps)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Restores known-good version or configuration&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;14&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Restart named service&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Guest OS)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Restarts &lt;STRONG&gt;only&lt;/STRONG&gt; &lt;CODE&gt;orders-service&lt;/CODE&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;15&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validate recovery&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;CPU, retries, latency, dependencies, health recover&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;16&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Return to rotation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Azure)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Restores traffic &lt;STRONG&gt;only after&lt;/STRONG&gt; successful validation&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;17&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update/resolve INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Drain, rollback, restart, approval, results&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;18&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create defect&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Retry, backoff, and circuit-breaker remediation&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;Steps 12 and 16 are the pattern to steal: &lt;STRONG&gt;drain before you touch, restore traffic only after validation passes.&lt;/STRONG&gt; Most homegrown automation restarts a service while it's still taking traffic and turns a degradation into an outage.&lt;/P&gt;
&lt;H4&gt;Recovery note&lt;/H4&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Release 7.3.1 introduced a retry loop when the inventory dependency failed. After approval, the VM was drained, the previous release was restored, and orders-service was restarted. CPU declined from 98% to 43%, and application latency returned to baseline.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;Try it yourself&lt;/H4&gt;
&lt;P&gt;&lt;STRONG&gt;Break it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;az vm run-command invoke -g $RG -n vm-payments-linux \
  --command-id RunShellScript --scripts "
    nohup bash -c 'while true; do curl -s -m 1 http://127.0.0.1:9/inventory &amp;gt;/dev/null 2&amp;gt;&amp;amp;1; done' &amp;amp;
    nohup bash -c 'while true; do :; done' &amp;amp;
    echo started
  "
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Ask it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="text"&gt;CPU on vm-payments-linux is at 98%. Read-only.
First: is this actually anomalous, or is it consistent with this VM's
historical pattern for this hour and day of week? Show me the baseline.
If anomalous: which process, which dependency is it calling, at what rate,
and what deployed immediately before?
Do NOT propose killing the top process or resizing the VM.
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Expect it to propose&lt;/STRONG&gt; a drain → restore → restart sequence, with the drain step as a separate approval from the restart.&lt;/P&gt;
&lt;H4&gt;Guardrails&lt;/H4&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Do not automatically terminate the highest-CPU process.&lt;/STRONG&gt; It is very often a legitimate workload, and occasionally it's a security incident you just destroyed the evidence for.&lt;/LI&gt;
&lt;LI&gt;Allow operations only for &lt;STRONG&gt;named services&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI&gt;Drain before service restart where possible.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Escalate unknown or suspicious processes to security&lt;/STRONG&gt; rather than remediating them.&lt;/LI&gt;
&lt;LI&gt;Limit restart attempts.&lt;/LI&gt;
&lt;LI&gt;Do not permanently resize the VM when the evidence points to faulty code. Resizing to survive a retry storm is buying hardware to host a bug.&lt;/LI&gt;
&lt;/UL&gt;
&lt;HR /&gt;
&lt;H3 id="mcetoc_blog_29"&gt;Use case #7 — Windows VM with IIS: memory leak&lt;/H3&gt;
&lt;H4&gt;What happened&lt;/H4&gt;
&lt;P&gt;Release ClaimsPortal 5.9.0 introduced a memory leak in the Claims IIS application pool. Memory consumption climbed over several hours, paging began, request queues grew, and IIS returned HTTP 503.&lt;/P&gt;
&lt;H4&gt;How Azure Monitor detected it&lt;/H4&gt;
&lt;P&gt;Azure Monitor Agent collected available memory, committed bytes, paging activity, process working set, IIS request queues, HTTP 500/503 responses, and availability-test results. The alert required &lt;STRONG&gt;sustained abnormal growth&lt;/STRONG&gt;, not a brief spike.&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;P1 – Memory exhaustion affecting production IIS application&lt;/STRONG&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;How the agent found the cause&lt;/H4&gt;
&lt;OL&gt;
&lt;LI&gt;Confirmed the VM and monitoring agent were healthy.&lt;/LI&gt;
&lt;LI&gt;Compared memory behavior with the historical baseline.&lt;/LI&gt;
&lt;LI&gt;Identified the relevant &lt;CODE&gt;w3wp.exe&lt;/CODE&gt; process.&lt;/LI&gt;
&lt;LI&gt;Mapped it to the Claims application pool.&lt;/LI&gt;
&lt;LI&gt;Reviewed paging and request queues.&lt;/LI&gt;
&lt;LI&gt;Correlated HTTP 503 errors with low available memory.&lt;/LI&gt;
&lt;LI&gt;Reviewed Windows Event Logs.&lt;/LI&gt;
&lt;LI&gt;Correlated the growth with release 5.9.0.&lt;/LI&gt;
&lt;LI&gt;Excluded antivirus and scheduled-reporting activity.&lt;/LI&gt;
&lt;LI&gt;Confirmed another instance could carry traffic.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;STRONG&gt;Agent finding:&lt;/STRONG&gt;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Available memory declined from 42% to 4% over three hours. The Claims application pool grew from 1.8 GB to 11.6 GB without releasing memory after traffic normalized. Paging and HTTP 503 errors followed. The pattern began after release 5.9.0.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;"Without releasing memory after traffic normalized" is the sentence that distinguishes a leak from load. Cache growth under load is normal; failure to return afterwards is not.&lt;/P&gt;
&lt;H4&gt;Agent action table&lt;/H4&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;#&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Action&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Type&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Detail&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;1&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Confirm anomaly&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Memory, committed bytes, paging vs. baseline&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;2&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Identify process&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Maps growing &lt;CODE&gt;w3wp.exe&lt;/CODE&gt; to the Claims pool&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;3&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Analyze IIS health&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Pools, queues, HTTP errors, availability&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;4&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Correlate release&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Links sustained growth to release 5.9.0&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;5&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Rule out other causes&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Excludes scheduled jobs, antivirus, maintenance&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;6&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Assess capacity&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Confirms another instance can serve traffic&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;7&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;P1 against the Claims Portal CI&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;8&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update INC evidence&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Memory, paging, pool, release, HTTP impact&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;9&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Classify response&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Targeted recycling&lt;/STRONG&gt;, not a full VM restart&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;10&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Request approval&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Approval to drain and recycle the named pool&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;11&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Drain VM&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Azure)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Removes VM from load-balancer rotation&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;12&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Capture diagnostics&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Guest OS)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Captures approved diagnostics to a protected location&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;13&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Recycle app pool&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Guest OS)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Recycles &lt;STRONG&gt;only&lt;/STRONG&gt; the Claims application pool&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;14&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validate recovery&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Memory, paging, queues, HTTP errors, availability recover&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;15&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Return to rotation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Azure)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Restores traffic after health checks pass&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;16&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update/resolve INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Memory before/after, approval, stability&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;17&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW PRB&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Memory-leak remediation&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;Step 12 before step 13 matters: recycling the pool destroys the evidence. Capture first.&lt;/P&gt;
&lt;H4&gt;Recovery note&lt;/H4&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Abnormal memory growth was isolated to the Claims application pool following release 5.9.0. The VM was drained, the approved application pool was recycled, and health checks passed before traffic was restored. Available memory increased from 4% to 61%, and HTTP 503 responses stopped.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;Try it yourself&lt;/H4&gt;
&lt;P&gt;&lt;STRONG&gt;Break it&lt;/STRONG&gt; &lt;EM&gt;(lab VM only)&lt;/EM&gt;:&lt;/P&gt;
&lt;LI-CODE lang="powershell"&gt;az vm run-command invoke -g $RG -n vm-claims-win `
  --command-id RunPowerShellScript --scripts "
    Install-WindowsFeature Web-Server -IncludeManagementTools
    New-WebAppPool -Name 'ClaimsPool'
    # Simulate the leak
    \$leak = New-Object System.Collections.ArrayList
    1..40 | ForEach-Object { [void]\$leak.Add((New-Object byte[] 100MB)) ; Start-Sleep -Milliseconds 200 }
  "
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Ask it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="text"&gt;vm-claims-win available memory is at 4%. Read-only.
Map the growing process to an IIS application pool.
Show me the memory curve for the last 6 hours and tell me whether memory
was released after traffic dropped. Correlate with deployment history.
Then propose the most targeted possible mitigation — I do not want a VM restart.
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Expect it to propose:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="powershell"&gt;Restart-WebAppPool -Name "ClaimsPool"
&lt;/LI-CODE&gt;
&lt;P&gt;…and nothing else on that machine.&lt;/P&gt;
&lt;H4&gt;Guardrails&lt;/H4&gt;
&lt;UL&gt;
&lt;LI&gt;Recycle &lt;STRONG&gt;only the named application pool&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI&gt;Avoid full VM restart as the first response — it's a bigger hammer with a longer outage and it destroys the leak evidence.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Do not copy memory dumps into ServiceNow.&lt;/STRONG&gt; They contain credentials, tokens, and customer data.&lt;/LI&gt;
&lt;LI&gt;Store diagnostics in a secured location; put the &lt;EM&gt;link&lt;/EM&gt; in the ticket.&lt;/LI&gt;
&lt;LI&gt;Prevent repeated automatic recycling — a pool that needs recycling every 40 minutes is an incident, not a routine.&lt;/LI&gt;
&lt;LI&gt;Escalate if the leak returns during the observation period.&lt;/LI&gt;
&lt;/UL&gt;
&lt;HR /&gt;
&lt;H3 id="mcetoc_blog_30"&gt;Use case #8 — Virtual Machine Scale Set: unhealthy instance&lt;/H3&gt;
&lt;P&gt;&lt;EM&gt;The most autonomy-ready use case in the list, and the reason is stateless workloads.&lt;/EM&gt;&lt;/P&gt;
&lt;H4&gt;What happened&lt;/H4&gt;
&lt;P&gt;A configuration extension failed while VMSS instance 17 was being provisioned. The VM was running but its application service never started. The instance failed application health probes and caused intermittent errors.&lt;/P&gt;
&lt;H4&gt;How Azure Monitor detected it&lt;/H4&gt;
&lt;UL&gt;
&lt;LI&gt;Reduced healthy backend count&lt;/LI&gt;
&lt;LI&gt;Application Health extension failure&lt;/LI&gt;
&lt;LI&gt;Backend health-probe failure&lt;/LI&gt;
&lt;LI&gt;VM extension provisioning failure&lt;/LI&gt;
&lt;LI&gt;Instance-specific errors&lt;/LI&gt;
&lt;LI&gt;Elevated &lt;STRONG&gt;intermittent&lt;/STRONG&gt; HTTP 5xx&lt;/LI&gt;
&lt;/UL&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;P2 – Unhealthy VM Scale Set instance causing intermittent API failures&lt;/STRONG&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;How the agent found the cause&lt;/H4&gt;
&lt;OL&gt;
&lt;LI&gt;Reviewed the VMSS healthy-instance count.&lt;/LI&gt;
&lt;LI&gt;Identified instance 17 as the only unhealthy instance.&lt;/LI&gt;
&lt;LI&gt;Reviewed backend health.&lt;/LI&gt;
&lt;LI&gt;Compared instance 17 with healthy instances.&lt;/LI&gt;
&lt;LI&gt;Checked the image and VMSS model.&lt;/LI&gt;
&lt;LI&gt;Reviewed VM extension state.&lt;/LI&gt;
&lt;LI&gt;Found the configuration extension failure.&lt;/LI&gt;
&lt;LI&gt;Reviewed boot and application diagnostics.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Confirmed the workload was stateless.&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;Confirmed eleven instances could carry production traffic.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;STRONG&gt;Agent finding:&lt;/STRONG&gt;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Instance 17 is the only unhealthy member of the 12-instance scale set. Its application health probe has failed since 09:18 UTC. The configuration extension failed during provisioning, and the application service never started. Eleven healthy instances can maintain service.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;Agent action table&lt;/H4&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;#&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Action&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Type&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Detail&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;1&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Investigate VMSS&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Instance health, provisioning state, healthy count&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;2&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Identify instance&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Instance 17 is the only unhealthy member&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;3&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Analyze backend health&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Confirms the instance fails application probes&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;4&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Compare instances&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Image, model, extensions, configuration&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;5&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Find extension failure&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Locates the failed configuration extension&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;6&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Review diagnostics&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Boot, extension, and application diagnostics&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;7&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Assess safe capacity&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Eleven instances can carry traffic&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;8&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Confirm statelessness&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Replacement won't destroy required local state&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;9&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;P2 against the VMSS/application CI&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;10&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update INC evidence&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Instance, extension error, health, capacity evidence&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;11&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Classify response&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Replacement or reimage per approved procedure&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;12&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Request approval&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Approval to isolate and replace instance 17&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;13&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Isolate instance&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Azure)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Ensures the instance receives no production traffic&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;14&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Preserve evidence&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read/Write&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Stores approved diagnostic evidence securely&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;15&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Replace instance&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Azure)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Reimages or replaces &lt;STRONG&gt;only&lt;/STRONG&gt; instance 17&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;16&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validate provisioning&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Image, model, and extensions deploy successfully&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;17&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validate service&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Backend health, capacity, customer errors recover&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;18&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update/resolve INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Replacement, approval, diagnostics, recovery&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;19&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW PRB&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Extension and image-validation improvement work&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;Steps 8 and 7 are the gate. Replacing a stateless instance with eleven healthy peers is genuinely low risk. Replacing a &lt;EM&gt;stateful&lt;/EM&gt; instance, or replacing one when you're already at minimum capacity, is an outage. Both must be &lt;EM&gt;confirmed&lt;/EM&gt;, not assumed.&lt;/P&gt;
&lt;H4&gt;Recovery note&lt;/H4&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;VMSS instance 17 was isolated after its configuration extension failed and the application service did not start. Diagnostic evidence was captured, and the instance was replaced after approval. The replacement passed extension, application, and backend health checks. The scale set has returned to 12 healthy instances.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;Try it yourself&lt;/H4&gt;
&lt;P&gt;&lt;STRONG&gt;Break it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;INSTANCE_ID=$(az vmss list-instances -g $RG -n vmss-api-demo \
  --query "[5].instanceId" -o tsv)

az vmss extension set -g $RG --vmss-name vmss-api-demo \
  --name CustomScript --publisher Microsoft.Azure.Extensions \
  --settings '{"commandToExecute":"exit 1"}'
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Ask it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="text"&gt;vmss-api-demo has an unhealthy instance. Read-only.
Identify which instance, since when, and the specific extension error.
Confirm for me: (a) the workload is stateless, (b) how many healthy instances
remain, and (c) whether remaining capacity can carry current traffic with
20% headroom. Only if all three are satisfied, propose a reimage of that
single instance. Capture diagnostics before proposing anything.
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Expect it to propose:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;az vmss reimage -g $RG -n vmss-api-demo --instance-id $INSTANCE_ID
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Verify it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;az vmss get-instance-view -g $RG -n vmss-api-demo --instance-id $INSTANCE_ID \
  --query "vmHealth.status.code"
az vmss list-instances -g $RG -n vmss-api-demo -o table
&lt;/LI-CODE&gt;
&lt;H4&gt;Guardrails&lt;/H4&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Confirm the workload is stateless&lt;/STRONG&gt; before any replacement.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Confirm sufficient healthy capacity first.&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;Preserve diagnostic evidence before replacement — the instance is your only copy of the failure.&lt;/LI&gt;
&lt;LI&gt;Restrict permissions to the named VMSS.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Limit simultaneous replacements.&lt;/STRONG&gt; One at a time. An agent that reimages six instances because six probes failed has just caused the outage it was investigating.&lt;/LI&gt;
&lt;LI&gt;Do not permit deletion of the entire scale set.&lt;/LI&gt;
&lt;/UL&gt;
&lt;HR /&gt;
&lt;H3 id="mcetoc_blog_31"&gt;Use case #9 — Application Gateway: HTTP 502 from unhealthy backends&lt;/H3&gt;
&lt;P&gt;&lt;EM&gt;The cross-component change nobody coordinated.&lt;/EM&gt;&lt;/P&gt;
&lt;H4&gt;What happened&lt;/H4&gt;
&lt;P&gt;Customer Portal release 9.2 changed the backend listener from port 443 to 8443. Application Gateway remained configured to connect on 443. All backend probes failed and customers received HTTP 502.&lt;/P&gt;
&lt;H4&gt;How Azure Monitor detected it&lt;/H4&gt;
&lt;UL&gt;
&lt;LI&gt;Increased Application Gateway HTTP 502 responses&lt;/LI&gt;
&lt;LI&gt;Increased failed requests&lt;/LI&gt;
&lt;LI&gt;Four unhealthy backends&lt;/LI&gt;
&lt;LI&gt;Reduced healthy-host count&lt;/LI&gt;
&lt;LI&gt;Failed synthetic availability tests&lt;/LI&gt;
&lt;/UL&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;P1 – Application Gateway returning HTTP 502 due to unhealthy backends&lt;/STRONG&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;How the agent found the cause&lt;/H4&gt;
&lt;OL&gt;
&lt;LI&gt;Reviewed Application Gateway metrics.&lt;/LI&gt;
&lt;LI&gt;Retrieved backend health.&lt;/LI&gt;
&lt;LI&gt;Identified the affected pool.&lt;/LI&gt;
&lt;LI&gt;Reviewed probe path, protocol, host header, and port.&lt;/LI&gt;
&lt;LI&gt;Reviewed backend settings.&lt;/LI&gt;
&lt;LI&gt;Confirmed the application responded on 8443.&lt;/LI&gt;
&lt;LI&gt;Confirmed the gateway used 443.&lt;/LI&gt;
&lt;LI&gt;Correlated the mismatch with release 9.2.&lt;/LI&gt;
&lt;LI&gt;Reviewed NSG, route, DNS, certificate, and WAF changes.&lt;/LI&gt;
&lt;LI&gt;Excluded networking, certificate, and platform-health issues.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;STRONG&gt;Agent finding:&lt;/STRONG&gt;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;HTTP 502 responses began at 16:07 UTC. All four Customer Portal backends are unhealthy. Release 9.2 changed the backend listener to port 8443, while Application Gateway continues to use port 443. No NSG, routing, or certificate issue correlates with the incident.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;Step 9 is what separates a real investigation from a lucky guess. A 502 has at least six plausible causes — NSG, UDR, DNS, expired cert, WAF rule, backend down. The agent has to &lt;EM&gt;exclude&lt;/EM&gt; them, in writing, before you trust the conclusion.&lt;/P&gt;
&lt;H4&gt;Agent action table&lt;/H4&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;#&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Action&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Type&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Detail&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;1&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Investigate gateway&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;HTTP 502, failed requests, latency, backend counts&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;2&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Inspect backend health&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Identifies four unhealthy Customer Portal backends&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;3&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Review configuration&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Settings, probes, protocol, port, TLS, routing&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;4&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Test backend state&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;App responds on 8443 while gateway uses 443&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;5&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Correlate changes&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Links mismatch to Customer Portal release 9.2&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;6&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Rule out networking&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;NSGs, routes, DNS, TLS, WAF, platform health&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;7&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Assess blast radius&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Confirms the Customer Portal pool is unavailable&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;8&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;P1 against the portal/gateway CI&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;9&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update INC evidence&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;502s, backend, port, deployment, known-good configuration&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;10&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Classify response&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Restore the known-good backend listener&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;11&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Request approval&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Approval to restore source-controlled configuration&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;12&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Restore listener&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Application)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Restores the application listener to approved port 443&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;13&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validate backend health&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;All four backends become healthy&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;14&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validate application&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;HTTP 502 declines; synthetic transactions pass&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;15&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Verify controls&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Confirms no WAF, TLS, routing, or NSG control was weakened&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;16&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update/resolve INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Cause, restoration, approval, recovery&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;17&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW CHG&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Coordinated change for the intended port migration&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;18&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW PRB&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Cross-component deployment-validation work&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;Step 12 is a genuinely important choice. There were two ways to fix this: change the &lt;EM&gt;application&lt;/EM&gt; back to 443, or change the &lt;EM&gt;gateway&lt;/EM&gt; to 8443. The agent restores the &lt;STRONG&gt;application to the known-good, source-controlled state&lt;/STRONG&gt; rather than mutating the gateway to match an unapproved change. One of those is a rollback; the other is ratifying an unreviewed change during an outage. Then step 17 files a proper change record for the migration the team clearly &lt;EM&gt;intended&lt;/EM&gt; to do.&lt;/P&gt;
&lt;P&gt;Step 15 exists because the fastest way to make a 502 disappear is to disable TLS validation. The agent must prove it didn't take the fast way.&lt;/P&gt;
&lt;H4&gt;Recovery note&lt;/H4&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Customer Portal release 9.2 changed the backend listener from port 443 to 8443 without a coordinated gateway change. The application listener was restored to the previous configuration. All four backends are healthy, HTTP 502 responses returned to baseline, and synthetic login tests passed for 15 minutes.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;Try it yourself&lt;/H4&gt;
&lt;P&gt;&lt;STRONG&gt;Ask it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="text"&gt;appgw-portal is returning 502s and all backends are unhealthy. Read-only.
Walk me through the exclusion, explicitly, for each of:
NSG, UDR/route table, DNS resolution, backend TLS certificate,
WAF rule blocking, backend process down, and probe configuration mismatch.
State which you ruled out and the evidence for each.
Then tell me the known-good configuration and where it is source-controlled.
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Verify it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;az network application-gateway show-backend-health \
  -g $RG -n appgw-portal \
  --query "backendAddressPools[].backendHttpSettingsCollection[].servers[].{addr:address,health:health}" -o table
&lt;/LI-CODE&gt;
&lt;H4&gt;Guardrails&lt;/H4&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Do not disable TLS validation.&lt;/STRONG&gt; Not to test, not temporarily, not "just to confirm."&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Do not weaken NSGs or WAF policies&lt;/STRONG&gt; as a mitigation.&lt;/LI&gt;
&lt;LI&gt;Use source-controlled configuration as the definition of "known-good."&lt;/LI&gt;
&lt;LI&gt;Restore a known-good state during the incident; migrate through a change record afterwards.&lt;/LI&gt;
&lt;LI&gt;Require approval for gateway or backend changes.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Validate full application transactions, not just health probes.&lt;/STRONG&gt; A probe returning 200 on &lt;CODE&gt;/health&lt;/CODE&gt; proves very little.&lt;/LI&gt;
&lt;/UL&gt;
&lt;HR /&gt;
&lt;H3 id="mcetoc_blog_32"&gt;Use case #10 — Azure Service Bus: queue and dead-letter backlog&lt;/H3&gt;
&lt;P&gt;&lt;EM&gt;The one that teaches the most important lesson in the entire post.&lt;/EM&gt;&lt;/P&gt;
&lt;H4&gt;What happened&lt;/H4&gt;
&lt;P&gt;Release &lt;CODE&gt;fulfillment-worker:6.4.0&lt;/CODE&gt; couldn't deserialize messages containing a new &lt;CODE&gt;deliveryWindow&lt;/CODE&gt; field. Consumer throughput dropped by 92%. The active backlog grew rapidly, and incompatible messages entered the dead-letter queue.&lt;/P&gt;
&lt;H4&gt;How Azure Monitor detected it&lt;/H4&gt;
&lt;UL&gt;
&lt;LI&gt;Increasing active-message count&lt;/LI&gt;
&lt;LI&gt;Increasing oldest-message age&lt;/LI&gt;
&lt;LI&gt;Dead-letter growth&lt;/LI&gt;
&lt;LI&gt;Reduced completed-message rate&lt;/LI&gt;
&lt;LI&gt;Consumer application errors&lt;/LI&gt;
&lt;LI&gt;Delayed downstream business processing&lt;/LI&gt;
&lt;/UL&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;P1 – Production order-event backlog delaying fulfillment&lt;/STRONG&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;How the agent found the cause&lt;/H4&gt;
&lt;OL&gt;
&lt;LI&gt;Reviewed active, incoming, outgoing, and dead-letter counts.&lt;/LI&gt;
&lt;LI&gt;Calculated message arrival and completion rates.&lt;/LI&gt;
&lt;LI&gt;Confirmed the backlog was growing.&lt;/LI&gt;
&lt;LI&gt;Reviewed consumer instance health.&lt;/LI&gt;
&lt;LI&gt;Reviewed consumer errors and restarts.&lt;/LI&gt;
&lt;LI&gt;Checked downstream dependency health.&lt;/LI&gt;
&lt;LI&gt;Checked Service Bus authentication and authorization.&lt;/LI&gt;
&lt;LI&gt;Correlated the throughput decline with release 6.4.0.&lt;/LI&gt;
&lt;LI&gt;Found deserialization errors for the new field.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Determined scaling more broken consumers would amplify failures.&lt;/STRONG&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;STRONG&gt;Agent finding:&lt;/STRONG&gt;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;The order-events queue grew from 3,000 to 185,000 active messages in 35 minutes. Consumer throughput dropped by 92% immediately after release fulfillment-worker:6.4.0. Application logs show deserialization failures involving the new &lt;CODE&gt;deliveryWindow&lt;/CODE&gt; field.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;Step 10 is the whole reason to use a reasoning agent instead of an autoscale rule. Every metric here screams "scale out the consumers." An HPA would have done exactly that, and every new replica would have dead-lettered messages faster.&lt;/P&gt;
&lt;H4&gt;Agent action table&lt;/H4&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;#&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Action&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Type&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Detail&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;1&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Investigate queue&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Active, incoming, outgoing, scheduled, DLQ counts&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;2&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Calculate flow rates&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Confirms arrival exceeds completion; backlog growing&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;3&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Inspect consumers&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Health, instance count, scaling, errors, dependencies&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;4&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Analyze failures&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Deserialization exceptions involving the new field&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;5&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Correlate release&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Links the 92% throughput reduction to release 6.4.0&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;6&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Rule out Service Bus&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Health, authorization, throttling, networking, service status&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;7&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Assess blast radius&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Backlog age and fulfillment impact&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;8&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Protect message data&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Prevents payloads or personal data entering ServiceNow&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;9&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;P1 against the fulfillment integration CI&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;10&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update INC evidence&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Backlog, age, flow, exception, release, business impact&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;11&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Classify response&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Rollback, not scaling broken consumers&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;12&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Request approval&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Approval to restore consumer 6.3.7&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;13&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Roll back consumer&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Azure/K8s)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Rolls back through the approved deployment platform&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;14&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Restore capacity&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (Azure/K8s)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Restores approved consumer instance count&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;15&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validate processing&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Consumer and downstream health&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;16&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validate backlog&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Completion exceeds arrival; DLQ growth stops&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;17&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Update INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Rollback, rates, estimated drain time, approval&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;18&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Maintain incident&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Decision&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Keeps the incident open until backlog age meets the objective&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;19&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW CHG&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Separate controlled change for DLQ replay&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;20&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create SNOW PRB&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Write (ITSM)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Message-contract compatibility remediation&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;H4&gt;Recovery note&lt;/H4&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Consumer release 6.4.0 could not deserialize messages containing the new &lt;CODE&gt;deliveryWindow&lt;/CODE&gt; field. The fulfillment worker was rolled back to 6.3.7. Consumer throughput recovered, new dead-letter growth stopped, and the active backlog is draining at approximately 7,500 messages per minute. Dead-letter replay requires a separately approved procedure.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H4&gt;The lesson: technical recovery is not business recovery&lt;/H4&gt;
&lt;P&gt;Step 18 is the most important row in this entire post.&lt;/P&gt;
&lt;P&gt;At the moment of rollback, every technical signal is green. Consumers are healthy. Throughput has recovered. The DLQ has stopped growing. An agent optimizing for metrics would resolve the incident right there and go back to sleep.&lt;/P&gt;
&lt;P&gt;But there are still 185,000 unshipped orders and a dead-letter queue full of messages that need a &lt;EM&gt;separately approved&lt;/EM&gt; replay procedure. Customers are still affected. &lt;STRONG&gt;The incident stays open until backlog age meets the business objective&lt;/STRONG&gt;, not until the graphs look nice.&lt;/P&gt;
&lt;P&gt;Encode this in the response plan explicitly:&lt;/P&gt;
&lt;LI-CODE lang="text"&gt;Do not resolve this incident when consumer health recovers.
Resolution criteria:
  1. Completion rate exceeds arrival rate for 15 consecutive minutes, AND
  2. Oldest active message age is under 5 minutes, AND
  3. Dead-letter count has not increased for 30 minutes.
DLQ replay is out of scope for this incident. File a separate change record.
&lt;/LI-CODE&gt;
&lt;H4&gt;Try it yourself&lt;/H4&gt;
&lt;P&gt;&lt;STRONG&gt;Break it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;# Flood the queue while no consumer is running
for i in $(seq 1 5000); do
  az servicebus queue message send -g $RG --namespace-name sb-sre-demo \
    -q order-events --body "{\"orderId\":$i,\"deliveryWindow\":\"2026-08-09T10:00Z\"}" 2&amp;gt;/dev/null
done

az servicebus queue show -g $RG --namespace-name sb-sre-demo -n order-events \
  --query "countDetails" -o json
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Ask it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="text"&gt;order-events on sb-sre-demo has a growing backlog. Read-only.
Give me: arrival rate, completion rate, current active count, oldest message age,
DLQ count and DLQ growth rate, and the projected drain time at current rates.
Then tell me why scaling out consumers is or is not the correct action here.
Do not include any message payloads or customer data in your answer.
&lt;/LI-CODE&gt;
&lt;P&gt;That last line is not optional. Message bodies routinely contain names, addresses, and payment references — and everything the agent writes goes into a ticket.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Verify it:&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;az servicebus queue show -g $RG --namespace-name sb-sre-demo -n order-events \
  --query "{active:countDetails.activeMessageCount, dlq:countDetails.deadLetterMessageCount}" -o json
&lt;/LI-CODE&gt;
&lt;H4&gt;Guardrails&lt;/H4&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Never automatically purge queues.&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Do not automatically replay dead-letter messages.&lt;/STRONG&gt; Replay without idempotency guarantees means duplicate charges and duplicate shipments.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Do not place message payloads in ServiceNow.&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;Confirm idempotency before replay.&lt;/LI&gt;
&lt;LI&gt;Scale consumers &lt;STRONG&gt;only when the processing path is healthy&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI&gt;Require separate approval for replay or queue configuration changes.&lt;/LI&gt;
&lt;LI&gt;Keep the incident open until &lt;STRONG&gt;business&lt;/STRONG&gt; recovery is confirmed.&lt;/LI&gt;
&lt;/UL&gt;
&lt;HR /&gt;
&lt;H2 id="mcetoc_blog_33"&gt;7. The ITSM integration model&lt;/H2&gt;
&lt;H3 id="mcetoc_blog_34"&gt;Recommended incident fields&lt;/H3&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Field&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Example&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Short description&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Production Orders API pods failing after deployment&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Configuration item&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;prod-aks-orders-api&lt;/CODE&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Assignment group&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Container Platform Operations&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Impact&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;High&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Urgency&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;High&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Environment&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Production&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Azure resource ID&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Full affected Azure resource ID&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Azure alert ID&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Azure Monitor alert correlation identifier&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;SRE investigation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Link to the SRE Agent investigation thread&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Current impact&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Eight of ten replicas unavailable&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Probable cause&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Missing configuration in latest release&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Confidence&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;High&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Proposed action&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Roll back to revision 41&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Approval&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Approver and UTC timestamp&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Ten replicas ready and synthetic tests passing&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Resolution&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Service restored through rollback&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;The &lt;STRONG&gt;Confidence&lt;/STRONG&gt; field earns its place. An agent that says "probable cause: missing configuration (confidence: low)" is far more useful than one that always sounds certain, because it tells the human how much to verify before approving.&lt;/P&gt;
&lt;P&gt;Fields the agent can set directly (&lt;STRONG&gt;preview&lt;/STRONG&gt;): &lt;CODE&gt;assignment_group&lt;/CODE&gt;, &lt;CODE&gt;category&lt;/CODE&gt;, &lt;CODE&gt;subcategory&lt;/CODE&gt;, &lt;CODE&gt;impact&lt;/CODE&gt;, &lt;CODE&gt;urgency&lt;/CODE&gt;, &lt;CODE&gt;priority&lt;/CODE&gt;, &lt;CODE&gt;short_description&lt;/CODE&gt;, and any custom &lt;CODE&gt;u_*&lt;/CODE&gt; field. It &lt;STRONG&gt;cannot&lt;/STRONG&gt; change incident state through field updates — acknowledge and resolve are separate, dedicated tools.&lt;/P&gt;
&lt;H3 id="mcetoc_blog_35"&gt;Recommended agent-generated timeline&lt;/H3&gt;
&lt;P&gt;Every work note should carry:&lt;/P&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Element&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Example&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Timestamp&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;CODE&gt;2026-03-18 14:26 UTC&lt;/CODE&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Observation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;HTTP 500 rate increased to 18%&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Evidence&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Link to the Application Insights query&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Change correlation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Incident started three minutes after deployment&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Probable cause&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Missing production application setting&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Confidence&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;High&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Proposed action&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Swap to previous healthy slot&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Risk&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Temporary deployment rollback&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Approval&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Incident commander and timestamp&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Execution&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Slot-swap operation and result&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;5xx below 1% for 15 minutes&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Follow-up&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Problem record for configuration validation&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;Note that &lt;STRONG&gt;Evidence is a link, not a paste&lt;/STRONG&gt;. This is a deliberate data-protection pattern: the ticket carries a pointer to the query, and the query results stay in the system that already has the right access controls.&lt;/P&gt;
&lt;H3 id="mcetoc_blog_36"&gt;Deduplication strategy&lt;/H3&gt;
&lt;P&gt;Do &lt;STRONG&gt;not&lt;/STRONG&gt; create one incident per alert, pod, queue, or VMSS instance. Correlate on:&lt;/P&gt;
&lt;LI-CODE lang="text"&gt;application/business service
  + Azure resource ID
  + environment (production)
  + alert-rule family
  + active incident time window
&lt;/LI-CODE&gt;
&lt;P&gt;Related alerts attach to the existing incident as evidence or child alerts. Azure Monitor already merges recurring alerts into a single thread when it's the bound platform; for ServiceNow, this correlation key is yours to implement.&lt;/P&gt;
&lt;P&gt;Get this wrong and your first AKS incident produces eight incidents, eight investigations, and eight rollback proposals for the same deployment.&lt;/P&gt;
&lt;H3 id="mcetoc_blog_37"&gt;Record responsibilities&lt;/H3&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Record&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Purpose&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Example&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Incident (INC)&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Restore service quickly and safely&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;App Service HTTP 500 outage&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Problem (PRB)&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Identify and remove the underlying cause&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Missing deployment configuration validation&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Change (CHG)&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Govern permanent or higher-risk production changes&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Coordinated Application Gateway port migration&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Engineering defect/task&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Correct application or automation behavior&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Add retry backoff to the Linux application&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;The agent must clearly distinguish &lt;STRONG&gt;temporary mitigation&lt;/STRONG&gt; from &lt;STRONG&gt;permanent correction&lt;/STRONG&gt;. Every single use case above ends with a follow-up record, and that's not bureaucratic theatre — an agent that mitigates flawlessly and never files a problem record is an agent that lets the same outage recur forever while making the metrics look great.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="mcetoc_blog_38"&gt;8. Reality check: where you have to build&lt;/H2&gt;
&lt;P&gt;This is the section that will save you a month.&lt;/P&gt;
&lt;P&gt;The PDF this post is built from is explicit that these are &lt;STRONG&gt;target response patterns, not guaranteed zero-configuration behavior&lt;/STRONG&gt;. Having now checked each pattern against the product documentation, here is exactly where the gaps are and what fills them.&lt;/P&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;#&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;The pattern assumes&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;What's actually documented&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;What you must build&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;1&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Agent &lt;STRONG&gt;creates&lt;/STRONG&gt; a ServiceNow INC&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;ServiceNow is an &lt;EM&gt;inbound&lt;/EM&gt; platform. Documented writes: &lt;STRONG&gt;post discussion entries, acknowledge, resolve&lt;/STRONG&gt;, plus field updates (preview)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Incidents should originate in ServiceNow (via its own Azure Monitor integration) and flow &lt;EM&gt;in&lt;/EM&gt;. If you truly need agent-initiated creation, add a &lt;STRONG&gt;Python tool or MCP server&lt;/STRONG&gt; against the ServiceNow Table API&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;2&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Agent creates &lt;STRONG&gt;PRB / CHG / defect&lt;/STRONG&gt; records&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Not a documented first-class ServiceNow action&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Same: a custom tool against &lt;CODE&gt;/api/now/table/problem&lt;/CODE&gt; and &lt;CODE&gt;/change_request&lt;/CODE&gt;. This is ~30 lines of Python and worth doing properly, with a least-privileged integration user&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;3&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;ServiceNow &lt;EM&gt;and&lt;/EM&gt; PagerDuty both connected&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Only one incident platform active at a time&lt;/STRONG&gt;; switching disconnects the other&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Bind the agent to your system of record (ServiceNow). Reach the pager through a connector, Teams/Slack, or a webhook&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;4&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Agent runs guest-OS cleanup (&lt;CODE&gt;rm&lt;/CODE&gt;, rotate, restart)&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;&lt;CODE&gt;delete&lt;/CODE&gt; and &lt;CODE&gt;remove&lt;/CODE&gt; commands are blocked outright.&lt;/STRONG&gt; &lt;CODE&gt;az keyvault&lt;/CODE&gt; blocked. Management locks respected&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Wrap all guest-OS work in &lt;STRONG&gt;fixed-purpose Azure Automation runbooks&lt;/STRONG&gt; or a constrained &lt;CODE&gt;az vm run-command&lt;/CODE&gt; script, invoked with allowlisted parameters. See use case #5&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;5&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Agent "pauses the release pipeline"&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Requires the GitHub or Azure DevOps connector, plus permissions on that pipeline&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Connect source control; grant pipeline permissions explicitly; test the pause path before you need it&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;6&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Approval gate on &lt;EM&gt;every&lt;/EM&gt; action&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Review mode shows Approve/Deny &lt;STRONG&gt;only for Azure infrastructure operations&lt;/STRONG&gt;. Emails, Teams posts, and external queries proceed on the agent's reasoning&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Use &lt;STRONG&gt;hooks&lt;/STRONG&gt; or &lt;STRONG&gt;tool access policies&lt;/STRONG&gt; to gate non-Azure actions&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;7&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Agent resolves incidents&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Supported — but during a pilot you don't want it&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Require human confirmation before resolve. Encode it in the response plan and enforce it with a &lt;CODE&gt;Stop&lt;/CODE&gt; hook&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;8&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;One agent handles everything&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Response plans route to &lt;STRONG&gt;custom agents&lt;/STRONG&gt;; skills cap at &lt;STRONG&gt;five concurrent active&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Build domain specialists (§4.6). A single mega-agent thrashes its skill budget&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;9&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Autonomous mode by default is fine&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Per-plan default &lt;STRONG&gt;is&lt;/STRONG&gt; Autonomous, and connecting a platform auto-creates an autonomous &lt;CODE&gt;quickstart_handler&lt;/CODE&gt;&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Delete the quickstart plan.&lt;/STRONG&gt; Set every plan to Review explicitly&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;10&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Agent has broad subscription rights&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;You can't remove individual permissions — only whole resource groups&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Design resource groups as blast-radius boundaries &lt;STRONG&gt;before&lt;/STRONG&gt; onboarding&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;None of these are blockers. All of them are a week of work you'd rather discover now than during your pilot readout.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="mcetoc_blog_39"&gt;9. Approval and autonomy policy&lt;/H2&gt;
&lt;P&gt;The policy I would actually ship:&lt;/P&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Action category&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Recommended initial policy&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Read metrics, logs, traces, resource health&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Automatic&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Correlate deployments and configuration changes&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Automatic&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create a ServiceNow incident&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Automatic after deduplication&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Add sanitized work notes&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Automatic&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Prepare a remediation plan&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Automatic&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create a draft ServiceNow change&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Automatic, but &lt;STRONG&gt;not&lt;/STRONG&gt; approve it&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Modify an Azure production resource&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Human approval required&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Execute a VM guest runbook&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Human approval required&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Roll back an application&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Human approval required&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Delete or replace a stateless VMSS instance&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Human approval required&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Resolve a ServiceNow incident&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Human confirmation during the pilot&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Delete data, purge queues, replay DLQ messages&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Separate explicit approval&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Autonomous remediation&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Only after a proven, bounded pilot&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;Two notes on making this real:&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Start in Review and stay there longer than feels necessary.&lt;/STRONG&gt; The documented recommendation is to observe for two to four weeks and then promote &lt;EM&gt;specific&lt;/EM&gt; triggers you consistently approve. Not the agent — the triggers. Promotion should be per-response-plan and evidence-based: "we approved this exact rollback proposal eleven times without modification" is a reason to go autonomous. "It seems good" is not.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Autonomy should be earned per action type, not per environment.&lt;/STRONG&gt; "Autonomous in staging" is a fine starting rule, but the durable version is "autonomous for VMSS single-instance reimage where the workload is stateless and healthy capacity exceeds 80%" — a narrow, well-characterized action with a mechanical precondition.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="mcetoc_blog_40"&gt;10. Cross-cutting security controls&lt;/H2&gt;
&lt;OL&gt;
&lt;LI&gt;Use a &lt;STRONG&gt;dedicated managed identity&lt;/STRONG&gt; for SRE Agent.&lt;/LI&gt;
&lt;LI&gt;Scope Azure roles to selected &lt;STRONG&gt;resources or resource groups&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI&gt;Avoid broad &lt;STRONG&gt;Contributor&lt;/STRONG&gt; and &lt;STRONG&gt;Owner&lt;/STRONG&gt; assignments.&lt;/LI&gt;
&lt;LI&gt;Use &lt;STRONG&gt;fixed-purpose Automation runbooks&lt;/STRONG&gt; for guest operations.&lt;/LI&gt;
&lt;LI&gt;Do &lt;STRONG&gt;not&lt;/STRONG&gt; provide unrestricted SSH, shell, or PowerShell execution.&lt;/LI&gt;
&lt;LI&gt;Use a dedicated &lt;STRONG&gt;least-privileged ServiceNow integration identity&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI&gt;Prefer &lt;STRONG&gt;OAuth&lt;/STRONG&gt; or a managed connector over stored credentials.&lt;/LI&gt;
&lt;LI&gt;Store required secrets in &lt;STRONG&gt;Key Vault&lt;/STRONG&gt; — never in prompts. (The agent blocks &lt;CODE&gt;az keyvault&lt;/CODE&gt; commands entirely, which helps.)&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Sanitize logs&lt;/STRONG&gt; before posting them into ServiceNow.&lt;/LI&gt;
&lt;LI&gt;Do not post &lt;STRONG&gt;tokens, personal data, SQL text, message payloads, or memory dumps&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI&gt;Record &lt;STRONG&gt;every approval and production action&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI&gt;Define &lt;STRONG&gt;remediation cooldowns&lt;/STRONG&gt; and maximum retry counts.&lt;/LI&gt;
&lt;LI&gt;Require &lt;STRONG&gt;post-action application validation&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI&gt;Retain existing &lt;STRONG&gt;manual runbooks&lt;/STRONG&gt; as fallback.&lt;/LI&gt;
&lt;LI&gt;Use &lt;STRONG&gt;Azure Cost Management alerts&lt;/STRONG&gt; for temporary scaling actions.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H3 id="mcetoc_blog_41"&gt;What the platform gives you for free&lt;/H3&gt;
&lt;P&gt;Worth knowing so you don't rebuild it:&lt;/P&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Layer&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Isolation model&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Compute&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Dedicated sandbox (micro VM) per agent; tool execution separate from the reasoning loop&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Database&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Separate database per agent&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Blob storage&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Separate blob storage per agent&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Network&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Per-agent proxy instance validating every outbound request&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Credentials&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Identity sidecar issues short-lived, per-call tokens; &lt;STRONG&gt;credentials never enter the reasoning context&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;Token lifetimes: managed identity ~1 hour (auto-refreshed), OAuth refreshed 20 minutes before expiry, per-tool-call action tokens are single-use, blob SAS 1 hour refreshed at 45 minutes. Each tool invocation launches a fresh process whose entire tree terminates on completion — there are no persistent process pools, so one tool call cannot see another's environment.&lt;/P&gt;
&lt;H3 id="mcetoc_blog_42"&gt;The audit trail&lt;/H3&gt;
&lt;P&gt;Every &lt;CODE&gt;az&lt;/CODE&gt; command is logged to &lt;STRONG&gt;your&lt;/STRONG&gt; Application Insights as an &lt;CODE&gt;AgentAzCliExecution&lt;/CODE&gt; custom event. This is your evidence for change management:&lt;/P&gt;
&lt;LI-CODE lang="kusto"&gt;customEvents
| where name == "AgentAzCliExecution"
| where timestamp &amp;gt; ago(30d)
| project timestamp,
          command   = tostring(customDimensions.command),
          resource  = tostring(customDimensions.resourceId),
          succeeded = tostring(customDimensions.success),
          thread    = tostring(customDimensions.threadId)
| order by timestamp desc
&lt;/LI-CODE&gt;
&lt;P&gt;Run that query in front of your auditor once and most of the "but can we prove what it did" conversation ends.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="mcetoc_blog_43"&gt;11. A 30/60/90 pilot that survives contact with your CAB&lt;/H2&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Phase&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Enabled capability&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;1&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Detect Azure Monitor alerts&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;2&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Create or correlate ServiceNow incidents&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;3&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Perform &lt;STRONG&gt;read-only&lt;/STRONG&gt; investigation&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;4&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Add sanitized findings to ServiceNow&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;5&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;&lt;STRONG&gt;Recommend&lt;/STRONG&gt; remediation without execution&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;6&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Execute bounded actions after approval in Review mode&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;7&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Validate technical &lt;STRONG&gt;and business&lt;/STRONG&gt; recovery&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;8&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Prepare incident resolution and follow-up records&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;9&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Consider autonomy only for proven low-risk actions&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;Phases 1–5 have &lt;STRONG&gt;zero production write risk&lt;/STRONG&gt; and deliver most of the MTTR reduction. Do not rush past them to get to the demo-friendly part.&lt;/P&gt;
&lt;H3 id="mcetoc_blog_44"&gt;The best first five candidates&lt;/H3&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;App Service deployment-slot rollback&lt;/STRONG&gt; — clean trigger, reversible action, unambiguous validation&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;AKS deployment rollback&lt;/STRONG&gt; — same shape, one GitOps caveat&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;VMSS unhealthy-instance replacement&lt;/STRONG&gt; — stateless, bounded, easy precondition check&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Restricted VM disk-recovery runbook&lt;/STRONG&gt; — high toil, high value, forces you to build the runbook pattern properly&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Automatic ServiceNow incident creation and timeline updates&lt;/STRONG&gt; — the compounding one; every incident from here on is better documented than any incident before it&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;These five share the properties you want: clear triggers, tightly bounded actions, measurable validation criteria, and practical escalation paths.&lt;/P&gt;
&lt;H3 id="mcetoc_blog_45"&gt;What to measure in week one&lt;/H3&gt;
&lt;P&gt;Before you enable a single write action, capture your baseline:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Median time from alert to &lt;STRONG&gt;first accurate human diagnosis&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;Percentage of incidents where the first hypothesis was wrong&lt;/LI&gt;
&lt;LI&gt;Median time from diagnosis to mitigation&lt;/LI&gt;
&lt;LI&gt;Percentage of incidents with a complete timeline in the ticket&lt;/LI&gt;
&lt;LI&gt;Percentage of incidents that produced a follow-up problem record&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The read-only phase moves the first, second, and fourth of those immediately. If it doesn't, your telemetry is the problem, not the agent — and that's a genuinely useful thing to discover in week one rather than week twelve.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="mcetoc_blog_46"&gt;12. Measuring whether it's actually working&lt;/H2&gt;
&lt;P&gt;Under &lt;STRONG&gt;Monitor → Incident metrics&lt;/STRONG&gt;:&lt;/P&gt;
&lt;TABLE style="border-collapse:collapse;border:0;width:100%;margin:18px 0;font-size:15px;line-height:1.45;"&gt;&lt;TBODY&gt;&lt;TR&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;Metric&lt;/STRONG&gt;&lt;/TD&gt;
&lt;TD style="text-align:left;padding:10px 14px;border:0;border-bottom:2px solid #0F6CBD;background:#F5F9FD;vertical-align:top;"&gt;&lt;STRONG&gt;What it shows&lt;/STRONG&gt;&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Incidents reviewed&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Total incidents the agent processes&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Mitigated by agent&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Resolved autonomously&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Assisted by agent&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Agent helped; a human completed it&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Mitigated by user&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Human resolved using agent-provided information&lt;/TD&gt;
&lt;/TR&gt;&lt;TR&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Pending user action&lt;/TD&gt;
&lt;TD style="padding:10px 14px;border:0;border-bottom:1px solid #E6E6E6;vertical-align:top;"&gt;Waiting on a human&lt;/TD&gt;
&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;
&lt;P&gt;The counter-intuitive read: &lt;STRONG&gt;"Assisted by agent" and "Mitigated by user" are the healthy numbers during a pilot.&lt;/STRONG&gt; A high "Mitigated by agent" count in month one means someone left autonomy on.&lt;/P&gt;
&lt;P&gt;Watch &lt;STRONG&gt;Pending user action&lt;/STRONG&gt; closely. A growing queue there means either your approval routing is broken or the agent is proposing things nobody is comfortable approving — both are important signals, and both are invisible without this dashboard.&lt;/P&gt;
&lt;P&gt;Also check &lt;STRONG&gt;Monitor → Session insights&lt;/STRONG&gt; periodically. Each insight card links back to the thread that generated it, so you can trace any learned pattern to its origin. If the agent has learned something wrong, this is where you find it — and &lt;CODE&gt;#forget&lt;/CODE&gt; is how you fix it.&lt;/P&gt;
&lt;HR /&gt;
&lt;H2 id="mcetoc_blog_47"&gt;13. Resources&lt;/H2&gt;
&lt;P&gt;&lt;STRONG&gt;Core documentation&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/sre-agent/overview" target="_blank"&gt;Overview of Azure SRE Agent&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/sre-agent/security-overview" target="_blank"&gt;Security and trust model&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/sre-agent/permissions" target="_blank"&gt;Agent permissions&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/sre-agent/run-modes" target="_blank"&gt;Run modes&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Incident response&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/sre-agent/incident-platforms" target="_blank"&gt;Incident management platforms&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/sre-agent/incident-response-plans" target="_blank"&gt;Incident response plans&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/sre-agent/servicenow-incidents" target="_blank"&gt;ServiceNow incident indexing&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/sre-agent/root-cause-analysis" target="_blank"&gt;Root cause analysis&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/sre-agent/execute-mitigations" target="_blank"&gt;Execute mitigations&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Extensibility&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/sre-agent/sub-agents" target="_blank"&gt;Custom agents&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/sre-agent/skills" target="_blank"&gt;Skills&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/sre-agent/agent-hooks" target="_blank"&gt;Agent hooks&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/sre-agent/scheduled-tasks" target="_blank"&gt;Scheduled tasks&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/sre-agent/memory" target="_blank"&gt;Memory and knowledge&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;HR /&gt;
&lt;H2 id="mcetoc_blog_48"&gt;Closing&lt;/H2&gt;
&lt;P&gt;The framing that makes this work isn't "AI runs my production." It's this:&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Your agent is the most junior person on the rotation — and the most thorough.&lt;/STRONG&gt; It will never skip the Resource Health check. It will never forget to compare against the previous slot. It will never write "restarted it, seems fine" in a work note at 4 AM. And it will never, ever be allowed to &lt;CODE&gt;rm -rf&lt;/CODE&gt; anything.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;Every guardrail in this post exists to keep it in that role. Scope the identity to resource groups. Keep production writes in Review. Wrap guest-OS work in runbooks with allowlists. Never let it purge a queue or replay a dead-letter message on its own. Keep the incident open until customers are actually served, not until the graphs look nice.&lt;/P&gt;
&lt;P&gt;Do that, and the ten workflows above stop being a slide deck and start being your Tuesday.&lt;/P&gt;
&lt;P&gt;Start with use case #1. One App Service, one slot, one alert rule, one response plan in Review mode. Watch it assemble an evidence chain you'd have spent twenty minutes building by hand, and then decide how much further you want to go.&lt;/P&gt;
&lt;HR /&gt;
&lt;P&gt;&lt;EM&gt;The ten scenarios in this post are target response patterns. Each one requires appropriate telemetry, scoped Azure RBAC, response-plan instructions, approved remediation tooling, and ITSM integration. Confirm current Azure SRE Agent and ServiceNow connector capabilities against Microsoft documentation before implementing — the product is moving quickly, and several capabilities referenced here are in preview.&lt;/EM&gt;&lt;/P&gt;</description>
      <pubDate>Wed, 26 Aug 2026 19:02:49 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-mission-critical-blog/your-on-call-rotation-has-a-new-member-10-production-incidents/ba-p/4545187</guid>
      <dc:creator>zmustafa</dc:creator>
      <dc:date>2026-08-26T19:02:49Z</dc:date>
    </item>
    <item>
      <title>Grok 4.6 comes to Microsoft Foundry Models: Built for long-horizon reasoning and complex workflows</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/grok-4-6-comes-to-microsoft-foundry-models-built-for-long/ba-p/4547578</link>
      <description>&lt;P&gt;The next wave of AI applications won't be defined by how quickly models answer questions. They'll be defined by how much work they can complete. Whether resolving repository issues, navigating codebases, generating engineering designs, or executing multi-step enterprise workflows, developers are increasingly building systems that need models capable of staying on task across complex, long-running objectives.&lt;/P&gt;
&lt;P&gt;Today, Grok 4.6 from SpaceXAI is available in Microsoft Foundry Models through public preview, bringing SpaceXAI's latest frontier model to developers through a unified platform for model discovery, evaluation, deployment, and governance.&lt;/P&gt;
&lt;P&gt;Built on SpaceXAI's 1.5T-scale model family, Grok 4.6 focuses on strong performance across coding, engineering, office productivity, research enablement, and inference optimization tasks. Rather than optimizing for a single benchmark category, Grok 4.6 is designed for sustained reasoning and execution across software engineering, agentic workflows, and technical problem solving.&lt;/P&gt;
&lt;H5&gt;Long-Running Agents That Stay On Task&lt;/H5&gt;
&lt;P&gt;Many AI systems perform well on isolated prompts. Real-world agents are different.&lt;/P&gt;
&lt;P&gt;They need to plan, recover from errors, navigate tools, and continue making progress over extended workflows with minimal supervision.&lt;/P&gt;
&lt;P&gt;Grok 4.6 is designed for this type of long-horizon agentic work, delivering strong performance on benchmarks that measure sustained software engineering and agent execution.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Benchmark highlights&lt;/STRONG&gt;&lt;/P&gt;
&lt;img&gt;&lt;STRONG&gt;Terminal-Bench 3.0&lt;/STRONG&gt;&lt;/img&gt;
&lt;P class="lia-align-center"&gt;&lt;EM&gt;* Benchmarks are from model provider, see &lt;A class="lia-external-url" href="https://media.x.ai/v1/website/card-4p6-4cd2dc57.pdf" target="_blank" rel="noopener"&gt;details&lt;/A&gt;.&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;Terminal-Bench 3.0&lt;STRONG&gt; &lt;/STRONG&gt;evaluates a model’s ability to complete multi-step, terminal-based execution that closely resemble real-world engineering workflows. For developers building coding agents, software engineering copilots, or autonomous workflows, Grok 4.6 is optimized not only for generating output, but for completing work.&lt;/P&gt;
&lt;H5&gt;Engineering-Grade Reasoning Beyond Software&lt;/H5&gt;
&lt;P&gt;Grok 4.6 extends beyond traditional coding into technical and engineering workflows. From computer-aided design (CAD) and procedural design, to complex engineering analysis, it is designed to reason through structured technical challenges and support multi-step problem solving.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Benchmark highlights&lt;/STRONG&gt;&lt;/P&gt;
&lt;img&gt;&lt;STRONG&gt;3DCodeBench&lt;/STRONG&gt;&lt;/img&gt;
&lt;P class="lia-align-center"&gt;&lt;EM&gt;* Benchmarks are from model provider, see &lt;A href="https://media.x.ai/v1/website/card-4p6-4cd2dc57.pdf" target="_blank" rel="noopener"&gt;details&lt;/A&gt;.&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;3DCodeBench evaluates agentic procedural 3D modeling via code, testing how effectively models can reason about and generate complex engineering designs through code. The result highlights Grok 4.6's potential beyond traditional software development, supporting technical workloads across manufacturing, robotics, product design, and industrial AI.&lt;/P&gt;
&lt;H5&gt;Enterprise Agents That Deliver Work, Not Just Answers&lt;/H5&gt;
&lt;P&gt;The most valuable AI workflows often don't end with an answer. They produce business-ready deliverables, such as reports, presentations, spreadsheets, recommendations, and other business artifacts that can be acted on immediately.&lt;/P&gt;
&lt;P&gt;Grok 4.6 demonstrates strong performance on evaluations designed to measure knowledge work and enterprise productivity.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Benchmark highlights&lt;/STRONG&gt;&lt;/P&gt;
&lt;img&gt;&lt;STRONG&gt;AA Briefcase&lt;/STRONG&gt;&lt;/img&gt;
&lt;P class="lia-align-center"&gt;&lt;EM&gt;* Benchmarks are from model provider, see &lt;A href="https://media.x.ai/v1/website/card-4p6-4cd2dc57.pdf" target="_blank" rel="noopener"&gt;details&lt;/A&gt;.&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;AA Briefcase (Artificial Analysis) evaluates agents on long-horizon, complex professional knowledge-work projects that culminate in deliverables such as spreadsheets, presentations, memos, financial models, and PDFs. The results highlights Grok 4.6’s ability to support enterprise assistants, research agents, and business workflow automation that transform information into actionable business outcomes.&lt;/P&gt;
&lt;H5&gt;Pricing&lt;/H5&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 73.4259%; height: 92px; border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr style="height: 39px;"&gt;&lt;td style="height: 39px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Model &lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 39px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Deployment &lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 39px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Input/1M Tokens &lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 39px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Output/1M Tokens &lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Cache/1M Tokens&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 39px;"&gt;&lt;td style="height: 39px;"&gt;
&lt;P&gt;Grok 4.6&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 39px;"&gt;
&lt;P&gt;Global Standard&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 39px;"&gt;
&lt;P&gt;$2.00&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 39px;"&gt;
&lt;P&gt;$6.00&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$0.50&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 17.298%" /&gt;&lt;col style="width: 18.0556%" /&gt;&lt;col style="width: 20.5808%" /&gt;&lt;col style="width: 22.3474%" /&gt;&lt;col style="width: 21.7183%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H5&gt;Why build with Grok 4.6 in Foundry?&lt;/H5&gt;
&lt;P&gt;As organizations adopt more frontier models, the challenge is no longer accessing models. It's operationalizing them.&lt;/P&gt;
&lt;P&gt;Foundry provides a unified platform to discover, evaluate, and deploy models while applying enterprise-grade governance, security, and operational controls throughout the AI lifecycle. Organizations can assess Grok 4.6 against their own workloads, business requirements, and production constraints, enabling confident model selection and faster deployment.&lt;/P&gt;
&lt;P&gt;With Grok 4.6 in Foundry, developers can:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Evaluate Grok 4.6 alongside other leading frontier models&lt;/LI&gt;
&lt;LI&gt;Run workload-specific evaluations before deployment&lt;/LI&gt;
&lt;LI&gt;Deploy through managed endpoints&lt;/LI&gt;
&lt;LI&gt;Integrate the model into agentic applications and workflows&lt;/LI&gt;
&lt;LI&gt;Operate with enterprise governance and security controls&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The combination of Grok 4.6's strengths in long-horizon execution, engineering reasoning, and enterprise productivity with Foundry's model platform capabilities gives developers another powerful option for building the next generation of AI applications.&lt;/P&gt;
&lt;H5&gt;Get started&lt;/H5&gt;
&lt;P&gt;Grok 4.6 is now available in Microsoft Foundry Models.&lt;STRONG&gt; &lt;/STRONG&gt;Explore the &lt;STRONG&gt;&lt;A class="lia-external-url" href="https://ai.azure.com/catalog/models/grok-4.6" target="_blank" rel="noopener"&gt;model card&lt;/A&gt;&lt;/STRONG&gt; to learn more and get started.&lt;/P&gt;
&lt;P&gt;If you're building coding agents, engineering copilots, research assistants, or enterprise automation solutions, Grok 4.6 offers a compelling combination of reasoning, execution, and technical depth, all through the unified Foundry platform experience.&lt;/P&gt;</description>
      <pubDate>Wed, 26 Aug 2026 17:04:55 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/grok-4-6-comes-to-microsoft-foundry-models-built-for-long/ba-p/4547578</guid>
      <dc:creator>MaheshBalachandran</dc:creator>
      <dc:date>2026-08-26T17:04:55Z</dc:date>
    </item>
    <item>
      <title>Azure DNS + Traffic Manager linked records</title>
      <link>https://techcommunity.microsoft.com/t5/azure-networking-blog/azure-dns-traffic-manager-linked-records/ba-p/4548221</link>
      <description>&lt;P&gt;&lt;STRONG&gt;Introduction&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Traffic Manager linked records creates a direct, managed link between an Azure DNS record set and a Traffic Manager profile. When a query arrives, Azure DNS evaluates the linked Traffic Manager profile internally and returns the appropriate endpoint response directly. For A and AAAA records, this means the client receives endpoint IP addresses without an intermediate CNAME response to trafficmanager.net.&lt;/P&gt;
&lt;P&gt;This practical guide uses a small multi-region Contoso scenario to explain the architecture, configure the feature in the Azure portal, validate the DNS response, and demonstrate endpoint failover. The feature is currently in public preview, so the walkthrough should be tested in a non-production environment before broader adoption.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Note&lt;/STRONG&gt;: DNS-based global routing is powerful, but the traditional integration between a custom domain and Traffic Manager adds an intermediate name that customers must understand, resolve, and govern.&lt;/P&gt;
&lt;H1&gt;The scenario and the problem&lt;/H1&gt;
&lt;P&gt;Contoso operates a public application across North America, Europe, and Asia. Users browse to contoso.com. Azure Traffic Manager monitors the regional endpoints and selects the preferred healthy destination according to the profile routing method.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Traditional CNAME integration can expose the intermediate trafficmanager.net name in the DNS response.&lt;/LI&gt;
&lt;LI&gt;A CNAME cannot be placed at the zone of apex, which makes root-domain routing harder.&lt;/LI&gt;
&lt;LI&gt;An unsigned intermediate DNS hop can interrupt end-to-end DNSSEC validation.&lt;/LI&gt;
&lt;LI&gt;Operations teams must understand which system is authoritative for the customer zone, and which system makes the routing decision.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Design objective&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Preserve Traffic Manager routing methods, endpoint health monitoring, and automatic failover while returning the effective endpoint answer directly from Azure DNS.&lt;/P&gt;
&lt;H1&gt;The solution: Traffic Manager linked records&lt;/H1&gt;
&lt;img&gt;
&lt;P&gt;&lt;EM&gt;Figure 1. Traditional CNAME resolution compared with Traffic Manager linked record.&lt;/EM&gt;&lt;/P&gt;
&lt;/img&gt;
&lt;P&gt;As you see in figure 1 above, the left side shows the traditional path. The client queries the application name in Azure DNS, receives a CNAME that points to a trafficmanager.net profile name, and then performs another lookup before receiving the selected endpoint address. The right side shows the linked-record path. Azure DNS remains authoritative for the Contoso zone, consults the linked Traffic Manager profile internally, and returns the selected endpoint directly.&lt;/P&gt;
&lt;P&gt;The key point is that the Traffic Manager still performs the routing decision and health evaluation. Linked records change the DNS integration and the answer returned to the client; they do not turn Traffic Manager into a reverse proxy or place it in the application's data path.&lt;/P&gt;
&lt;H1&gt;Reference architecture&lt;/H1&gt;
&lt;img&gt;
&lt;P&gt;&lt;EM&gt;Figure 2. Practical multi-region architecture for testing direct resolution and failover.&lt;/EM&gt;&lt;/P&gt;
&lt;/img&gt;
&lt;P&gt;Reference architecture in figure 2 above shows the client queries contoso.com. Azure Public DNS hosts the authoritative zone and contains a record set linked to the tm-profile Traffic Manager profile. The profile uses a configured routing method and continuously evaluates the health of the regional endpoints. Azure DNS uses the Traffic Manager decision and returns the effective endpoint answer to the client.&lt;/P&gt;
&lt;P&gt;For a simple lab, two endpoints are enough. Configure the first endpoint with priority 1 and the second with priority 2. The third regional endpoint in the diagram illustrates how the same design scales to additional regions. The record set can be created at the zone apex by leaving the record name empty, or for a subdomain such as www.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Note:&lt;/STRONG&gt; Traffic Manager is DNS based. After name resolution, the client connects directly to the selected public endpoint. Application traffic does not pass through the Traffic Manager.&lt;/P&gt;
&lt;H1&gt;Prerequisites&lt;/H1&gt;
&lt;UL&gt;
&lt;LI&gt;An Azure subscription and permissions to create or update Azure DNS and Traffic Manager resources.&lt;/LI&gt;
&lt;LI&gt;A public domain hosted in Azure DNS and delegated to the Azure DNS name servers.&lt;/LI&gt;
&lt;LI&gt;Two public application endpoints that can return visibly different responses for failover testing.&lt;/LI&gt;
&lt;LI&gt;The Microsoft.Network resource provider is registered in every subscription that contains the DNS zone or Traffic Manager profile.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H1&gt;Implementation in the Azure portal&lt;/H1&gt;
&lt;P&gt;The portal workflow has two main parts: create a strictly typed Traffic Manager profile, then create an Azure DNS record set that links to that profile.&lt;/P&gt;
&lt;H2&gt;Create the Traffic Manager profile&lt;/H2&gt;
&lt;P&gt;&lt;STRONG&gt;Step1:&lt;/STRONG&gt; Open Traffic Manager profiles, select create.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;STRONG&gt;Step 2:&lt;/STRONG&gt;&amp;nbsp;Configure the profile and instance details, fill out the subscription, resource group, name(tm-profile), routing method: priority, record type: A, and resource group location.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;STRONG&gt;Step 3:&lt;/STRONG&gt;&amp;nbsp;Add endpoints (Add tmendpoint-1 as the public IP address and external endpoint type with priority 1.&lt;/P&gt;
&lt;img /&gt;&lt;img /&gt;
&lt;P&gt;&lt;STRONG&gt;Step 4:&lt;/STRONG&gt;&amp;nbsp;Repeat Step 3 to add the secondary endpoint (Add tmendpoint-2 as the public IP address and external endpoint type with priority 2)&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Step 5:&lt;/STRONG&gt; Confirm endpoint health (Wait until both endpoints report online before testing DNS behavior.)&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;STRONG&gt;Note:&lt;/STRONG&gt;&amp;nbsp;The record type selection creates a strictly typed profile. Select A for IPv4 answers, AAAA for IPv6 answers, or CNAME for canonical-name answers. The type is used to validate compatibility between the Traffic Manager profile and the linked DNS record.&lt;/P&gt;
&lt;H2&gt;Create the linked record in Azure DNS&lt;/H2&gt;
&lt;P&gt;&lt;STRONG&gt;Step 1: &lt;/STRONG&gt;Open the Azure DNS zone, select the zone that hosts the application domain, for example contoso.com.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;STRONG&gt;Step 2:&amp;nbsp;&lt;/STRONG&gt;Add DNS record sets, select + record set.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;STRONG&gt;Step 3:&lt;/STRONG&gt;&amp;nbsp;Choose the record name, leave name empty for the zone apex or enter www for a subdomain.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;STRONG&gt;Step 4:&lt;/STRONG&gt; Choose the type, select the record type that matches the Traffic Manager profile, for example, A.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Step 5: &lt;/STRONG&gt;Enable Traffic Management, select Enable Traffic Management (Preview).&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Step 6: &lt;/STRONG&gt;Select the profile, choose the profile subscription and tm-profile, then save the record set.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Note&lt;/STRONG&gt;: The record set contains a managed association to the Traffic Manager profile. The linked record inherits its TTL from the Traffic Manager profile.&lt;/P&gt;
&lt;P&gt;In the Azure DNS record list, the new record should appear as a linked record associated with the Traffic Manager profile. For an apex A record, the conceptual result is:&lt;/P&gt;
&lt;img /&gt;
&lt;H1&gt;Validate the DNS response&lt;/H1&gt;
&lt;P&gt;Use both the Azure portal and direct DNS query. The portal confirms the resource association; the DNS query confirms the serving behavior.&lt;/P&gt;
&lt;H2&gt;Portal validation&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;The record exists in the expected Azure DNS zone.&lt;/LI&gt;
&lt;LI&gt;The record type matches the strictly typed Traffic Manager profile.&lt;/LI&gt;
&lt;LI&gt;The linked profile is tm-profile.&lt;/LI&gt;
&lt;LI&gt;The Traffic Manager endpoints show their expected health state.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Command-line validation&lt;/H2&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;/DIV&gt;
&lt;P&gt;A successful A-record test returns the selected healthy endpoint IP address and does not include a trafficmanager.net CNAME in the answer. Querying for an Azure DNS authoritative name server directly is useful during testing because it reduces ambiguity from local recursive resolver caches.&lt;/P&gt;
&lt;P&gt;The specific endpoint returned depends on the routing method, endpoint health, source resolver location, and cached TTL behavior. Validate the result against the profile configuration rather than expecting one universal IP address.&lt;/P&gt;
&lt;H1&gt;Prove automatic failover&lt;/H1&gt;
&lt;P&gt;A practical demonstration should show that the linked record is not static. The returned answer continues to reflect Traffic Manager health decisions.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN lia-align-center"&gt;&lt;table class="lia-background-color-16 lia-border-color-21 lia-border-style-solid" border="1" style="width: 74.6296%; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;&lt;STRONG&gt;Step&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;&lt;STRONG&gt;Portal action&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;&lt;STRONG&gt;What to configure&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;&lt;STRONG&gt;1&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;Browse to contoso.com&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;Confirm the priority 1 endpoint serves the application.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;&lt;STRONG&gt;2&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;Stop priority 1 endpoint&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;Uncheck the enable endpoint config in tm-profile or stop the VM&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;&lt;STRONG&gt;3&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;Monitor Traffic Manager&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;Wait until the endpoint status changes to unhealthy.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;&lt;STRONG&gt;4&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;Clear local cache&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;On Windows, run ipconfig /flushdns before querying again.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;&lt;STRONG&gt;5&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;Query or browse again&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;Confirm the priority 2 endpoint is now selected.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;&lt;STRONG&gt;6&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;Restore the primary endpoint&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;Start the endpoint and confirm it returns to online.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 7.45168%" /&gt;&lt;col style="width: 33.4344%" /&gt;&lt;col style="width: 59.0906%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H1&gt;Why this matters for real architecture&lt;/H1&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table class="lia-background-color-16 lia-border-color-21 lia-border-style-solid" border="1" style="width: 85.0926%; height: 442px; border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;&lt;STRONG&gt;Cleaner DNS answers&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The client receives the effective endpoint result directly instead of resolving an intermediate trafficmanager.net name.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;&lt;STRONG&gt;Zone-apex routing&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Linked records can be used at the root domain, where DNS standards do not allow a CNAME record.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;&lt;STRONG&gt;DNSSEC compatibility&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Keeping resolution inside Azure DNS removes the unsigned intermediate trafficmanager.net hop from the customer response chain.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;&lt;STRONG&gt;Operational type safety&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Strictly typed profiles help ensure that the linked DNS record type and Traffic Manager endpoint response type remain compatible.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-border-color-21"&gt;
&lt;P&gt;&lt;STRONG&gt;Preserved routing intelligence&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Traffic Manager routing methods, endpoint monitoring, and automatic failover continue to determine the selected endpoint.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 100.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H1&gt;Preview and operational considerations&lt;/H1&gt;
&lt;UL&gt;
&lt;LI&gt;Treat the feature as preview and review the Microsoft Azure preview supplemental terms before production use.&lt;/LI&gt;
&lt;LI&gt;Confirm that the DNS zone and Traffic Manager profile subscriptions have the required resource provider registration.&lt;/LI&gt;
&lt;LI&gt;Plan the record type before creating the Traffic Manager profile because the strictly typed profile value cannot be changed after creation.&lt;/LI&gt;
&lt;LI&gt;Remember that the record TTL is inherited from the Traffic Manager profile.&lt;/LI&gt;
&lt;LI&gt;Test routing, endpoint health, DNSSEC behavior, zone-apex resolution, and rollback in a controlled environment.&lt;/LI&gt;
&lt;LI&gt;Document the previous DNS record so the change can be reversed if the validation plan fails.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H1&gt;Conclusion&lt;/H1&gt;
&lt;P&gt;Traffic Manager linked records is a focused DNS improvement with meaningful architectural impact.&lt;/P&gt;
&lt;P&gt;The feature keeps Azure DNS authoritative for the customer domain, preserves Traffic Manager health-aware routing, and removes the need for the client to follow an intermediate trafficmanager.net hop. That creates a cleaner DNS response path and enables scenarios that were difficult with a traditional CNAME, especially zone-apex routing and DNSSEC-protected domains.&lt;/P&gt;
&lt;P&gt;The strongest way to evaluate the feature is through a small, observable lab: create a strictly typed profile, link an Azure DNS record, query the authoritative DNS server, and deliberately fail the primary endpoint. When the DNS response changes to the secondary healthy destination without exposing an intermediate Traffic Manager name, the value becomes immediately clear.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Final takeaway&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;A small change in DNS integration can simplify the client experience without giving up Traffic Manager routing intelligence, monitoring, or failover.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;I hope you enjoyed it!&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H1&gt;References&lt;/H1&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/dns/dns-traffic-manager-linked-records" target="_blank" rel="noopener" data-lia-auto-title-active="0" data-lia-auto-title="Traffic Manager Linked Records overview - Azure DNS"&gt;Traffic Manager linked records overview - Azure DNS&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/dns/tutorial-traffic-manager-linked-records-portal" target="_blank" rel="noopener" data-lia-auto-title-active="0" data-lia-auto-title="Tutorial: Create a Traffic Manager Linked Record - Azure portal - Azure DNS"&gt;Tutorial: Create a Traffic Manager linked record - Azure portal - Azure DNS&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/dns/tutorial-traffic-manager-linked-records-cli" target="_blank" rel="noopener"&gt;Create a Traffic Manager linked record using Azure CLI&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/traffic-manager/traffic-manager-overview" target="_blank" rel="noopener"&gt;Azure Traffic Manager overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/traffic-manager/traffic-manager-how-it-works" target="_blank" rel="noopener"&gt;How Azure Traffic Manager works&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/dns/dns-faq" target="_blank" rel="noopener"&gt;Azure DNS FAQ&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;EM&gt;&lt;STRONG&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; Visual note:&lt;/STRONG&gt; Figures in this document were generated with Microsoft Copilot for explanatory use.&lt;/EM&gt;&lt;/P&gt;</description>
      <pubDate>Wed, 26 Aug 2026 15:59:25 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-networking-blog/azure-dns-traffic-manager-linked-records/ba-p/4548221</guid>
      <dc:creator>atiy</dc:creator>
      <dc:date>2026-08-26T15:59:25Z</dc:date>
    </item>
    <item>
      <title>Kick Off the Fall by Building the Future: Join Microsoft's Government Agent-a-thon in Arlington, VA</title>
      <link>https://techcommunity.microsoft.com/t5/public-sector-blog/kick-off-the-fall-by-building-the-future-join-microsoft-s/ba-p/4550532</link>
      <description>&lt;P&gt;On &lt;STRONG&gt;September 29, 2026&lt;/STRONG&gt;, Microsoft is launching the first event in our &lt;STRONG&gt;FY27 Frontier Innovation Series&lt;/STRONG&gt; with an immersive, hands-on experience exclusively designed for &lt;STRONG&gt;Government, Defense, and Education organizations in the Washington DC region&lt;/STRONG&gt;. Hosted at the Microsoft Innovation Hub in Arlington, Virginia, this collaborative workshop will help attendees gain practical skills, build working agents, and discover new ways to unlock value from Microsoft 365 Copilot.&lt;/P&gt;
&lt;H3&gt;The Next Evolution of AI is Here&lt;/H3&gt;
&lt;P&gt;Most organizations have begun exploring generative AI. Many are already seeing value with Microsoft 365 Copilot. But now the conversation is shifting toward something even more powerful: &lt;STRONG&gt;agents&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P&gt;Agents can help automate recurring tasks, deliver information in context, support decision-making, and streamline everyday work. The opportunity is significant, but many organizations are still asking:&lt;/P&gt;
&lt;P&gt;&lt;EM&gt;When should I use Copilot? When should I use an agent? And how do I actually build one?&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;This Agent-a-thon is designed to answer those questions through hands-on learning and real-world application.&lt;/P&gt;
&lt;H3&gt;Learn by Building&lt;/H3&gt;
&lt;P&gt;Rather than spending the day listening to presentations, participants will actively engage in a team-based experience focused on designing, testing, and refining agents using Microsoft 365 Copilot Agent Builder and Copilot Studio.&lt;/P&gt;
&lt;P&gt;Working in small groups, you'll explore practical scenarios, experiment with natural language instructions, and leverage Microsoft 365 data to create solutions that support productivity and mission outcomes.&lt;/P&gt;
&lt;P&gt;No advanced development experience is required.&lt;/P&gt;
&lt;P&gt;Whether you're an IT leader, innovation champion, business stakeholder, or Microsoft 365 Copilot user looking to take the next step, you'll have the opportunity to learn by doing.&lt;/P&gt;
&lt;H3&gt;Real-World Outcomes, Not Just Theory&lt;/H3&gt;
&lt;P&gt;One of the reasons Agent-a-thons have become so popular is their focus on tangible outcomes.&lt;/P&gt;
&lt;P&gt;The goal isn't simply to learn about agents.&lt;/P&gt;
&lt;P&gt;The goal is to leave with ideas, prototypes, and repeatable patterns that can be brought back to your organization.&lt;/P&gt;
&lt;P&gt;Throughout the day you'll discover:&lt;/P&gt;
&lt;P&gt;How to decide when Microsoft 365 Copilot is the right solution and when agents provide additional value&lt;/P&gt;
&lt;P&gt;How to design personal productivity agents using natural language and organizational knowledge&lt;/P&gt;
&lt;P&gt;How to identify opportunities to streamline workflows and reduce repetitive work&lt;/P&gt;
&lt;P&gt;How to collaborate with peers from other government organizations and learn from their approaches&lt;/P&gt;
&lt;P&gt;How to create reusable solutions that can inspire future AI initiatives&lt;/P&gt;
&lt;P&gt;By the end of the event, you'll understand not only what agents are, but how they can be applied to solve meaningful challenges within your own environment.&lt;/P&gt;
&lt;H3&gt;Connect with Peers Across Government&lt;/H3&gt;
&lt;P&gt;One of the most valuable aspects of this experience is the opportunity to collaborate with others who are navigating similar AI adoption journeys.&lt;/P&gt;
&lt;P&gt;Participants will share ideas, exchange lessons learned, receive feedback from Microsoft experts, and discover new approaches that can accelerate success.&lt;/P&gt;
&lt;P&gt;Often the most impactful innovation comes from seeing how others are solving similar problems.&lt;/P&gt;
&lt;H3&gt;Why You Should Attend&lt;/H3&gt;
&lt;P&gt;If you're:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Curious about agents and agentic AI&lt;/LI&gt;
&lt;LI&gt;Looking to get more value from Microsoft 365 Copilot&lt;/LI&gt;
&lt;LI&gt;Hoping to improve prompting and AI solution design skills&lt;/LI&gt;
&lt;LI&gt;Exploring ways to accelerate productivity across your organization&lt;/LI&gt;
&lt;LI&gt;Seeking practical, low-code and no-code approaches to innovation&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Then this event is for you.&lt;/P&gt;
&lt;P&gt;Whether you're just beginning your AI journey or looking to move from experimentation to implementation, the Agent-a-thon offers a unique opportunity to gain hands-on experience and build confidence with emerging AI capabilities.&lt;/P&gt;
&lt;H3&gt;Event Agenda&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;9:00 AM – 9:30 AM&lt;/STRONG&gt;&lt;BR /&gt;Arrival &amp;amp; Networking&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;9:30 AM – 10:00 AM&lt;/STRONG&gt;&lt;BR /&gt;Generative AI and Government Use Cases&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;10:00 AM – 10:45 AM&lt;/STRONG&gt;&lt;BR /&gt;Microsoft 365 Copilot &amp;amp; Agents Masterclass&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;11:00 AM – 3:00 PM&lt;/STRONG&gt;&lt;BR /&gt;Team-Based Labs, Building, and Hacking&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;3:00 PM – 4:00 PM&lt;/STRONG&gt;&lt;BR /&gt;Team Presentations, Showcase &amp;amp; Awards&lt;/P&gt;
&lt;H3&gt;Be Part of the Future of Government Innovation&lt;/H3&gt;
&lt;P&gt;Organizations that succeed with AI won't simply consume technology. They'll learn how to shape it around their mission, people, and processes.&lt;/P&gt;
&lt;P&gt;This Agent-a-thon provides a safe, collaborative, and highly interactive environment to do exactly that.&lt;/P&gt;
&lt;P&gt;Join Microsoft and your government peers as we kick off FY27 by exploring what's possible with Microsoft 365 Copilot, Agent Builder, and Copilot Studio.&lt;/P&gt;
&lt;P&gt;Come curious. Leave inspired. Leave with an agent.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;We can't wait to see you in Arlington.&lt;/STRONG&gt; 🚀&lt;/P&gt;
&lt;P&gt;&lt;A href="https://nam06.safelinks.protection.outlook.com/?url=https%3A%2F%2Fmsevents.microsoft.com%2Fevent%3Fid%3D2571040340&amp;amp;data=05%7C02%7CAbby.Quinn%40microsoft.com%7C9f2ba80102fe43e2a8b908df0207e0e4%7C72f988bf86f141af91ab2d7cd011db47%7C1%7C0%7C639231904997302707%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&amp;amp;sdata=snvjOPCublctAS1BbQFl9CvczrM%2FT8ntPALBK5%2FtTo0%3D&amp;amp;reserved=0" target="_blank"&gt;&lt;STRONG&gt;Register here.&lt;/STRONG&gt;&lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Wed, 26 Aug 2026 15:57:29 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/public-sector-blog/kick-off-the-fall-by-building-the-future-join-microsoft-s/ba-p/4550532</guid>
      <dc:creator>Abby_Quinn</dc:creator>
      <dc:date>2026-08-26T15:57:29Z</dc:date>
    </item>
    <item>
      <title>M365 Copilot Government Roadshow is BACK! Kicking off in Atlanta with our First Ever Agent Bootcamp!</title>
      <link>https://techcommunity.microsoft.com/t5/public-sector-blog/m365-copilot-government-roadshow-is-back-kicking-off-in-atlanta/ba-p/4550531</link>
      <description>&lt;P&gt;This is exactly the kind of event government customers have been asking for: less theory, more building. If you've been hearing about AI agents and wondering how to turn the promise into practical outcomes for your agency, this is your opportunity to roll up your sleeves and create something real.&lt;/P&gt;
&lt;H3&gt;M365 Copilot Government Roadshow: Agent Bootcamp&lt;/H3&gt;
&lt;P&gt;&lt;A class="lia-external-url" href="https://ms-workshops.cloudevents.ai/ms-innovation-workshops/events/3705C44B-900D-42CF-AF9E-888391EC9D84" target="_blank"&gt;Register Here&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;Build Your First AI Agent in a Day&lt;/P&gt;
&lt;P&gt;The next wave of AI innovation isn't just about asking better questions. It's about creating intelligent agents that can help automate work, streamline processes, and support mission outcomes. On September 29 in Atlanta, Microsoft is bringing together government leaders, innovators, and Microsoft 365 Copilot users for an immersive, hands-on Agent Bootcamp designed to help you go from concept to working prototype in a single day.&lt;/P&gt;
&lt;P&gt;This isn't another presentation-heavy event where you sit through hours of slides. It's a collaborative build experience where you'll learn, design, experiment, and create alongside Microsoft experts and government peers facing many of the same challenges as your organization.&lt;/P&gt;
&lt;H3&gt;Why Agents Matter&lt;/H3&gt;
&lt;P&gt;Across government, organizations are exploring how AI can help reduce administrative burden, improve employee productivity, accelerate decision making, and enhance citizen services. Agents represent one of the most exciting opportunities to achieve these outcomes.&lt;/P&gt;
&lt;P&gt;Think of agents as purpose-built AI assistants that can be grounded in your organization's knowledge, follow specific instructions, and help users complete real work. Whether you're looking to streamline internal processes, improve knowledge access, or support frontline teams, agents can play a critical role in your AI strategy.&lt;/P&gt;
&lt;P&gt;The question is no longer if agents will impact government operations.&lt;/P&gt;
&lt;P&gt;The question is: How do you get started?&lt;/P&gt;
&lt;P&gt;That's exactly what this bootcamp is designed to answer.&lt;/P&gt;
&lt;H3&gt;What You'll Experience&lt;/H3&gt;
&lt;P&gt;Throughout the day, you'll be guided through a structured journey that combines learning with hands-on application.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;You'll discover:&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;What agents are and how they differ from traditional AI experiences&lt;/LI&gt;
&lt;LI&gt;How to identify high-value government use cases&lt;/LI&gt;
&lt;LI&gt;Best practices for instructions, grounding, and knowledge sources&lt;/LI&gt;
&lt;LI&gt;How to create agents using Microsoft 365 Copilot Agent Builder and Copilot Studio Lite&lt;/LI&gt;
&lt;LI&gt;Approaches for testing, refining, and scaling agent solutions&lt;/LI&gt;
&lt;LI&gt;Most importantly, you'll apply everything you learn immediately.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;From Idea to Prototype&lt;/P&gt;
&lt;P&gt;The heart of the event is an interactive build sprint where participants work in teams to solve real government-inspired scenarios.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;You Will:&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Explore practical use cases&lt;/P&gt;
&lt;P&gt;Collaborate with peers&lt;/P&gt;
&lt;P&gt;Design an agent solution&lt;/P&gt;
&lt;P&gt;Build and test a working prototype&lt;/P&gt;
&lt;P&gt;Share your work during team showcases&lt;/P&gt;
&lt;P&gt;Learn from the creativity and innovation of other groups&lt;/P&gt;
&lt;P&gt;By the end of the day, you'll leave with more than knowledge. You'll leave with hands-on experience, actionable frameworks, and a tangible example of what's possible when you combine government expertise with AI innovation.&lt;/P&gt;
&lt;P&gt;No Coding Required&lt;/P&gt;
&lt;P&gt;One of the biggest misconceptions about agents is that they require deep technical expertise.&lt;/P&gt;
&lt;P&gt;They don't!&lt;/P&gt;
&lt;P&gt;Whether you're a business leader, mission owner, IT professional, innovation champion, or existing Microsoft 365 Copilot user, this workshop is designed for you. The focus is on problem solving, design thinking, and practical application rather than software development.&lt;/P&gt;
&lt;P&gt;If you can identify a challenge worth solving, you can participate in building an agent.&lt;/P&gt;
&lt;H3&gt;Who Should Attend?&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;This event is ideal for:&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Government IT leaders&lt;/P&gt;
&lt;P&gt;Digital transformation teams&lt;/P&gt;
&lt;P&gt;Program and mission owners&lt;/P&gt;
&lt;P&gt;AI champions&lt;/P&gt;
&lt;P&gt;Business decision makers&lt;/P&gt;
&lt;P&gt;Technical decision makers&lt;/P&gt;
&lt;P&gt;Microsoft 365 Copilot users interested in expanding into agents&lt;/P&gt;
&lt;P&gt;Whether your agency is just beginning its AI journey or already seeing success with Microsoft 365 Copilot, you'll gain valuable insights and practical experience that can help shape your next phase of adoption.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Event Details&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;September 29, 2026&lt;/P&gt;
&lt;P&gt;9:00 AM – 4:00 PM EDT&lt;/P&gt;
&lt;P&gt;Microsoft Atlanta South Building&lt;/P&gt;
&lt;P&gt;200 17th Street NW&lt;/P&gt;
&lt;P&gt;1st Floor MPR 1.1A&lt;/P&gt;
&lt;P&gt;Atlanta, GA 30363&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Agenda Highlights&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Agent Foundations &amp;amp; Live Demonstration&lt;/P&gt;
&lt;P&gt;Team Formation &amp;amp; Use Case Selection&lt;/P&gt;
&lt;P&gt;Hands-On Agent Build Sprint&lt;/P&gt;
&lt;P&gt;Team Showcase &amp;amp; Awards&lt;/P&gt;
&lt;P&gt;Take the Next Step in Your AI Journey&lt;/P&gt;
&lt;P&gt;The most successful organizations are moving beyond AI exploration and into AI creation. This bootcamp provides a unique opportunity to learn directly from experts, collaborate with peers, and experience firsthand how agents can support government missions.&lt;/P&gt;
&lt;P&gt;Seats are intentionally limited to ensure an interactive, high-value experience for every attendee.&lt;/P&gt;
&lt;P&gt;If you've been waiting for the right moment to start building with agentic AI, this is it.&lt;/P&gt;
&lt;P&gt;Register today and join us in Atlanta to build your first AI agent in a day.&lt;/P&gt;
&lt;P&gt;&lt;A href="https://ms-workshops.cloudevents.ai/ms-innovation-workshops/events/3705C44B-900D-42CF-AF9E-888391EC9D84" target="_blank" rel="noopener"&gt;Register here.&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;Registration closes September 27, 2026.&lt;/P&gt;</description>
      <pubDate>Wed, 26 Aug 2026 15:56:45 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/public-sector-blog/m365-copilot-government-roadshow-is-back-kicking-off-in-atlanta/ba-p/4550531</guid>
      <dc:creator>Abby_Quinn</dc:creator>
      <dc:date>2026-08-26T15:56:45Z</dc:date>
    </item>
    <item>
      <title>Stop paying for idle VMs — safely: ringed start/stop waves for your Azure estate</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-mission-critical-blog/stop-paying-for-idle-vms-safely-ringed-start-stop-waves-for-your/ba-p/4542066</link>
      <description>&lt;P&gt;&lt;EM&gt;Turning non-production VMs off at night is the easiest Azure saving there is. Doing it without causing an outage is the part nobody writes about. This is an open-source scheduler that takes that saving as an&amp;nbsp;&lt;STRONG&gt;application&lt;/STRONG&gt;, in&amp;nbsp;&lt;STRONG&gt;ordered rings&lt;/STRONG&gt;, with&amp;nbsp;&lt;STRONG&gt;two independent safety gates&lt;/STRONG&gt;&amp;nbsp;between you and a real power action — and it runs entirely in your own subscription.&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;It is 3am. Somewhere in your subscription, a dev estate is fully powered and doing absolutely nothing. Nobody has logged in since 18:40. The bill does not care.&lt;/P&gt;
&lt;P&gt;The fix looks trivial for about ten minutes. Write a script, stop everything at 19:00, start everything at 07:00, collect the applause. Then someone runs it against an environment where the database tier and the app tier are just two more entries in the same resource group, the app tier comes up first, spends four minutes failing its health probe, and the platform team spends the morning explaining why the cost-saving initiative caused an incident.&lt;/P&gt;
&lt;P&gt;The saving is easy.&amp;nbsp;&lt;STRONG&gt;The sequencing is what breaks.&lt;/STRONG&gt;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;The whole application, end to end. Everything you are about to see runs against the built-in demo estate with both safety gates off — note the mock-mode banners.&lt;/P&gt;
&lt;P&gt;&lt;A href="https://github.com/zmustafa/AzureVMScheduler" target="_blank" rel="noopener"&gt;&lt;STRONG&gt;Azure VM Scheduler&lt;/STRONG&gt;&lt;/A&gt;&amp;nbsp;is an open-source, MIT-licensed scheduler that models your estate the way you actually talk about it —&amp;nbsp;&lt;EM&gt;applications hold rings, rings hold virtual machines&lt;/EM&gt;&amp;nbsp;— and fans every scheduled occurrence out into an&amp;nbsp;&lt;STRONG&gt;ordered, staggered wave&lt;/STRONG&gt;&amp;nbsp;of per-VM actions. Starts walk the rings forward. Stops unwind them in reverse. And until you deliberately flip two separate switches, none of it touches a real machine.&lt;/P&gt;
&lt;H2&gt;Two problems, not one&lt;/H2&gt;
&lt;H3&gt;The money&lt;/H3&gt;
&lt;P&gt;Deallocating a VM stops the compute meter. That is the whole trick, and it is a good one: a machine that only needs to run 07:00–19:00 on weekdays is genuinely needed for&amp;nbsp;&lt;STRONG&gt;60 of the 168 hours in a week — about 36%&lt;/STRONG&gt;. The other 64% is compute you are buying and nobody is using.&lt;/P&gt;
&lt;P&gt;Be precise about what you keep paying for, because this is where "we'll save 64%" quietly becomes a credibility problem in your next FinOps review:&amp;nbsp;&lt;STRONG&gt;deallocation stops compute charges, not storage&lt;/STRONG&gt;. Managed disks keep billing whether the VM is running or not, as do reserved public IP addresses and anything else with its own meter. The saving is real and it is large, but it is a saving on the compute line, not on the invoice total.&lt;/P&gt;
&lt;H3&gt;The shape&lt;/H3&gt;
&lt;P&gt;Azure already gives you several ways to turn a machine off on a timer, and they work. The catch is that they are shaped like a&amp;nbsp;&lt;EM&gt;machine&lt;/EM&gt;, or like a&amp;nbsp;&lt;EM&gt;tag&lt;/EM&gt;, and applications are not shaped like either. An application has an order. The data tier comes up before the app tier; the front door comes up last; and on the way down it all has to happen in reverse. A canary ring exists precisely so that it is the&amp;nbsp;&lt;EM&gt;first&lt;/EM&gt;&amp;nbsp;thing to come up and the&amp;nbsp;&lt;EM&gt;last&lt;/EM&gt;&amp;nbsp;thing to go down.&lt;/P&gt;
&lt;P&gt;None of that is expressible as "shut this VM down at 19:00", because the unit is wrong. What a large estate needs is a scheduler whose unit of scheduling is the&amp;nbsp;&lt;STRONG&gt;application&lt;/STRONG&gt;, with ordered stages inside it — and that is a modelling problem, sitting on top of primitives Azure already exposes very well.&lt;/P&gt;
&lt;H2&gt;Where this fits alongside what Azure already gives you&lt;/H2&gt;
&lt;P&gt;Azure is not short of building blocks here. Azure Resource Graph will tell you what is deployed across every subscription you can see. Azure Resource Manager exposes&amp;nbsp;start,&amp;nbsp;deallocate&amp;nbsp;and&amp;nbsp;powerOff&amp;nbsp;as first-class operations with proper RBAC around them. Azure Policy, Azure Automation and the platform's own auto-shutdown all sit on that same foundation. This project is not an alternative to any of that — it is a client of it. Discovery goes through Azure Resource Graph, power actions go through ARM, and identity goes through Microsoft Entra ID.&lt;/P&gt;
&lt;P&gt;What it adds is a shape:&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Tool&lt;/th&gt;&lt;th&gt;Designed for&lt;/th&gt;&lt;th&gt;Unit of action&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;VM / DevTest Labs&amp;nbsp;&lt;STRONG&gt;Auto-shutdown&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;A single VM's daily shutdown time&lt;/td&gt;&lt;td&gt;One VM&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Azure Automation&lt;/STRONG&gt;&amp;nbsp;runbooks&lt;/td&gt;&lt;td&gt;Arbitrary automation you author and maintain&lt;/td&gt;&lt;td&gt;Whatever you script&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Start/Stop VMs v2&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Scheduled, sequenced and CPU-triggered start-stop across whole scopes&lt;/td&gt;&lt;td&gt;Subscription, resource group or VM list&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Azure VM Scheduler&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Ordered, application-aware start/stop waves with per-action safety gates&lt;/td&gt;&lt;td&gt;Application → ring → VM&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;Read that table honestly.&amp;nbsp;&lt;STRONG&gt;If a resource group is a good enough boundary for your estate, use the built-in option&lt;/STRONG&gt;&amp;nbsp;— it is less to run, less to patch and less to explain to your auditor. Plenty of dev estates genuinely are flat, and for those this project is over-engineering.&lt;/P&gt;
&lt;P&gt;This exists for the estates where it isn't. Where boot order matters, where a canary ring means something, where "stop everything in&amp;nbsp;rg-dev" would take down the one machine that runs overnight settlement, and where the blast radius of a scheduling mistake is measured in incidents rather than in dollars.&lt;/P&gt;
&lt;H3&gt;A closer look at Start/Stop VMs v2&lt;/H3&gt;
&lt;P&gt;Of the built-in options,&amp;nbsp;&lt;A href="https://learn.microsoft.com/azure/azure-functions/start-stop-v2/overview" target="_blank" rel="noopener"&gt;Start/Stop VMs v2&lt;/A&gt;&amp;nbsp;is the nearest neighbour, so it deserves a proper description rather than a table row. It is a Microsoft-published solution you deploy into your own subscription: an Azure Functions app holding a managed identity in Microsoft Entra ID, five Azure Logic Apps that carry the schedules and call that app with a JSON payload, Azure Storage for execution metadata and queues, and Application Insights behind a shared Azure dashboard with optional email through an action group. Each action is scoped to one or more subscriptions, resource groups, or an explicit VM list — with wildcard exclusions — and machines are ordered within a scope by tagging them&amp;nbsp;sequencestart&amp;nbsp;and&amp;nbsp;sequencestop&amp;nbsp;with values from 1 to N. It also does something this project does not:&amp;nbsp;&lt;STRONG&gt;AutoStop&lt;/STRONG&gt;&amp;nbsp;watches CPU utilisation and stops idle machines via an alert.&lt;/P&gt;
&lt;P&gt;One thing to know if you are choosing today: Microsoft's guidance states that no further development or enhancements are planned for Start/Stop VMs v2 beyond what is needed to keep its components on supported versions. That is ordinary platform evolution, not a warning — it still deploys, it was moved to the .NET 8 isolated worker model in 2024, and it still does what it does well.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;&amp;nbsp;&lt;/th&gt;&lt;th&gt;Start/Stop VMs v2&lt;/th&gt;&lt;th&gt;Azure VM Scheduler&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Where the schedule lives&lt;/td&gt;&lt;td&gt;Azure Logic Apps you deploy and manage&lt;/td&gt;&lt;td&gt;The application's own database, edited in a UI or over its API&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;How machines are grouped&lt;/td&gt;&lt;td&gt;Subscriptions, resource groups or a VM list, with wildcard exclusions&lt;/td&gt;&lt;td&gt;Application → ring → VM, inherited down the tree&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;How order is expressed&lt;/td&gt;&lt;td&gt;A&amp;nbsp;sequencestart&amp;nbsp;/&amp;nbsp;sequencestop&amp;nbsp;tag per machine, processed in ascending order&lt;/td&gt;&lt;td&gt;The ring's sequence, inherited by every machine in it&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Where results land&lt;/td&gt;&lt;td&gt;Application Insights, a shared Azure dashboard and action-group email&lt;/td&gt;&lt;td&gt;Per-run and per-attempt records in the app, with per-attempt retry&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Where it runs&lt;/td&gt;&lt;td&gt;Azure&lt;/td&gt;&lt;td&gt;Azure, or a laptop on SQLite&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;Three differences follow from that, and they are differences of design centre rather than of quality.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Grouping is modelled, not enumerated.&lt;/STRONG&gt;&amp;nbsp;A scope in Start/Stop VMs v2 is written into the payload of the Logic App that owns it. Here the scope is a thing in the model: schedule the application and every ring and machine underneath inherits it, so adding a VM to a ring puts it in the right wave with nothing else to edit and no tag to remember.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Order and overlap are properties of that model.&lt;/STRONG&gt;&amp;nbsp;Sequencing tags are processed in ascending order for both directions, so a reverse-order shutdown is something you author machine by machine in the&amp;nbsp;sequencestop&amp;nbsp;values. Here reverse&amp;nbsp;&lt;EM&gt;is&lt;/EM&gt;&amp;nbsp;the default for stops — the wave diagram further down this post is two rows of the same list read in opposite directions. And because schedules can be inherited, they can overlap, which is why resolution is per action and nearest-wins: a property you need only once inheritance exists.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;The guards are sized to the unit.&lt;/STRONG&gt;&amp;nbsp;When the thing you point a schedule at is an entire application, a mistake is proportionally bigger — so the stop path gets two independent gates,&amp;nbsp;never_stop&amp;nbsp;that a machine inherits from any ancestor rather than an exclusion list per schedule, and an exact-count confirmation. That machinery is the price of the larger scoping unit, not a criticism of a smaller one.&lt;/P&gt;
&lt;P&gt;So: if you want CPU-triggered auto-stop, availability in the US Government cloud, a Microsoft support path, or simply nothing extra to run beyond Azure resources themselves, Start/Stop VMs v2 is the better answer. This project is for the estates where ordering, inheritance and blast radius are the hard part.&lt;/P&gt;
&lt;P&gt;&lt;EM&gt;Start/Stop VMs v2 details verified against its&amp;nbsp;&lt;A href="https://learn.microsoft.com/azure/azure-functions/start-stop-v2/overview" target="_blank" rel="noopener"&gt;overview documentation&lt;/A&gt;, July 2026.&lt;/EM&gt;&lt;/P&gt;
&lt;H2&gt;Model the estate the way you talk about it&lt;/H2&gt;
&lt;P&gt;The hierarchy is deliberately, aggressively shallow:&amp;nbsp;&lt;STRONG&gt;an application holds rings, a ring holds virtual machines, and that is the entire tree.&lt;/STRONG&gt;&amp;nbsp;Exactly two levels, enforced everywhere — on create, on move, on CSV import, on settings import. There is no way to end up with a ring inside a ring inside a ring, which means there is never a debate about what a given schedule actually covers.&lt;/P&gt;
&lt;P&gt;The built-in demo estate shows the idea in about thirty seconds.&amp;nbsp;&lt;STRONG&gt;Zava Commerce&lt;/STRONG&gt;&amp;nbsp;has three rings —&amp;nbsp;Canary&amp;nbsp;(one VM),&amp;nbsp;Pilot&amp;nbsp;(two),&amp;nbsp;Production&amp;nbsp;(four).&amp;nbsp;&lt;STRONG&gt;Zava Analytics&lt;/STRONG&gt;&amp;nbsp;has&amp;nbsp;Batch&amp;nbsp;and&amp;nbsp;Interactive.&amp;nbsp;&lt;STRONG&gt;Zava Intranet&lt;/STRONG&gt;&amp;nbsp;has&amp;nbsp;Pilot&amp;nbsp;and&amp;nbsp;Production. If your estate is tiered rather than ringed, name them&amp;nbsp;Data,&amp;nbsp;App&amp;nbsp;and&amp;nbsp;Web&amp;nbsp;instead; the model does not care what a stage is called, only that stages are&amp;nbsp;&lt;EM&gt;ordered&lt;/EM&gt;.&lt;/P&gt;
&lt;P&gt;Attach a schedule to the application and every ring inherits it. Override one ring and only that ring changes. Override a single VM and only that machine changes. That is the whole inheritance story, and it is short on purpose.&lt;/P&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;The ring board for one application. The sequence number on each ring is the whole ordering model — there is nothing else to configure.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;Waves: what actually happens at 06:30&lt;/H2&gt;
&lt;P&gt;One occurrence produces one&amp;nbsp;&lt;STRONG&gt;run&lt;/STRONG&gt;. One run fans out into one&amp;nbsp;&lt;STRONG&gt;attempt per virtual machine&lt;/STRONG&gt;, ordered by ring sequence, with a configurable&amp;nbsp;&lt;STRONG&gt;stagger&lt;/STRONG&gt;&amp;nbsp;between machines so you never hand ARM several hundred simultaneous power requests and watch it start throttling you.&lt;/P&gt;
&lt;P&gt;Zava Commerce's start wave fires at&amp;nbsp;&lt;STRONG&gt;06:30&lt;/STRONG&gt;&amp;nbsp;with a&amp;nbsp;&lt;STRONG&gt;60-second stagger&lt;/STRONG&gt;. The canary machine goes first. A minute later the two pilot machines. A minute after that the four production machines, one per minute. Seven machines, seven attempts, in a defined order, each recorded.&lt;/P&gt;
&lt;P&gt;The stop wave at&amp;nbsp;&lt;STRONG&gt;20:00&lt;/STRONG&gt;&amp;nbsp;does the same thing backwards.&amp;nbsp;&lt;STRONG&gt;Stops default to&amp;nbsp;reverse&lt;/STRONG&gt;&amp;nbsp;— the last ring in the sequence is the first one down, and the canary ring, which came up first, goes down last. That default is not a convenience; it is the single most important behaviour in the product, because reverse-order shutdown is exactly what a hand-rolled script forgets.&lt;/P&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;One application, two waves. Read the stop row right to left and you have the start row — that is the point. If you only take one picture from this post, take this one.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3&gt;Nearest schedule wins — per action&lt;/H3&gt;
&lt;P&gt;Here is the guarantee that makes overlapping schedules safe to live with. For&amp;nbsp;&lt;STRONG&gt;each action independently&lt;/STRONG&gt;, a VM is acted on by the&amp;nbsp;&lt;EM&gt;nearest&lt;/EM&gt;&amp;nbsp;schedule that targets it: a schedule on the VM beats one on its ring, which beats one on its application. Deeper shadows shallower.&lt;/P&gt;
&lt;P&gt;Because start and stop are resolved separately, a machine ends up with&amp;nbsp;&lt;STRONG&gt;at most one effective start and at most one effective stop&lt;/STRONG&gt;. Stack an application-wide 06:30 start, a ring-level 07:15 start and a per-VM 05:00 start on the same machine and it still starts exactly once, at 05:00. It cannot be started twice. It cannot be stopped twice. That property holds no matter how many schedules sit above it, which is what lets you layer overrides without keeping a mental model of the whole tree.&lt;/P&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;A wave running. Each row appears as its machine's turn comes up, in ring order, one stagger interval apart.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3&gt;Every wave leaves a paper trail&lt;/H3&gt;
&lt;P&gt;Every wave is fully reconstructable afterwards. A run carries a rolled-up status —&amp;nbsp;succeeded,&amp;nbsp;partially_failed,&amp;nbsp;failed,&amp;nbsp;timed_out&amp;nbsp;or&amp;nbsp;cancelled&amp;nbsp;— and each attempt records its own status, whether it ran in real or mock mode, a message, its attempt number, its sequence position and a correlation id. When someone asks "what happened at 06:30 on Tuesday", the answer is a page, not an investigation. Failed attempts can be retried individually, or the whole run can be retried and it will pick up only the ones that failed.&lt;/P&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;One wave, after the fact: the mode it ran in, the correlation id, the roll-up, and every machine it touched. This page is the answer to "what happened at 06:30 on Tuesday".&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;Recurrence you can check before you trust it&lt;/H2&gt;
&lt;P&gt;Schedules come in four flavours:&amp;nbsp;&lt;STRONG&gt;one-time&lt;/STRONG&gt;,&amp;nbsp;&lt;STRONG&gt;daily&lt;/STRONG&gt;,&amp;nbsp;&lt;STRONG&gt;weekly&lt;/STRONG&gt;&amp;nbsp;and full&amp;nbsp;&lt;STRONG&gt;five-field cron&lt;/STRONG&gt;&amp;nbsp;— lists, ranges, steps, month and day names, Sunday as 0. Daily and weekly are stored as friendly fields and translated to cron behind the scenes, so there is exactly one occurrence engine in the codebase rather than one engine and three special cases that disagree at the edges.&lt;/P&gt;
&lt;P&gt;Everything is timezone-aware and stored against an IANA zone (the demo estate runs on&amp;nbsp;America/New_York). Times are&amp;nbsp;&lt;STRONG&gt;wall-clock&lt;/STRONG&gt;: 08:00 stays 08:00 across a daylight-saving change, which is what an operations team means when they say "eight in the morning". And an occurrence that lands in a spring-forward gap — an 02:30 job on the morning the clocks jump from 02:00 to 03:00 — is&amp;nbsp;&lt;STRONG&gt;skipped, not silently shifted&lt;/STRONG&gt;&amp;nbsp;to 03:30. That is a decision, made once, applied consistently, and covered by tests.&lt;/P&gt;
&lt;H3&gt;The preview is the point&lt;/H3&gt;
&lt;P&gt;Cron is where scheduling bugs live. Everybody has confidently written&amp;nbsp;0 0 * * 0&amp;nbsp;and then argued about which day that is.&lt;/P&gt;
&lt;P&gt;So the editor does not ask you to trust it. On every keystroke it asks the&amp;nbsp;&lt;STRONG&gt;server&lt;/STRONG&gt;&amp;nbsp;to describe the recurrence in plain English and return its&amp;nbsp;&lt;STRONG&gt;next five occurrences&lt;/STRONG&gt;&amp;nbsp;— and the server answers using the same engine that will actually fire the schedule. Not a browser re-implementation of cron that agrees with the backend most of the time. The same code path. What the editor shows you is, by construction, what the scheduler will do.&lt;/P&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Watch the right-hand panel as the cron expression is typed. That description and those five dates come back from the server on every keystroke — from the same engine that fires the schedule.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3&gt;Schedules that know when to stop&lt;/H3&gt;
&lt;P&gt;A schedule can carry&amp;nbsp;start_date&amp;nbsp;and&amp;nbsp;end_date&amp;nbsp;as local calendar bounds, and a&amp;nbsp;run_limit&amp;nbsp;budget. When the budget is spent or the end date passes, the schedule flips to status&amp;nbsp;completed&amp;nbsp;and stops producing a next run — rather than sitting there looking enabled while silently never firing again, which is the failure mode that has you debugging a scheduler that is working perfectly. Manual runs do not consume the budget; only scheduler-triggered ones do.&lt;/P&gt;
&lt;H2&gt;Why you can point this at production without flinching&lt;/H2&gt;
&lt;P&gt;This is the section the rest of the post exists for. Seven guarantees, each one a mechanism rather than a promise.&lt;/P&gt;
&lt;H3&gt;The gates&lt;/H3&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt; Two independent gates.&lt;/STRONG&gt;A real Azure start requires the globalENABLE_REAL_AZURE_STARTS&amp;nbsp;setting&amp;nbsp;&lt;STRONG&gt;and&lt;/STRONG&gt;&amp;nbsp;the target tenant's&amp;nbsp;allow_vm_start&amp;nbsp;permission. A real stop requires&amp;nbsp;ENABLE_REAL_AZURE_STOPS&amp;nbsp;&lt;STRONG&gt;and&lt;/STRONG&gt;&amp;nbsp;allow_vm_stop. They are entirely separate. Turning on starts does not, and cannot, enable stops.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt; Re-evaluated on every single attempt.&lt;/STRONG&gt;The permission decision is deliberately kept apart from the expensive work of building an ARM credential. Adapters get cached because acquiring a token is costly; the&lt;EM&gt;decision&lt;/EM&gt;&amp;nbsp;is remade for every attempt. Revoke a tenant's stop permission while a stop wave is halfway through and the remaining attempts stop being real — immediately, not whenever a cached adapter happens to expire.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt; Mock-first.&lt;/STRONG&gt;Until both switches are on, every wave runs against a deterministic mock adapter that records a simulated result. You can build an entire estate, wire up every schedule, preview the next month of occurrences, run the waves, inspect the attempts and show the whole thing to your change board without touching a single machine. A fresh deployment&lt;STRONG&gt;arrives inert&lt;/STRONG&gt;&amp;nbsp;— both gates ship&amp;nbsp;false.&lt;/LI&gt;
&lt;/OL&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;The gates are a visible posture, not a buried environment variable. This is what a fresh deployment looks like, and it is what it keeps looking like until somebody changes it.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3&gt;The guards&lt;/H3&gt;
&lt;OL start="4"&gt;
&lt;LI&gt;&lt;STRONG&gt;never_stop.&lt;/STRONG&gt;Set it on a VM, or on any ancestor, and that machine is removed from every stop wave and every on-demand stop. Zava Payments demonstrates it: its&amp;nbsp;Production&amp;nbsp;ring runs overnight settlement, so it is marked&amp;nbsp;never_stop&amp;nbsp;and the application has a start wave but no stop wave at all. Nothing you do to the parent schedule can pull those machines into a shutdown.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt; Exact-count confirmation.&lt;/STRONG&gt;Selecting machines and pressing&lt;STRONG&gt;Stop now&lt;/STRONG&gt;&amp;nbsp;stops nothing. The dialog names the tenant, states whether it will deallocate or power off, and refuses to arm until you type the machine count back. It is unglamorous and it has saved somebody's afternoon.&lt;/LI&gt;
&lt;/OL&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Eighteen machines selected, but the dialog offers to stop sixteen — two are&amp;nbsp;never_stop&amp;nbsp;and it says so. The confirm button stays dead until&amp;nbsp;16&amp;nbsp;is typed. Three of the seven guarantees, in one dialog.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;OL start="6"&gt;
&lt;LI&gt;&lt;STRONG&gt; Conflict guard.&lt;/STRONG&gt;The scheduler skips an attempt whose VM already has an opposite-action attempt in flight. If a start wave is still working on a machine, a stop wave will not race it.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt; Read-only tenants can never power anything.&lt;/STRONG&gt;A connection marked read-only refuses start and stop outright, gate or no gate, and a disabled connection refuses everything.stop_mode&amp;nbsp;is&amp;nbsp;deallocate&amp;nbsp;by default — the one that actually stops the compute meter — with&amp;nbsp;power_off&amp;nbsp;available per schedule when you need the machine halted but not released.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2&gt;Coverage gaps: the report you didn't know you needed&lt;/H2&gt;
&lt;P&gt;Any scheduling system's second-order failure is&amp;nbsp;&lt;EM&gt;partial&lt;/EM&gt;&amp;nbsp;coverage. Not "the scheduler broke" — the scheduler is fine — but "eleven machines quietly fell outside it". So the overview goes looking for exactly that, in both directions:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Starts but never stops&lt;/STRONG&gt;&amp;nbsp;— machines with an effective start wave and no effective stop wave. These come up every morning and never go down. They are the ones burning money while everyone congratulates themselves on the savings.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Stops but never starts&lt;/STRONG&gt;&amp;nbsp;— the inverse, and the more expensive mistake. These go down tonight and stay down, and you find out at 09:00 tomorrow.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Stop protected&lt;/STRONG&gt;&amp;nbsp;— machines excluded by&amp;nbsp;never_stop, listed explicitly so the exclusion is a visible decision rather than something you rediscover during an incident.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;It also flags start and stop waves that&amp;nbsp;&lt;STRONG&gt;overlap in time&lt;/STRONG&gt;, including the stagger tail, so you find out at design time that your 20:00 stop wave for a forty-machine application is still running when the 20:30 start wave for something else begins.&lt;/P&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Gaps in both directions, named machine by machine. The middle one — three machines that stop tonight and are never started again — is the one that ruins a Monday.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;The same page carries the pre-flight readiness checks (both global gates, tenants that are read-only or missing the permission a live schedule needs, credentials approaching expiry), windowed KPIs with a previous-period delta, a fourteen-bucket trend, the next-24-hours strip, the rollout plan, an application health matrix and reliability statistics.&lt;/P&gt;
&lt;H2&gt;Getting a real estate in without a spreadsheet&lt;/H2&gt;
&lt;P&gt;Nobody types four hundred resource IDs. Three routes in.&lt;/P&gt;
&lt;H3&gt;Paste bare VM names&lt;/H3&gt;
&lt;P&gt;Give it&amp;nbsp;vm-commerce-prod-01&amp;nbsp;and it resolves the subscription and resource group for you through Azure Resource Graph. If a name is ambiguous across subscriptions, it surfaces the candidates and asks you to pick rather than guessing. Duplicates are blocked. You can also&amp;nbsp;&lt;STRONG&gt;browse a subscription&lt;/STRONG&gt;&amp;nbsp;and select machines directly, or&amp;nbsp;&lt;STRONG&gt;paste full resource IDs&lt;/STRONG&gt;&amp;nbsp;if you already have them.&lt;/P&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Four names in. Three are already filed and it tells you exactly which ring holds each; the fourth is unknown and gets offered a tenant to resolve against.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3&gt;Import VMs&lt;/H3&gt;
&lt;P&gt;Point it at a UTF-8 inventory CSV and it runs a validating preview that reports exactly what would be created, updated and skipped before anything is written. The preview is bound to an encrypted, expiring token, and the commit is atomic — a stale or tampered preview is rejected rather than half-applied. The simplest valid file is a single column of VM names.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;Nothing is written during a preview. The commit is all-or-nothing, and a preview that has gone stale is refused rather than partly applied.&lt;/P&gt;
&lt;H2&gt;Who gets to press the button&lt;/H2&gt;
&lt;P&gt;Access control is capability-based, and it is enforced on the API rather than hidden in the UI.&lt;/P&gt;
&lt;P&gt;Users hold roles directly&amp;nbsp;&lt;STRONG&gt;or&lt;/STRONG&gt;&amp;nbsp;through access groups; their effective permissions are the union of both, resolved once per request. Five roles ship built in —&amp;nbsp;admin,&amp;nbsp;operator,&amp;nbsp;auditor,&amp;nbsp;viewer,&amp;nbsp;noaccess&amp;nbsp;— and they are re-seeded on every start, so a capability added in a new version is never left unusable on an existing deployment. Custom roles are free-form on top.&lt;/P&gt;
&lt;P&gt;Sign-in is local, or SSO, or both. SSO is multi-provider: any number of OIDC issuers configured from their discovery documents (Microsoft Entra ID is just the case where the issuer is derived from a directory id), and SAML 2.0 verified with&amp;nbsp;signxml&amp;nbsp;— where the code reads&amp;nbsp;&lt;STRONG&gt;only the signed subtree&lt;/STRONG&gt;&amp;nbsp;the library returns, never the raw document, which is how a whole family of SAML signature-wrapping attacks stops being interesting.&lt;/P&gt;
&lt;P&gt;Three walls are enforced for every request: a&amp;nbsp;noaccess&amp;nbsp;allowlist, a forced-password-change allowlist, and a database-backed per-IP brute-force throttle that catches one attacker spraying many usernames rather than only many attempts at one account. Two lock-out guards return&amp;nbsp;409&amp;nbsp;rather than letting you strand yourself: you cannot remove the last enabled account that can manage users, and you cannot disable local sign-in when no SSO provider is enabled.&lt;/P&gt;
&lt;P&gt;Every action lands in an audit log.&lt;/P&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Users, roles, access groups, live sessions, sign-in policy and the SSO providers — one page, one permission to reach it.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;Where it runs&lt;/H2&gt;
&lt;P&gt;One container image. The built single-page app is served by FastAPI at the same origin, so there is no CORS story, no second container and no reverse-proxy configuration to get wrong. The static mount is registered last, so it can never shadow the API.&lt;/P&gt;
&lt;P&gt;The&amp;nbsp;&lt;A href="https://portal.azure.com/#create/Microsoft.Template/uri/https%3A%2F%2Fraw.githubusercontent.com%2Fzmustafa%2FAzureVMScheduler%2Fmain%2Fdeploy%2Fmain.json" target="_blank" rel="noopener"&gt;&lt;STRONG&gt;Deploy to Azure&lt;/STRONG&gt;&lt;/A&gt;&amp;nbsp;button provisions, in your subscription: a&amp;nbsp;&lt;STRONG&gt;Container App&lt;/STRONG&gt;&amp;nbsp;running the public image,&amp;nbsp;&lt;STRONG&gt;Azure Database for PostgreSQL flexible server&lt;/STRONG&gt;&amp;nbsp;(Burstable B1ms), an&amp;nbsp;&lt;STRONG&gt;Azure Files&lt;/STRONG&gt;&amp;nbsp;share mounted at&amp;nbsp;/app/.data&amp;nbsp;for the encryption key and connection registry, a Container Apps environment and a Log Analytics workspace. Replicas are pinned to one, because the scheduler is in-process by design. Setting&amp;nbsp;privateNetworking&amp;nbsp;at create time injects the environment into a VNet and puts both the database and the storage account behind Private Endpoints with public access disabled.&lt;/P&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Everything one deployment creates. The private-networking switch is create-time only — an existing public deployment has to be redeployed, not flipped.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3&gt;Connecting Azure without storing a secret&lt;/H3&gt;
&lt;P&gt;The Container App gets a&amp;nbsp;&lt;STRONG&gt;system-assigned managed identity&lt;/STRONG&gt;. Grant it Reader on the scope you want to manage plus&amp;nbsp;start&amp;nbsp;/&amp;nbsp;deallocate&amp;nbsp;/&amp;nbsp;powerOff&amp;nbsp;on the target VMs, add a tenant in the app using the&amp;nbsp;default_chain&amp;nbsp;auth method, and&amp;nbsp;&lt;STRONG&gt;no credential is stored anywhere at all&lt;/STRONG&gt;. Where you do store client secrets, they are Fernet-encrypted at rest and never leave the host.&lt;/P&gt;
&lt;img /&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Two tenants, one writable and one read-only, both authenticating as the host identity. The per-tenant permissions are the second half of every gate.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;Or run the whole thing on a laptop against SQLite.&amp;nbsp;DATABASE_URL&amp;nbsp;picks the engine; nothing else about the application changes.&lt;/P&gt;
&lt;H2&gt;How it's judged&lt;/H2&gt;
&lt;P&gt;The claims above are mechanisms, and each one is testable:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;
&lt;BLOCKQUOTE&gt;A VM is never started twice nor stopped twice, whatever the schedule overlap — resolution is per action, and the nearest schedule wins.&lt;/BLOCKQUOTE&gt;
&lt;/LI&gt;
&lt;LI&gt;
&lt;BLOCKQUOTE&gt;The recurrence preview is produced by the&amp;nbsp;&lt;STRONG&gt;same&lt;/STRONG&gt;&amp;nbsp;engine that fires the schedule, so what the editor shows is what the scheduler does.&lt;/BLOCKQUOTE&gt;
&lt;/LI&gt;
&lt;LI&gt;
&lt;BLOCKQUOTE&gt;An occurrence in a DST spring-forward gap is skipped, never silently shifted.&lt;/BLOCKQUOTE&gt;
&lt;/LI&gt;
&lt;LI&gt;
&lt;BLOCKQUOTE&gt;A real Azure power action requires two independent switches, re-checked on&amp;nbsp;&lt;STRONG&gt;every&lt;/STRONG&gt;&amp;nbsp;attempt — so revoking a permission halts a wave already in progress.&lt;/BLOCKQUOTE&gt;
&lt;/LI&gt;
&lt;LI&gt;
&lt;BLOCKQUOTE&gt;never_stop&amp;nbsp;is honoured by scheduled stops and on-demand stops alike.&lt;/BLOCKQUOTE&gt;
&lt;/LI&gt;
&lt;LI&gt;
&lt;BLOCKQUOTE&gt;Every wave is reconstructable after the fact: run status, per-VM attempt, mode, message, attempt number, sequence position and correlation id.&lt;/BLOCKQUOTE&gt;
&lt;/LI&gt;
&lt;LI&gt;
&lt;BLOCKQUOTE&gt;A fresh deployment cannot touch Azure until a human deliberately arms it.&lt;/BLOCKQUOTE&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;What it deliberately isn't&lt;/H2&gt;
&lt;P&gt;Worth knowing before you invest an afternoon:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Single-replica by design.&lt;/STRONG&gt;&amp;nbsp;The scheduler runs in-process and the Bicep pins replicas to 1/1. That is a conscious trade — no leader election, no distributed lock, no split-brain — but it does mean this is not a highly-available control plane.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Virtual machines only.&lt;/STRONG&gt;&amp;nbsp;Not scale sets, not AKS node pools, not App Service plans, not SQL.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Not a cost tool.&lt;/STRONG&gt;&amp;nbsp;It will not price your savings, forecast them or chargeback them. It turns machines off; your existing cost tooling reports on the result.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Two levels, permanently.&lt;/STRONG&gt;&amp;nbsp;If you need a five-deep hierarchy, this is the wrong shape and no amount of configuration will change that.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Try it&lt;/H2&gt;
&lt;P&gt;The fastest honest test is to run it against nothing. Deploy it with the button below — it arrives with both action gates off, so it physically cannot touch a machine — then open&amp;nbsp;&lt;STRONG&gt;Settings → Demo data&lt;/STRONG&gt;&amp;nbsp;and load the sample&amp;nbsp;&lt;STRONG&gt;Zava&lt;/STRONG&gt;&amp;nbsp;estate: four applications, nine rings, eighteen virtual machines and seven start/stop waves, complete with a&amp;nbsp;never_stop&amp;nbsp;ring and an application that starts but never stops, so the coverage-gap detection has something to find. It loads in one click and removes exactly what it created.&lt;/P&gt;
&lt;P&gt;Build your real rollout on top of that, rehearse every wave against the mock adapter, and only then decide whether to arm anything.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H3&gt;Let GitHub Copilot do the setup&lt;/H3&gt;
&lt;P&gt;Open an empty folder, put Copilot Chat in&amp;nbsp;&lt;STRONG&gt;agent mode&lt;/STRONG&gt;, and hand it the repository:&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Clone&amp;nbsp;https://github.com/zmustafa/AzureVMScheduler&amp;nbsp;into this folder, set it up for local development, then start the backend and the frontend.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;That is enough. The repository ships a README with the setup steps and VS Code tasks for both servers, so the agent has something to follow rather than a command line to invent. The API comes up on&amp;nbsp;127.0.0.1:8000&amp;nbsp;and Vite on&amp;nbsp;127.0.0.1:5173&amp;nbsp;proxying&amp;nbsp;/api&amp;nbsp;to it; with no&amp;nbsp;DATABASE_URL&amp;nbsp;set, the whole thing runs on a SQLite file under&amp;nbsp;.data/&amp;nbsp;with nothing else to install. Sign in as&amp;nbsp;admin, change the password when it makes you, and load&amp;nbsp;&lt;STRONG&gt;Settings → Demo data&lt;/STRONG&gt;.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Repository:&lt;/STRONG&gt;&amp;nbsp;&lt;A href="https://github.com/zmustafa/AzureVMScheduler" target="_blank" rel="noopener"&gt;https://github.com/zmustafa/AzureVMScheduler&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Container image:&lt;/STRONG&gt;&amp;nbsp;&lt;A href="https://hub.docker.com/r/zmustafa/azure-vm-scheduler" target="_blank" rel="noopener"&gt;https://hub.docker.com/r/zmustafa/azure-vm-scheduler&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Licence:&lt;/STRONG&gt;&amp;nbsp;MIT — issues and pull requests welcome&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;If you try it, I would genuinely like to hear which part of the safety model you found excessive and which part you found insufficient. That is the argument worth having.&lt;/P&gt;
&lt;P&gt;&lt;A class="lia-external-url" href="https://zeeshan.net/stop-paying-for-idle-vms-safely-ringed-start-stop-waves-for-your-azure-estate" target="_blank" rel="noopener"&gt;Original post&lt;/A&gt;.&lt;/P&gt;
&lt;P&gt;&lt;EM&gt;This is a community open-source project and is not affiliated with or endorsed by Microsoft.&lt;/EM&gt;&lt;/P&gt;</description>
      <pubDate>Wed, 26 Aug 2026 14:49:01 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-mission-critical-blog/stop-paying-for-idle-vms-safely-ringed-start-stop-waves-for-your/ba-p/4542066</guid>
      <dc:creator>zmustafa</dc:creator>
      <dc:date>2026-08-26T14:49:01Z</dc:date>
    </item>
    <item>
      <title>Navigating the Azure Databricks CSP Mandate</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-mission-critical-blog/navigating-the-azure-databricks-csp-mandate/ba-p/4550172</link>
      <description>&lt;UL&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;Mandatory Deadline:&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt; The Compliance Security Profile (CSP) becomes mandatory for processing&amp;nbsp;HIPAA, HITRUST, and IRAP regulated data on Azure Databricks by &lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;September 1, 2026&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;Permanent Change:&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt; Enabling CSP is an irreversible, workspace-level configuration. Reversion requires&amp;nbsp;deleting&amp;nbsp;and recreating the workspace, emphasizing the need for meticulous pre-enablement testing.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;Architectural Impact:&lt;SPAN style="color: rgb(30, 30, 30);" data-contrast="auto"&gt; CSP introduces significant changes, notably requiring&amp;nbsp;VNet&amp;nbsp;encryption which can disrupt existing on-premises connections, and disabling partner-powered AI features by default.&lt;/SPAN&gt;&lt;SPAN style="color: rgb(30, 30, 30);" data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;The evolving regulatory landscape demands rigorous security and compliance measures for data processing in the cloud. For organizations&amp;nbsp;leveraging&amp;nbsp;Azure Databricks, the introduction and upcoming mandatory enforcement of the Compliance Security Profile (CSP) represent a pivotal shift. This profile is not merely a feature toggle;&amp;nbsp;it's&amp;nbsp;a foundational architectural commitment designed to harden your Databricks environment against stringent compliance standards.&amp;nbsp;Understanding its implications and preparing proactively is crucial to ensure uninterrupted operations and continued regulatory adherence.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2 aria-level="2"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;Why CSP Matters&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:160,&amp;quot;335559739&amp;quot;:80}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;The Compliance Security Profile (CSP) is a robust, platform-hardening configuration applied at the workspace level within Azure Databricks. Its primary&amp;nbsp;objective&amp;nbsp;is to&amp;nbsp;facilitate&amp;nbsp;adherence to a broad spectrum of compliance standards by enforcing&amp;nbsp;additional&amp;nbsp;monitoring,&amp;nbsp;utilizing&amp;nbsp;hardened compute images, ensuring inter-node encryption with specific instance types, and implementing stricter controls across Databricks workspaces. This provides a secure baseline for the data plane, incorporating enhanced security monitoring capabilities.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;For organizations handling sensitive data, such as Protected Health Information (PHI) under HIPAA, or data subject to HITRUST and IRAP regulations, CSP is not just beneficial—it is becoming mandatory. Starting &lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;September 1, 2026&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;, all Azure Databricks workspaces&amp;nbsp;processing data under HIPAA, HITRUST, or IRAP will&amp;nbsp;be required&amp;nbsp;to have CSP enabled. Failure to&amp;nbsp;comply&amp;nbsp;will impede the ability to process such regulated workloads.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;img&gt;&lt;SPAN data-contrast="auto"&gt;Conceptual diagram illustrating data protection and complianceonDatabricks.&lt;/SPAN&gt;&lt;/img&gt;
&lt;H4 aria-level="4"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 4"&gt;Key Compliance Standards Supported by CSP&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:80,&amp;quot;335559739&amp;quot;:40}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H4&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;The compliance security profile supports or is&amp;nbsp;required&amp;nbsp;for processing data under a wide array of global and industry-specific compliance standards. This comprehensive coverage underscores its importance for organizations&amp;nbsp;operating&amp;nbsp;in highly regulated sectors:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;C5&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;CCCS Medium (Protected B)&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;FedRAMP&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;HIPAA (Health Insurance Portability and Accountability Act)&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;HITRUST (Health Information Trust Alliance)&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;IRAP (Infosec Registered Assessors Program)&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;ISMAP&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;Korean Financial Security Institute (K-FSI)&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;PCI-DSS (Payment Card Industry Data Security Standard)&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;TISAX&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;UK Cyber Essentials Plus&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2 aria-level="2"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;Understanding CSP's Architectural Implications&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:160,&amp;quot;335559739&amp;quot;:80}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;One of the most critical aspects of CSP to grasp is its permanence. Once enabled on a workspace, or once any regulated data has been processed within a CSP-enabled workspace, the profile cannot be simply toggled off. To revert to a non-CSP state, you would need to&amp;nbsp;delete&amp;nbsp;the existing workspace and provision a new one without the CSP activated. This irreversible nature&amp;nbsp;necessitates&amp;nbsp;a thorough planning and validation phase before implementing CSP in production environments.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H4 aria-level="4"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 4"&gt;Cost Considerations: The Enhanced Security &amp;amp; Compliance Add-on&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:80,&amp;quot;335559739&amp;quot;:40}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H4&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Enabling CSP automatically activates the "Enhanced Security and Compliance add-on" billing, which will incur&amp;nbsp;additional&amp;nbsp;costs. Organizations must factor these increased expenditures into their budgeting and financial planning.&amp;nbsp;Reviewing the Azure Databricks pricing page is essential to understand the full financial impact.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H4 aria-level="4"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 4"&gt;Networking Challenges:&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="heading 4"&gt;VNet&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="heading 4"&gt;&amp;nbsp;Encryption and On-Premises Connectivity&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:80,&amp;quot;335559739&amp;quot;:40}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H4&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;A significant "gotcha" highlighted in various sources is the impact of CSP on network architecture. CSP enforces stricter network security policies, often requiring&amp;nbsp;VNet&amp;nbsp;encryption. If your Azure Databricks workspaces utilize&amp;nbsp;VNet&amp;nbsp;Injection and connect to on-premises systems via ExpressRoute or VPN, this&amp;nbsp;VNet&amp;nbsp;encryption can&amp;nbsp;potentially disrupt existing connections. This&amp;nbsp;necessitates&amp;nbsp;close collaboration with networking teams to verify compatibility, test connectivity in pre-production environments, and potentially reconfigure network components such as routing tables, security groups, or&amp;nbsp;firewall&amp;nbsp;rules.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H4 aria-level="4"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 4"&gt;Feature Impact: AI Assist and Public Preview Limitations&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:80,&amp;quot;335559739&amp;quot;:40}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H4&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;To&amp;nbsp;maintain&amp;nbsp;stringent compliance boundaries, CSP workspaces alter default behaviors and restrict certain features:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;Partner-powered AI Features:&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt; By default, "Partner-powered AI features," which include some AI-powered coding assistants like Genie Code, are disabled in CSP workspaces. While Databricks-hosted models may power Genie Code in CSP workspaces for specific compliance standards (e.g., HIPAA, FedRAMP Moderate, PCI-DSS, HITRUST) in certain regions, organizations relying on these features should assess alternatives or plan for explicit enablement, understanding the potential constraints.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/databricks/security/privacy/security-profile" target="_blank" rel="noopener"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;Preview Features&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="auto"&gt;:&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt; Only a&amp;nbsp;limited&amp;nbsp;set of Public Preview features are supported in CSP workspaces. Workloads dependent on unsupported preview features may face limitations or require adjustments.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2 aria-level="2"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;Preparing for the CSP Mandate&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:160,&amp;quot;335559739&amp;quot;:80}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;To ensure a smooth transition to CSP, a detailed and phased approach is recommended. This involves thorough assessment, strategic planning, rigorous testing, and clear communication across various teams.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;The radar chart above visualizes the current readiness level (Pre-CSP Readiness) versus the desired target state (Post-CSP Target State) across critical preparation areas for the Compliance Security Profile. It underscores the significant effort required in areas like network reconfiguration and thorough testing to achieve full compliance and operational stability.&lt;/SPAN&gt;&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;SPAN style="color: rgb(30, 30, 30); font-size: 24px;"&gt;Inventory and Assessment&lt;/SPAN&gt;
&lt;UL&gt;
&lt;LI&gt;Identify Regulated Workspaces:&lt;SPAN style="color: rgb(30, 30, 30);" data-contrast="auto"&gt; Catalogue all existing Azure Databricks workspaces.&amp;nbsp;Determine&amp;nbsp;which ones currently process, or are planned to process, data subject to HIPAA, HITRUST, or IRAP.&lt;/SPAN&gt;&lt;SPAN style="color: rgb(30, 30, 30);" data-ccp-props="{}"&gt; &lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;Review Data Pipelines:&lt;SPAN style="color: rgb(30, 30, 30);" data-contrast="auto"&gt; Map out all data ingress and egress points for these identified workspaces, including connections to on-premises data sources, other cloud services, and external APIs. This helps&amp;nbsp;identify&amp;nbsp;potential network impacts.&lt;/SPAN&gt;&lt;SPAN style="color: rgb(30, 30, 30);" data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;/OL&gt;
&lt;OL start="2"&gt;
&lt;LI aria-level="4"&gt;
&lt;H4&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 4"&gt; Network Architecture Review&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/H4&gt;
&lt;/LI&gt;
&lt;/OL&gt;
&lt;UL&gt;
&lt;LI style="list-style-type: none;"&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="6" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Evaluate&amp;nbsp;VNet&amp;nbsp;Encryption:&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt; If your workspace uses&amp;nbsp;VNet&amp;nbsp;Injection and has private connectivity to an on-premises network (e.g., via ExpressRoute or VPN), verify that these connections will remain functional under CSP's new encryption requirements.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="6" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;Plan for Reconfiguration:&lt;SPAN style="color: rgb(30, 30, 30);" data-contrast="auto"&gt; Allocate resources for potential network re-architecting, including updating routing tables, security groups, or&amp;nbsp;firewall&amp;nbsp;rules to accommodate encrypted traffic flows.&lt;/SPAN&gt;&lt;SPAN style="color: rgb(30, 30, 30);" data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;OL start="3"&gt;
&lt;LI aria-level="4"&gt;
&lt;H4&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 4"&gt; Feature Impact Analysis&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/H4&gt;
&lt;/LI&gt;
&lt;/OL&gt;
&lt;UL&gt;
&lt;LI style="list-style-type: none;"&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="7" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;AI &amp;amp; Assist Features:&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt; Assess the reliance of your development teams on "Partner-powered AI features" like&amp;nbsp;Genie&amp;nbsp;Code.&amp;nbsp;Determine&amp;nbsp;if Databricks-hosted alternatives meet their needs or if explicit enablement (and its associated compliance considerations) is&amp;nbsp;required.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="7" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;Preview Features:&lt;SPAN style="color: rgb(30, 30, 30);" data-contrast="auto"&gt; Consult official documentation to check for limitations on Public Preview features in CSP workspaces and plan accordingly if your workloads depend on any unsupported features.&lt;/SPAN&gt;&lt;SPAN style="color: rgb(30, 30, 30);" data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;OL start="4"&gt;
&lt;LI aria-level="4"&gt;
&lt;H4&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 4"&gt; Data Handling Responsibilities&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/H4&gt;
&lt;/LI&gt;
&lt;/OL&gt;
&lt;UL&gt;
&lt;LI style="list-style-type: none;"&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="8" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Sensitive Information Fields:&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt; Implement policies to prevent sensitive data from being entered into customer-defined input fields (e.g., workspace names, compute resource names, tags, job names, URLs), as these might be stored or processed outside the compliance boundary.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="8" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;Encryption at Rest:&lt;SPAN style="color: rgb(30, 30, 30);" data-contrast="auto"&gt; Verify that all data&amp;nbsp;containing&amp;nbsp;PHI or other regulated information is encrypted at rest in any storage location interacting with Databricks, including workspace storage accounts.&lt;/SPAN&gt;&lt;SPAN style="color: rgb(30, 30, 30);" data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;OL start="5"&gt;
&lt;LI aria-level="4"&gt;
&lt;H4&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 4"&gt; Execution and Validation Plan&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/H4&gt;
&lt;/LI&gt;
&lt;/OL&gt;
&lt;UL&gt;
&lt;LI style="list-style-type: none;"&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="9" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Staging Environment First:&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt; Always enable CSP on a clone of your production environment in a staging subscription before rolling out to production.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="9" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;Thorough Testing:&lt;SPAN style="color: rgb(30, 30, 30);" data-contrast="auto"&gt; Conduct comprehensive end-to-end tests on the staging workspace to&amp;nbsp;validate&amp;nbsp;data pipelines, network connectivity, user workflows, and performance.&lt;/SPAN&gt;&lt;SPAN style="color: rgb(30, 30, 30);" data-ccp-props="{}"&gt; &lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="9" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;Enablement Method:&lt;SPAN style="color: rgb(30, 30, 30);" data-contrast="auto"&gt; Choose the&amp;nbsp;appropriate tooling&amp;nbsp;for enablement—Azure Portal, Azure CLI, PowerShell, ARM templates, or&amp;nbsp;Terraform—to ensure consistency and automation.&lt;/SPAN&gt;&lt;SPAN style="color: rgb(30, 30, 30);" data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2 aria-level="2"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;Detailed Enablement Methods&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:160,&amp;quot;335559739&amp;quot;:80}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Enabling CSP can be done for both new and existing workspaces using various Azure management tools. Regardless of the method chosen,&amp;nbsp;it's&amp;nbsp;crucial to select the applicable compliance standards during enablement.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H4 aria-level="4"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 4"&gt;Enabling CSP on New Workspaces&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:80,&amp;quot;335559739&amp;quot;:40}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H4&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;For new workspaces, the compliance security profile settings at the account level can control whether&amp;nbsp;it's&amp;nbsp;enabled by default. Account administrators can enable CSP individually during deployment using:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;Azure Portal&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;Azure CLI (e.g., az&amp;nbsp;databricks&amp;nbsp;workspace create --enable-compliance-security-profile --compliance-standards '["HIPAA"]')&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;PowerShell&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;ARM templates&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;Terraform (using&amp;nbsp;the databricks_compliance_security_profile_workspace_setting resource)&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H4 aria-level="4"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 4"&gt;Enabling CSP on Existing Workspaces&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:80,&amp;quot;335559739&amp;quot;:40}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H4&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Existing workspaces can also have CSP enabled through similar methods:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;Azure Portal:&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt; Navigate to Workspace &amp;gt; Settings &amp;gt; Security &amp;amp; compliance &amp;gt; Enable compliance security profile; select standards and save.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;Azure CLI / PowerShell / ARM templates / Terraform:&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt; Utilize&amp;nbsp;these tools to&amp;nbsp;modify&amp;nbsp;the workspace settings and enable CSP, specifying the relevant compliance standards.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2 aria-level="2"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;Understanding Key Technical Considerations&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:160,&amp;quot;335559739&amp;quot;:80}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Beyond the general planning, specific technical details must be understood and addressed to ensure a successful CSP implementation.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H4 aria-level="4"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 4"&gt;Account-Level Requirements&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:80,&amp;quot;335559739&amp;quot;:40}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H4&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Before enabling CSP, ensure your Databricks account meets these prerequisites:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;Your Databricks account must include the Enhanced Security and Compliance add-on.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="auto"&gt;Your Databricks workspace needs to be on the&amp;nbsp;Premium&amp;nbsp;pricing tier.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H4 aria-level="4"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 4"&gt;Security Configuration Table&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:80,&amp;quot;335559739&amp;quot;:40}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H4&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;The table below summarizes critical security configuration aspects influenced by CSP:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="auto"&gt;Security Aspect&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="auto"&gt;Pre-CSP State&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="auto"&gt;Post-CSP State&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="auto"&gt;Preparation Action&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Compute Images&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Standard images&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Hardened compute images&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;No direct action&amp;nbsp;required, platform handles.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Inter-node Encryption&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Configurable/Optional&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Enforced (for specific instance types)&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Ensure use of supported instance types.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Network Encryption (VNet)&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Standard&amp;nbsp;VNet&amp;nbsp;configuration&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;VNet&amp;nbsp;encryption&amp;nbsp;required&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Test on-premises connectivity, reconfigure network routes/firewalls if needed.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;AI Assist Features&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;"Partner-powered AI features" typically enabled&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;"Partner-powered AI features" disabled by default&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Assess reliance, plan explicit&amp;nbsp;enablement&amp;nbsp;or alternative Databricks-hosted models.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Monitoring&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Standard monitoring&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Enhanced security monitoring,&amp;nbsp;additional&amp;nbsp;agents&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Familiarize with new monitoring capabilities and data.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Compliance Scope&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Self-managed controls&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Platform-enforced controls for selected standards&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Select&amp;nbsp;appropriate compliance&amp;nbsp;standards during CSP enablement.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{}"&gt;&lt;SPAN data-contrast="auto"&gt;To further understand the broader context of security and compliance on Databricks, including how CSP fits into the larger ecosystem, watch this video from Databricks:&lt;/SPAN&gt; &lt;/SPAN&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 3"&gt;&lt;A class="lia-external-url" href="https://youtu.be/I9lxH6pzask?si=oXkOgR3o05PJtKYM" target="_blank" rel="noopener"&gt;Databricks Security &amp;amp; Compliance&lt;/A&gt;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:160,&amp;quot;335559739&amp;quot;:80}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2 aria-level="2"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;Re&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;ferences&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:160,&amp;quot;335559739&amp;quot;:80}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;SPAN data-contrast="none"&gt;Compliance security profile:&amp;nbsp;&lt;/SPAN&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/databricks/security/privacy/security-profile" target="_blank" rel="noopener"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;https://learn.microsoft.com/en-us/azure/databricks/security/privacy/security-profile&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="none"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134233117&amp;quot;:false,&amp;quot;134233118&amp;quot;:false,&amp;quot;335551550&amp;quot;:1,&amp;quot;335551620&amp;quot;:1,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559738&amp;quot;:0,&amp;quot;335559739&amp;quot;:0}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="none"&gt;Configure enhanced security and compliance settings:&amp;nbsp;&lt;/SPAN&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/databricks/security/privacy/enhanced-security-compliance" target="_blank" rel="noopener"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;https://learn.microsoft.com/en-us/azure/databricks/security/privacy/enhanced-security-compliance&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="none"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134233117&amp;quot;:false,&amp;quot;134233118&amp;quot;:false,&amp;quot;335551550&amp;quot;:1,&amp;quot;335551620&amp;quot;:1,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559738&amp;quot;:0,&amp;quot;335559739&amp;quot;:0}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="none"&gt;HIPAA, Azure Databricks, Microsoft Learn:&amp;nbsp;&lt;/SPAN&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/databricks/security/privacy/hipaa" target="_blank" rel="noopener"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;https://learn.microsoft.com/en-us/azure/databricks/security/privacy/hipaa&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="none"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134233117&amp;quot;:false,&amp;quot;134233118&amp;quot;:false,&amp;quot;335551550&amp;quot;:1,&amp;quot;335551620&amp;quot;:1,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559738&amp;quot;:0,&amp;quot;335559739&amp;quot;:0}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN data-contrast="none"&gt;What is Azure Virtual Network encryption:&amp;nbsp;&lt;/SPAN&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/virtual-network/virtual-network-encryption-overview" target="_blank" rel="noopener"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;https://learn.microsoft.com/en-us/azure/virtual-network/virtual-network-encryption-overview&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-ccp-props="{&amp;quot;134233117&amp;quot;:false,&amp;quot;134233118&amp;quot;:false,&amp;quot;335551550&amp;quot;:1,&amp;quot;335551620&amp;quot;:1,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559738&amp;quot;:0,&amp;quot;335559739&amp;quot;:0}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;</description>
      <pubDate>Wed, 26 Aug 2026 14:47:42 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-mission-critical-blog/navigating-the-azure-databricks-csp-mandate/ba-p/4550172</guid>
      <dc:creator>anishekkamal</dc:creator>
      <dc:date>2026-08-26T14:47:42Z</dc:date>
    </item>
    <item>
      <title>Connectivity Errors caused by slow logins in a Microsoft Fabric SQL Database</title>
      <link>https://techcommunity.microsoft.com/t5/azure-database-support-blog/connectivity-errors-caused-by-slow-logins-in-a-microsoft-fabric/ba-p/4548281</link>
      <description>&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;By Tanayankar Chakraborty &amp;amp; Sukhwant Kaur&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{}"&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;Issue&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559738&amp;quot;:360,&amp;quot;335559739&amp;quot;:120,&amp;quot;335559740&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;A customer recently reported that&amp;nbsp;their&amp;nbsp;Azure App&amp;nbsp;Service intermittently&amp;nbsp;cannot&amp;nbsp;establish&amp;nbsp;a TDS connection to its Microsoft Fabric SQL database within 15 seconds.&amp;nbsp;Also,&amp;nbsp;Review Queue and Worklists returned HTTP 500 after approximately 15 seconds.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This occurred before any SQL statement executed and was not an Entra login/permission error.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;Error&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559738&amp;quot;:360,&amp;quot;335559739&amp;quot;:120,&amp;quot;335559740&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559738&amp;quot;:360,&amp;quot;335559739&amp;quot;:120,&amp;quot;335559740&amp;quot;:240}"&gt;Here is the error generated in the App service logs:&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN class="lia-text-color-8"&gt;At 2026-07-29 22:42:15 UTC, App Insights recorded:&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN class="lia-text-color-8"&gt;ConnectionError&amp;nbsp;/ ETIMEOUT:&amp;nbsp;Failed to&amp;nbsp;connect to the Fabric SQL endpoint in&amp;nbsp;15000ms.&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN class="lia-text-color-8"&gt;GET /api/review/queue returned HTTP 500 in&amp;nbsp;15,100ms.&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{}"&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;Troubleshooting&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559738&amp;quot;:360,&amp;quot;335559739&amp;quot;:120,&amp;quot;335559740&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{}"&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559738&amp;quot;:360,&amp;quot;335559739&amp;quot;:120,&amp;quot;335559740&amp;quot;:240}"&gt;As stated in our public documentation, throttling by Fabric platform is expected behavior (requests getting delayed by 20 seconds for initial throttling, and requests getting rejected outright if it gets worse) if usage is heavy, especially on a heavy SQL workload on workspace/capacity shared with other artifacts, with a small capacity size (F4, etc).&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;Assuming&amp;nbsp;that's&amp;nbsp;what the issue is, general guidance is to:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;SPAN data-contrast="none"&gt;Decrease the max&amp;nbsp;vCore&amp;nbsp;limit on the affected SQL DB to limit the&amp;nbsp;amount&amp;nbsp;of CUs it can&amp;nbsp;consume ,&amp;nbsp;but obviously there would be a performance tradeoff here though, and throttling can still be caused by other items in the workspace/capacity&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;multilevel&amp;quot;}" data-aria-posinset="2" data-aria-level="1"&gt;&lt;SPAN data-contrast="none"&gt;Increase the capacity (F16+,&amp;nbsp;etc)&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{}"&gt;&lt;SPAN data-contrast="none"&gt;Fabric allows operations to temporarily use more&amp;nbsp;compute&amp;nbsp;than the provisioned SKU through bursting and&amp;nbsp;distributes&amp;nbsp;that usage over future capacity through smoothing. When&amp;nbsp;smoothed&amp;nbsp;usage exceeds the capacity's available future compute, the excess becomes carry-forward usage. Fabric then progressively protects the capacity by delaying or rejecting new operations.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-ccp-props="{}"&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&lt;SPAN data-contrast="none"&gt;Workaround&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559738&amp;quot;:360,&amp;quot;335559739&amp;quot;:120,&amp;quot;335559740&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{}"&gt;&lt;SPAN data-contrast="none"&gt;### Mitigate Active Throttling&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;Microsoft's documented mitigation options include:&lt;/SPAN&gt; &lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;- Wait for the capacity to self-heal while unused capacity burns down carry-forward usage.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;- Temporarily increase the F SKU size. The larger SKU provides more idle capacity and accelerates burndown.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;- Pause and resume&amp;nbsp;the capacity. This creates a billing event for accumulated future usage, and the resumed capacity starts with zero future usage.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;- Consider capacity overage billing where&amp;nbsp;appropriate. Microsoft documents this option as being billed at three times the normal capacity rate.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;Any pause/resume or billing change should be reviewed by the capacity administrator for operational and cost impact.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;### Improve Application Resilience&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;- Use bounded retry with jitter for transient connection failures where the operation is safe to retry.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;- Align SQL connection timeout and retry settings with the application's end-to-end request budget.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;- Capture UTC timestamps, connection/request IDs, operation names, and SQL error numbers for future incidents.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&lt;SPAN data-contrast="none"&gt;Monitoring&lt;/SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;It is recommended to&amp;nbsp;monitor&amp;nbsp;and right size the&amp;nbsp;capacity:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;Use the Microsoft Fabric Capacity Metrics app to&amp;nbsp;monitor:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;- Utilization&amp;nbsp;by workload and item&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;- Interactive and background CU consumption&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;- Carry-forward and burndown&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;- Throttling events and rejected operations&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;- Minutes to&amp;nbsp;burndown&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;A capacity that is throttled might appear in the Fabric Capacity Metrics app as shown below:&lt;/SPAN&gt;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;In the Capacity Metrics app:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;1. Use the&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-contrast="none"&gt;**Utilization**&lt;/SPAN&gt;&lt;SPAN data-contrast="none"&gt; chart and drill through to individual timepoints to&amp;nbsp;identify&amp;nbsp;operations contributing to overages.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;2. Use the&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-contrast="none"&gt;**Throttling**&lt;/SPAN&gt;&lt;SPAN data-contrast="none"&gt; chart to&amp;nbsp;determine&amp;nbsp;whether an overage has crossed a throttling threshold.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;3. Use the&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-contrast="none"&gt;**Overages**&lt;/SPAN&gt;&lt;SPAN data-contrast="none"&gt; tab to review carry-forward, additions, burndown, and cumulative&amp;nbsp;utilization.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;4. Configure capacity&amp;nbsp;alerts&amp;nbsp;so administrators are notified when provisioned CU usage reaches 100%.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{}"&gt;&lt;SPAN data-contrast="none"&gt;### Reduce Peak Concurrent Usage&lt;/SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;- Stagger high-CU SQL and Data Integration operations where possible.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;- Review expensive or highly concurrent SQL database operations.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;- Distribute workloads across capacities if sustained demand requires it.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="none"&gt;- Use historical Capacity Metrics data to select a SKU with sufficient peak headroom.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-ccp-props="{}"&gt;&lt;SPAN data-contrast="none"&gt;References&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-ccp-props="{}"&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&lt;A href="https://learn.microsoft.com/en-us/fabric/enterprise/throttling#throttle-triggers-and-throttle-stages" target="_blank"&gt;Understand capacity throttling and smoothing - Microsoft Fabric | Microsoft Learn&lt;/A&gt;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-ccp-props="{}"&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&lt;A href="https://learn.microsoft.com/en-us/fabric/enterprise/surge-protection#key-capabilities" target="_blank"&gt;Surge protection - Microsoft Fabric | Microsoft Learn&lt;/A&gt;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-ccp-props="{}"&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&lt;A href="https://learn.microsoft.com/en-us/fabric/database/sql/usage-reporting#compute-usage-reporting" target="_blank"&gt;Billing and Utilization Reporting - Microsoft Fabric | Microsoft Learn&lt;/A&gt;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-ccp-props="{}"&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&lt;A href="https://learn.microsoft.com/en-us/fabric/database/sql/control-compute-usage" target="_blank"&gt;Control Compute Usage - Microsoft Fabric | Microsoft Learn&lt;/A&gt;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-ccp-props="{}"&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&lt;A href="https://learn.microsoft.com/en-us/fabric/enterprise/metrics-app" target="_blank"&gt;What is the Microsoft Fabric Capacity Metrics app? - Microsoft Fabric | Microsoft Learn&lt;/A&gt;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-ccp-props="{}"&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335557856&amp;quot;:16777215,&amp;quot;335559739&amp;quot;:240,&amp;quot;335559740&amp;quot;:240}"&gt;&lt;A href="https://learn.microsoft.com/en-us/fabric/real-time-hub/set-alerts-fabric-capacity-overview-events" target="_blank"&gt;Set alerts on Fabric capacity overview events in Real-Time hub - Microsoft Fabric | Microsoft Learn&lt;/A&gt;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 26 Aug 2026 14:04:18 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-database-support-blog/connectivity-errors-caused-by-slow-logins-in-a-microsoft-fabric/ba-p/4548281</guid>
      <dc:creator>Tancy</dc:creator>
      <dc:date>2026-08-26T14:04:18Z</dc:date>
    </item>
    <item>
      <title>Customer review: Earth Knowledge enables compliance and prioritized action for a changing climate.</title>
      <link>https://techcommunity.microsoft.com/t5/marketplace-blog/customer-review-earth-knowledge-enables-compliance-and/ba-p/4540401</link>
      <description>&lt;P&gt;Earth Knowledge's &lt;A href="https://marketplace.microsoft.com/en-us/product/earthknowledgeinc1583873079308.ek3_2021v003?ocid=GTMRewards_Review_ek3_2021v003_56893" target="_blank" rel="noopener"&gt;Integrated Planetary Intelligence&lt;/A&gt;, a cloud-native platform published to &lt;A href="https://marketplace.microsoft.com/en-US/" target="_blank" rel="noopener"&gt;Microsoft Marketplace&lt;/A&gt;, models climate and natural impacts for government and private sector organizations. By analyzing cascading and interrelated impacts globally, Integrated Planetary Intelligence transforms complex, multi-dimensional data about the planet and the environment into actionable data enabling decision-makers to assess current and future risks and make choices that are better for their organizations and the planet. Staffed by world-class scientists, engineers, and analysts, Earth Knowledge guides organizations toward reduced risk and optimized operations.&lt;/P&gt;
&lt;P&gt;&lt;BR /&gt;Microsoft Marketplace interviewed Karen Lynn, VP Sustainability &amp;amp; EHS Management Systems at &lt;A href="https://www.eaton.com/us/en-us.html" target="_blank" rel="noopener"&gt;Eaton&lt;/A&gt;, a global power management company headquartered in the United States, to learn what she had to say about the platform.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;&lt;STRONG&gt;What do you like best about Earth Knowledge Integrated Planetary Intelligence?&lt;/STRONG&gt;&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;The solution is technically rigorous and is supported with knowledgeable technical and scientific personnel. The Earth Knowledge (EK) team is very willing to answer questions and explain the results.&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;&amp;nbsp;&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;&lt;STRONG&gt;How has the product helped your organization?&lt;/STRONG&gt;&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;The solution has enabled Eaton to meet our voluntary and regulatory requirements related to climate modeling and physical risk to our company locations. The data also then helps us prioritize actions needed to prepare for a changing climate.&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;&amp;nbsp;&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;&lt;STRONG&gt;How are customer service and support?&lt;/STRONG&gt;&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;Customer service and support are excellent. The EK team is always responsive. Assistance is also thoughtful and technically supported.&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;&amp;nbsp;&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;&lt;STRONG&gt;Any recommendations to other users considering Earth Knowledge Integrated Planetary Intelligence?&lt;/STRONG&gt;&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;The model is technically rigorous and can be used at various levels of detail based on subscription level.&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;&amp;nbsp;&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;&lt;STRONG&gt;What is your overall rating for this product?&lt;/STRONG&gt;&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;5 out of 5 stars.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Cloud marketplaces are transforming the way businesses find, try, and deploy applications to help their digital transformation. &lt;A href="https://marketplace.microsoft.com/" target="_blank" rel="noopener"&gt;Learn more about Microsoft Marketplace&lt;/A&gt; and find ways to discover the right application for your business needs.&lt;/P&gt;</description>
      <pubDate>Wed, 26 Aug 2026 13:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/marketplace-blog/customer-review-earth-knowledge-enables-compliance-and/ba-p/4540401</guid>
      <dc:creator>Nikhil_Viswanathan</dc:creator>
      <dc:date>2026-08-26T13:00:00Z</dc:date>
    </item>
    <item>
      <title>ICYMI: New Azure SQL Foundations video series with GitHub samples</title>
      <link>https://techcommunity.microsoft.com/t5/azure-sql-blog/icymi-new-azure-sql-foundations-video-series-with-github-samples/ba-p/4550489</link>
      <description>&lt;P&gt;Bob Ward and I recently released a series of videos, read more in the &lt;A href="https://devblogs.microsoft.com/blog/start-here-azure-sql-foundations-series/" target="_blank"&gt;original blog post&lt;/A&gt; or go directly to the &lt;A href="https://www.youtube.com/watch?v=Gbb4vTh0Lts&amp;amp;list=PLde0EPBCqsc0&amp;amp;index=1" target="_blank"&gt;series on YouTube&lt;/A&gt;.&lt;/P&gt;
&lt;P&gt;The&amp;nbsp;&lt;STRONG&gt;Azure SQL Database Foundations series&lt;/STRONG&gt; are four videos that take you from your first Hyperscale database to AI features running against your own operational data. We also included how to assess and migrate (with AI and skills!) to Hyperscale in the first place, and the common optimizations you should consider. Every episode ships with a repo, so you can follow along in your own environment instead of watching someone else’s terminal.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 26 Aug 2026 12:41:16 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-sql-blog/icymi-new-azure-sql-foundations-video-series-with-github-samples/ba-p/4550489</guid>
      <dc:creator>Anna Hoffman</dc:creator>
      <dc:date>2026-08-26T12:41:16Z</dc:date>
    </item>
    <item>
      <title>Streamlining Enterprise Collaboration - ACP - SharePoint Partner Spotlight</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-sharepoint-blog/streamlining-enterprise-collaboration-acp-sharepoint-partner/ba-p/4550384</link>
      <description>&lt;P&gt;Organizations continue to look for faster ways to modernize intranets, improve internal communications, and increase employee engagement. Partners in the Microsoft ecosystem are helping customers accelerate that journey by extending SharePoint and Microsoft 365 with ready-to-deploy solutions that deliver immediate business value.&lt;/P&gt;
&lt;P&gt;As organizations increasingly adopt Microsoft 365 Copilot, a well-structured and engaging intranet becomes even more important. High-quality content, clear information architecture, and effective content management help employees get greater value from AI-powered experiences.&lt;/P&gt;
&lt;P data-line="5"&gt;This partner showcase highlights&amp;nbsp;&lt;STRONG&gt;ACP&lt;/STRONG&gt;, one of the leading IT service providers in Austria, Germany, and Switzerland.&lt;/P&gt;
&lt;H2 data-line="9"&gt;About ACP&lt;/H2&gt;
&lt;P data-line="11"&gt;&lt;A class="lia-external-url" href="https://www.acp-gruppe.com/en/" target="_blank" rel="noopener"&gt;ACP&lt;/A&gt; is one of the leading&lt;STRONG&gt; IT service providers in Austria, Germany, and Switzerland&lt;/STRONG&gt;. As a leading technology and digital transformation partner, ACP delivers modern IT solutions, cloud services, digital workplaces, cybersecurity, and managed services to businesses and public sector organizations. With more than 2,600 specialists and a strong local presence, ACP helps customers unlock innovation and turn technology into measurable business value.&lt;/P&gt;
&lt;P data-line="13"&gt;Within ACP, the&amp;nbsp;&lt;STRONG&gt;SCOPE team&lt;/STRONG&gt; (SharePoint &amp;amp; Copilot Experts) serves as the &lt;STRONG&gt;AI-first center of excellence for Microsoft 365, SharePoint Online and Copilot&lt;/STRONG&gt;. Through consulting, implementation, training, adoption programs, and product innovations, the team enables organizations to become future-ready and achieve sustainable long-term success. &lt;SPAN style="color: rgb(30, 30, 30);"&gt;With more than 15 years of experience helping organizations build digital workplaces with SharePoint, the SCOPE team has helped organizations build intelligent digital workspaces, intranets, and business solutions that deliver real value - turning customer-specific use cases into practical and impactful solutions.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2 data-line="15"&gt;The ACP Intranet Accelerator Kit&lt;/H2&gt;
&lt;P data-line="17"&gt;To address recurring needs and challenges, ACP transformed years of project experience into a scalable product. In 2024, the company launched the&amp;nbsp;&lt;STRONG&gt;ACP Intranet Accelerator Kit&lt;/STRONG&gt;, a subscription-based add-on for SharePoint Online that combines the most requested intranet capabilities into a ready-to-use solution.&lt;/P&gt;
&lt;img /&gt;
&lt;P data-line="21"&gt;The &lt;A class="lia-external-url" href="https://www.acp-gruppe.com/de-at/modern-workplace/intranet-accelerator-kit" target="_blank" rel="noopener"&gt;ACP Intranet Accelerator Kit&lt;/A&gt; &lt;STRONG&gt;extends SharePoint Online with a comprehensive set of configurable capabilities&lt;/STRONG&gt; designed to enhance the employee experience while reducing implementation effort and cost. Instead of building custom solutions from scratch, organizations can quickly deploy proven functionality that addresses common intranet requirements while staying aligned with Microsoft 365 governance and standards.&lt;/P&gt;
&lt;P data-line="23"&gt;The solution improves several key areas of the intranet experience, including:&lt;/P&gt;
&lt;UL data-line="25"&gt;
&lt;LI&gt;Personalized information delivery through synchronized user profile data&lt;/LI&gt;
&lt;LI&gt;Enhanced navigation and content discoverability&lt;/LI&gt;
&lt;LI&gt;Extended search and filtering capabilities&lt;/LI&gt;
&lt;LI&gt;Flexible design and content presentation options&lt;/LI&gt;
&lt;LI&gt;Ready-to-use intranet components for common employee scenarios&lt;/LI&gt;
&lt;LI&gt;Improved content management experience for editors and site owners&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;These capabilities help organizations accelerate intranet modernization while reducing implementation complexity and cost. Organizations can rapidly introduce capabilities such as people directories, event hubs, FAQs, wikis, targeted news experiences, quick links, polls, and employee profile enhancements — without extensive custom development.&lt;/P&gt;
&lt;P data-line="34"&gt;For&amp;nbsp;&lt;STRONG&gt;employees&lt;/STRONG&gt;, this means faster access to relevant information, better discoverability, and a more engaging digital content experience. For&amp;nbsp;&lt;STRONG&gt;content owners and communication teams&lt;/STRONG&gt;, it means greater flexibility, lower maintenance effort, and the ability to launch valuable intranet experiences significantly faster.&lt;/P&gt;
&lt;P data-line="36"&gt;The ACP Intranet Accelerator Kit reflects ACP's broader innovation approach: transforming real customer challenges into reusable solutions that scale across organizations. By combining Microsoft SharePoint with practical product innovation, &lt;STRONG&gt;ACP helps customers accelerate their digital workplace journey&lt;/STRONG&gt; while continuously benefiting from ongoing enhancements, support, and product evolution.&lt;/P&gt;
&lt;H2 data-line="38"&gt;Watch the showcase&lt;/H2&gt;
&lt;div data-video-id="https://www.youtube.com/watch?v=3HW0qBi7tIE/1787735509819" data-video-remote-vid="https://www.youtube.com/watch?v=3HW0qBi7tIE/1787735509819" class="lia-video-container lia-media-is-center lia-media-size-large"&gt;&lt;iframe src="https://cdn.embedly.com/widgets/media.html?src=https%3A%2F%2Fwww.youtube.com%2Fembed%2F3HW0qBi7tIE%3Ffeature%3Doembed&amp;amp;display_name=YouTube&amp;amp;url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3D3HW0qBi7tIE&amp;amp;image=https%3A%2F%2Fi.ytimg.com%2Fvi%2F3HW0qBi7tIE%2Fhqdefault.jpg&amp;amp;type=text%2Fhtml&amp;amp;schema=youtube" allowfullscreen="" style="max-width: 100%"&gt;&lt;/iframe&gt;&lt;/div&gt;
&lt;P&gt;In this SharePoint Partner Showcase, ACP demonstrates how organizations can rapidly enhance SharePoint Online with proven intranet capabilities that improve communications, employee engagement, and knowledge sharing while reducing implementation effort. By optimizing corporate communications, streamlining document workflows, and boosting employee engagement, ACP helps enterprises maximize their Microsoft 365 investment and drive long-term business value.&lt;/P&gt;
&lt;P data-line="44"&gt;&lt;STRONG&gt;Key highlights covered in this video:&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL data-line="46"&gt;
&lt;LI&gt;&lt;STRONG&gt;Modern Corporate Communications&lt;/STRONG&gt;&amp;nbsp;– How to centralize news, announcements, and leadership messaging across distributed and hybrid teams.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Intranet Optimization&lt;/STRONG&gt;&amp;nbsp;– Building engaging, targeted communication sites and hub portals that eliminate information silos.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Enhanced Productivity &amp;amp; Collaboration&lt;/STRONG&gt;&amp;nbsp;– Streamlining daily workflows, content governance, and knowledge management within Microsoft 365.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Business ROI &amp;amp; Scalability&lt;/STRONG&gt; – Accelerating deployment with proven intranet capabilities while maintaining flexibility for future growth.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="51"&gt;Whether you're looking to upgrade your legacy intranet, improve internal comms engagement, or build a scalable employee portal, see how ACP delivers tailored solutions aligned with your core business objectives.&lt;/P&gt;
&lt;P data-line="53"&gt;&lt;STRONG&gt;Agenda:&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL data-line="54"&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://www.youtube.com/watch?v=3HW0qBi7tIE" target="_blank" rel="noopener"&gt;00:00&lt;/A&gt; – Introduction&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://www.youtube.com/watch?v=3HW0qBi7tIE&amp;amp;t=198s" target="_blank" rel="noopener"&gt;03:18&lt;/A&gt; – Live demo&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://www.youtube.com/watch?v=3HW0qBi7tIE&amp;amp;t=767s" target="_blank" rel="noopener"&gt;12:47&lt;/A&gt; – Discussion&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;When asked &lt;STRONG&gt;where Microsoft could further invest&lt;/STRONG&gt; for partners and customers, ACP highlighted two key areas:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Improved predictability and governance controls for new feature rollouts to support organizational change management.&lt;/LI&gt;
&lt;LI&gt;Greater opportunities to integrate Copilot capabilities directly into custom SharePoint solutions and experiences.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2 data-line="64"&gt;Presenters&lt;/H2&gt;
&lt;P data-line="66"&gt;In this SharePoint Partner Showcase,&amp;nbsp;&lt;STRONG&gt;Vesa Juvonen (Microsoft)&lt;/STRONG&gt;&amp;nbsp;sat down with&amp;nbsp;&lt;STRONG&gt;Julia Kaufmann-Hailu (ACP)&lt;/STRONG&gt;&amp;nbsp;and&amp;nbsp;&lt;STRONG&gt;Martin Ettl (ACP)&lt;/STRONG&gt; to discuss ACP's SCOPE offerings and how the ACP Intranet Accelerator Kit helps organizations turn SharePoint Online into a modern, engaging digital workplace.&lt;/P&gt;
&lt;UL data-line="70"&gt;
&lt;LI&gt;Julia Kaufmann-Hailu (ACP) –&amp;nbsp;&lt;A class="lia-external-url" href="https://www.linkedin.com/in/kaufmannjulia/" target="_blank" rel="noopener"&gt;LinkedIn&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Martin Ettl (ACP) –&amp;nbsp;&lt;A class="lia-external-url" href="https://www.linkedin.com/in/martin-ettl/" target="_blank" rel="noopener"&gt;LinkedIn&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Host: Vesa Juvonen (Microsoft) –&amp;nbsp;&lt;A class="lia-external-url" href="https://www.linkedin.com/in/vesajuvonen/" target="_blank" rel="noopener"&gt;LinkedIn&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559739&amp;quot;:200,&amp;quot;335559740&amp;quot;:259}"&gt;&amp;nbsp;&lt;/SPAN&gt;Learn more about ACP&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://www.acp-gruppe.com/en/" target="_blank" rel="noopener"&gt;ACP Web site&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://www.linkedin.com/company/acp-gruppe/home/" target="_blank" rel="noopener"&gt;ACP Linkedin&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Microsoft resources&lt;/H2&gt;
&lt;P&gt;Microsoft 365 and SharePoint provide powerful out-of-the-box capabilities that &lt;STRONG&gt;can be extended and tailored&amp;nbsp;to your user experience goals using no-code, low-code, or pro-code approaches&lt;/STRONG&gt;.&amp;nbsp;This flexibility enables customers to configure and build unique and powerful experiences at Microsoft 365 for corporate communications, employee experiences, and business processes.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://www.microsoft.com/en-us/microsoft-365/sharepoint" target="_blank" rel="noopener"&gt;Microsoft SharePoint&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://aka.ms/spat25" target="_blank" rel="noopener"&gt;SharePoint 25th anniversary announcements&lt;/A&gt;&amp;nbsp;- new AI features and more&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://support.microsoft.com/en-us/sharepoint" target="_blank" rel="noopener"&gt;SharePoint help &amp;amp; learning&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/sharepoint/intelligent-intranet-overview" target="_blank" rel="noopener"&gt;Intelligent intranet overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://aka.ms/spfx" target="_blank" rel="noopener"&gt;SharePoint Framework&lt;/A&gt;&amp;nbsp;provides a secure, modern way to build custom and AI-enhanced experiences in SharePoint and Microsoft 365.&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/sharepoint/dev/embedded/overview" target="_blank" rel="noopener"&gt;SharePoint Embedded&lt;/A&gt;&amp;nbsp;enables building solutions outside of Microsoft 365 with a custom user interface, but continues storing files in SharePoint.&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://aka.ms/community/youtube" target="_blank" rel="noopener"&gt;Microsoft Community Learning YouTube channel&lt;/A&gt; for videos and guidance&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2 data-line="87"&gt;Showcase your solution&lt;/H2&gt;
&lt;P data-line="89"&gt;Do you have offerings or products built with SharePoint and would be interested in creating a similar showcase video and blog post with Microsoft? Please fill in the following form to get connected with us:&amp;nbsp;&lt;A class="lia-external-url" href="https://aka.ms/sharepoint/partner/showcase" target="_blank" rel="noopener"&gt;https://aka.ms/sharepoint/partner/showcase.&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;The Microsoft 365 ecosystem continues to thrive through partners who innovate, build solutions that solve real business challenges, and share their expertise with customers around the world.&lt;/P&gt;</description>
      <pubDate>Wed, 26 Aug 2026 09:53:02 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-sharepoint-blog/streamlining-enterprise-collaboration-acp-sharepoint-partner/ba-p/4550384</guid>
      <dc:creator>VesaJuvonen</dc:creator>
      <dc:date>2026-08-26T09:53:02Z</dc:date>
    </item>
    <item>
      <title>Windows Update Failures in Japanese-Localized Azure VM SQL2022-WS2022 Images: Cause and Workaround</title>
      <link>https://techcommunity.microsoft.com/t5/sql-server-support-blog/windows-update-failures-in-japanese-localized-azure-vm-sql2022/ba-p/4535801</link>
      <description>&lt;P&gt;This blog post is primarily targeted at customers in the Japan region and is therefore written in Japanese as well. If you would like to read it in English, please use the link &lt;A class="lia-internal-link" href="#community--1-EnglishVersion" target="_blank" rel="noopener" data-lia-auto-title="Jump to English " data-lia-auto-title-active="0"&gt;Jump to English &lt;/A&gt; below to navigate directly to the English section.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;こんにちは、 SQL Server サポート チームです。&lt;/P&gt;
&lt;P&gt;Azure&amp;nbsp;上で SQL Server&amp;nbsp;を含むマーケットプレイスイメージ（SQL2022-WS2022）からデプロイした OS&amp;nbsp;を日本語化した場合、2026&amp;nbsp;年 5&amp;nbsp;月以降の&amp;nbsp;Windows Update&amp;nbsp;などによる OS更新プログラムの適用が 0x800f081f&amp;nbsp;や 0x800f0831 で失敗する事例について解説します。&lt;/P&gt;
&lt;H2&gt;■事象&lt;/H2&gt;
&lt;P&gt;OS更新プログラムの適用が以下のようなメッセージにより失敗します。&lt;/P&gt;
&lt;P&gt;エラー："更新プログラムのインストール中に問題が発生しましたが、後で自動的に再試行されます。この問題が引き続き発生し、 Web&amp;nbsp;検索やサポートへの問い合わせを通じて情報を集める必要がある場合は、次のエラー&amp;nbsp;コードが役立つ可能性があります。: (0x800f081f) "&lt;/P&gt;
&lt;P&gt;※ 表示されるエラー コードは、0x800f0831 の場合もあります。&lt;/P&gt;
&lt;H2&gt;■原因&lt;/H2&gt;
&lt;P&gt;本問題は、マーケットプレイスイメージの&amp;nbsp;SQL Server を含むイメージの問題により発生する場合があります。&lt;/P&gt;
&lt;P&gt;最新のイメージでは問題を修正済みのため、現時点で新規にデプロイしたサーバーに対しては関連しません。&lt;/P&gt;
&lt;H2&gt;■確認方法&lt;/H2&gt;
&lt;P&gt;C:\Windows\Logs\CBS\cbs.log&amp;nbsp;に以下のような amd64_microsoft-hyper-v-ram-parser.resources_31bf3856ad364e35_10.0.20348.1_ja-jp_531076835e0e2f98.manifest が存在しないという警告が出力されている場合、本事象に合致すると判断できます。&lt;/P&gt;
&lt;P&gt;//参考&lt;/P&gt;
&lt;P&gt;C:\Windows\Logs\CBS\cbs.log&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;2026-05-19 12:11:57, Info&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; CBS&amp;nbsp;&amp;nbsp;&amp;nbsp; Exec: Warning - Manifest doesn't exist for: \\?\C:\Windows\CbsTemp\31254331_3251161130\Windows10.0-KB5087545-x64.cab\amd64_microsoft-hyper-v-ram-parser.resources_31bf3856ad364e35_10.0.20348.1_ja-jp_531076835e0e2f98.manifest&amp;nbsp;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;■対応手順&lt;/H2&gt;
&lt;P&gt;＃＃注意事項＃＃&lt;/P&gt;
&lt;P&gt;・事前に Azure Backup で Azure VM のバックアップを取得し、作業前の状態にリストアできる状態にします。&lt;/P&gt;
&lt;P&gt;・当該作業時、 SQL Server サービスは事前に停止状態（スタートアップ種類を無効）にします。作業が完了した後で、起動状態（スタートアップ種類を自動や自動（遅延開始）など元の設定）にします。&lt;/P&gt;
&lt;P&gt;・日本語の言語パッケージの再追加後は更新プログラムの再適用を行うまでは決して OS&amp;nbsp;を再起動しないようにします。&lt;/P&gt;
&lt;P&gt;　言語リソースの追加直後は OS&amp;nbsp;のバイナリと追加された言語リソースとでバージョンが不一致のため、そのまま再起動を行った場合、リモート デスクトップ接続が行えなくなるなどの問題が発生する可能性があります。&lt;/P&gt;
&lt;P&gt;＃＃＃＃＃＃＃＃&lt;/P&gt;
&lt;H3&gt;1.日本語の言語パッケージの削除&lt;/H3&gt;
&lt;P&gt;-----------------------------------&lt;/P&gt;
&lt;P&gt;1-1. [コマンド プロンプト]&amp;nbsp;を [管理者として実行]&amp;nbsp;にて開始します。&lt;/P&gt;
&lt;P&gt;1-2. 以下のコマンドを実行し、日本語の言語パックをアンインストールします。&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;lpksetup /u ja-JP&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;1-3. コンピューターを再起動します。&lt;/P&gt;
&lt;P&gt;1-4. 日本語のユーザー インターフェイスを使用していたアカウントでサインインし、英語で表示されていることを確認します。&lt;/P&gt;
&lt;P&gt;※ lpksetup&amp;nbsp;コマンドをオプションなしで起動すると、GUI&amp;nbsp;からのアンインストールが可能です。&lt;/P&gt;
&lt;P&gt;※ 日本語が優先言語となっている場合、 アンインストール実行後も設定アプリなどに残存する場合がありますが、Preferred languages&amp;nbsp;設定で "Japanese"&amp;nbsp;を下に移動させて [Remove]&amp;nbsp;を押すことで、アンインストールすることが可能です。&lt;/P&gt;
&lt;H3&gt;2.日本語の言語パッケージの再追加&lt;/H3&gt;
&lt;P&gt;-----------------------------------&lt;/P&gt;
&lt;P&gt;日本語の言語パッケージの再追加と更新プログラムの再適用を行います。&lt;/P&gt;
&lt;P&gt;なお、言語パッケージの再追加方法は、Windows Update エンドポイント宛の通信可否によって手順が変わるため、それぞれの手順について記載します。&lt;/P&gt;
&lt;H4&gt;&amp;lt;A.Windows Update エンドポイント宛の通信可能な環境での再追加手順&amp;gt;&lt;/H4&gt;
&lt;P&gt;2-1.弊社カタログ サイトより、現在適用されている更新プログラムと同じバージョンの msu ファイルをダウンロードしておきます。&lt;/P&gt;
&lt;P&gt;　Microsoft Update Catalog&lt;/P&gt;
&lt;P&gt;　 &lt;A href="https://www.catalog.update.microsoft.com/home.aspx" target="_blank" rel="noopener"&gt;https://www.catalog.update.microsoft.com/home.aspx&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;参考）現在適用されている更新プログラムの確認方法&lt;/P&gt;
&lt;P&gt;　[ファイル名を指定して実行] で "winver" を実行して表示される [Windows のバージョン情報] ウインドウにて現在の OS ビルドの番号を確認します。&lt;/P&gt;
&lt;P&gt;　次にWindows Server 2022 の更新履歴から同じビルド番号の更新プログラムを特定します。&lt;/P&gt;
&lt;P&gt;　Windows Server 2022 更新履歴 - Microsoft サポート&lt;/P&gt;
&lt;P&gt;　&lt;A href="https://support.microsoft.com/ja-jp/topic/windows-server-2022-%E6%9B%B4%E6%96%B0%E5%B1%A5%E6%AD%B4-e1caa597-00c5-4ab9-9f3e-8212fe80b2ee" target="_blank" rel="noopener"&gt;https://support.microsoft.com/ja-jp/topic/windows-server-2022-%E6%9B%B4%E6%96%B0%E5%B1%A5%E6%AD%B4-e1caa597-00c5-4ab9-9f3e-8212fe80b2ee&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;2-2. 下記弊社ブログの 「&lt;STRONG&gt;オンラインでインストールが可能な場合&lt;/STRONG&gt;」 の項を参考に日本語の言語パッケージを再追加します。&lt;/P&gt;
&lt;P&gt;　Windows Server 2022&amp;nbsp;の日本語化&lt;/P&gt;
&lt;P&gt;　&lt;A href="https://jpwinsup.github.io/blog/2025/12/14/Performance/LangPack/WindowsServerJapanize/" target="_blank" rel="noopener"&gt;https://jpwinsup.github.io/blog/2025/12/14/Performance/LangPack/WindowsServerJapanize/&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;2-3. 言語パッケージ追加後、[コマンド プロンプト] を [管理者として実行] にて開始します。&lt;/P&gt;
&lt;P&gt;2-4. 以下の一連のコマンドを実行し、作業フォルダー作成の上、事前にダウンロードしておいた msu ファイルの展開 (2 回) を行います。&lt;/P&gt;
&lt;P&gt;2-4-1. 作業フォルダー作成&lt;/P&gt;
&lt;P&gt;構文)&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;md &amp;lt;フォルダー (1 つ目)&amp;gt;&lt;/P&gt;
&lt;P&gt;md &amp;lt;フォルダー (2 つ目)&amp;gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;実行例)&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;md C:\KBTemp\cab&lt;/P&gt;
&lt;P&gt;md C:\KBTemp\cab\cab2&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;2-4-2. 1 回目の展開&lt;/P&gt;
&lt;P&gt;構文)&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;expand -f:Windows*.cab &amp;lt;msu ファイルの場所&amp;gt; &amp;lt;抽出先 (1 つ目のフォルダー)&amp;gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;実行例)&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;expand -f:Windows*.cab C:\windows10.0-kb5082142-x64_0d2677e2b8f6550a978bd89e8991c4d3813ccf1f.msu C:\KBTemp\cab&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;2-4-3. 2 回目の展開&lt;/P&gt;
&lt;P&gt;構文)&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;expand -f:* &amp;lt;1 回目で展開された cab&amp;gt; &amp;lt;抽出先 (2 つ目のフォルダー)&amp;gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;実行例)&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;expand -f:* C:\KBTemp\cab\Windows10.0-KB5082142-x64.cab C:\KBTemp\cab\cab2&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;2-5.以下のコマンドを実行し、展開された cab ファイルを使って更新プログラムの再適用を行います。&lt;/P&gt;
&lt;P&gt;構文)&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;DISM /Online /Add-Package /PackagePath:&amp;lt;2 つ目の作業フォルダーに展開された cab ファイル (※ KB 番号を冠するもの)&amp;gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;実行例)&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;DISM /Online /Add-Package /PackagePath:C:\KBTemp\cab\cab2\Windows10.0-KB5082142-x64.cab&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;2-6. 言語パッケージ追加後、改めて更新プログラムの適用を行い、事象が解消することを確認します。&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H4&gt;&amp;lt;B.Windows Update エンドポイント宛の通信ができない環境での再追加手順&amp;gt;&lt;/H4&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;2-1.弊社カタログ サイトより、現在適用されている更新プログラムと同じバージョンの msu ファイルをダウンロードしておきます。&lt;/P&gt;
&lt;P&gt;　Microsoft Update Catalog&lt;/P&gt;
&lt;P&gt;　 &lt;A href="https://www.catalog.update.microsoft.com/home.aspx" target="_blank" rel="noopener"&gt;https://www.catalog.update.microsoft.com/home.aspx&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;参考）現在適用されている更新プログラムの確認方法&lt;/P&gt;
&lt;P&gt;　[ファイル名を指定して実行] で "winver" を実行して表示される [Windows のバージョン情報] ウインドウにて現在の OS ビルドの番号を確認します。&lt;/P&gt;
&lt;P&gt;　次にWindows Server 2022 の更新履歴から同じビルド番号の更新プログラムを特定します。&lt;/P&gt;
&lt;P&gt;　Windows Server 2022 更新履歴 - Microsoft サポート&lt;/P&gt;
&lt;P&gt;　&lt;A href="https://support.microsoft.com/ja-jp/topic/windows-server-2022-%E6%9B%B4%E6%96%B0%E5%B1%A5%E6%AD%B4-e1caa597-00c5-4ab9-9f3e-8212fe80b2ee" target="_blank" rel="noopener"&gt;https://support.microsoft.com/ja-jp/topic/windows-server-2022-%E6%9B%B4%E6%96%B0%E5%B1%A5%E6%AD%B4-e1caa597-00c5-4ab9-9f3e-8212fe80b2ee&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;2-2.下記弊社ブログの 「&lt;STRONG&gt;オンラインでインストールできない場合&lt;/STRONG&gt;」 の項を参考に日本語の言語パッケージを再追加します。&lt;/P&gt;
&lt;P&gt;　Windows Server 2022&amp;nbsp;の日本語化&lt;/P&gt;
&lt;P&gt;　&lt;A href="https://jpwinsup.github.io/blog/2025/12/14/Performance/LangPack/WindowsServerJapanize/" target="_blank" rel="noopener"&gt;https://jpwinsup.github.io/blog/2025/12/14/Performance/LangPack/WindowsServerJapanize/&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;2-3. 言語パッケージ追加後、[コマンド プロンプト] を [管理者として実行] にて開始します。&lt;/P&gt;
&lt;P&gt;2-4. 以下の一連のコマンドを実行し、作業フォルダー作成の上、事前にダウンロードしておいた msu ファイルの展開 (2 回) を行います。&lt;/P&gt;
&lt;P&gt;2-4-1. 作業フォルダー作成&lt;/P&gt;
&lt;P&gt;構文)&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;md &amp;lt;フォルダー (1 つ目)&amp;gt;&lt;/P&gt;
&lt;P&gt;md &amp;lt;フォルダー (2 つ目)&amp;gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;実行例)&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;md C:\KBTemp\cab&lt;/P&gt;
&lt;P&gt;md C:\KBTemp\cab\cab2&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;2-4-2. 1 回目の展開&lt;/P&gt;
&lt;P&gt;構文)&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;expand -f:Windows*.cab &amp;lt;msu ファイルの場所&amp;gt; &amp;lt;抽出先 (1 つ目のフォルダー)&amp;gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;実行例)&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;expand -f:Windows*.cab C:\windows10.0-kb5082142-x64_0d2677e2b8f6550a978bd89e8991c4d3813ccf1f.msu C:\KBTemp\cab&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;2-4-3. 2 回目の展開&lt;/P&gt;
&lt;P&gt;構文)&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;expand -f:* &amp;lt;1 回目で展開された cab&amp;gt; &amp;lt;抽出先 (2 つ目のフォルダー)&amp;gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;実行例)&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;expand -f:* C:\KBTemp\cab\Windows10.0-KB5082142-x64.cab C:\KBTemp\cab\cab2&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;2-5.以下のコマンドを実行し、展開された cab ファイルを使って更新プログラムの再適用を行います。&lt;/P&gt;
&lt;P&gt;構文)&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;DISM /Online /Add-Package /PackagePath:&amp;lt;2 つ目の作業フォルダーに展開された cab ファイル (※ KB 番号を冠するもの)&amp;gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;実行例)&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;DISM /Online /Add-Package /PackagePath:C:\KBTemp\cab\cab2\Windows10.0-KB5082142-x64.cab&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;2-6. 言語パッケージ追加後、改めて更新プログラムの適用を行い、事象が解消することを確認します。&lt;/P&gt;
&lt;H1 id="eng" class="lia-linked-item"&gt;&lt;a id="community--1-EnglishVersion" class="lia-anchor"&gt;&lt;/a&gt;English &amp;nbsp;&lt;/H1&gt;
&lt;P&gt;Hello, this is the SQL Server Support Team.&lt;/P&gt;
&lt;P&gt;In this article, we explain an issue where operating system updates (such as Windows Updates released in and after May 2026) may fail with error 0x800f081f or 0x800f0831 after localizing an Azure Marketplace image containing SQL Server (for example, SQL2022-WS2022) to Japanese.&lt;/P&gt;
&lt;H2&gt;Issue&lt;/H2&gt;
&lt;P&gt;Installation of Windows updates may fail with a message similar to the following:&lt;/P&gt;
&lt;P&gt;"更新プログラムのインストール中に問題が発生しましたが、後で自動的に再試行されます。この問題が引き続き発生し、 Web&amp;nbsp;検索やサポートへの問い合わせを通じて情報を集める必要がある場合は、次のエラー&amp;nbsp;コードが役立つ可能性があります。: (0x800f081f)."&lt;/P&gt;
&lt;P&gt;Note: In some cases, the reported error code may be 0x800f0831 instead.&lt;/P&gt;
&lt;H2&gt;Cause&lt;/H2&gt;
&lt;P&gt;This issue may occur due to a problem in certain Marketplace images that include SQL Server.&lt;/P&gt;
&lt;P&gt;The issue has already been fixed in the latest Marketplace images. Therefore, it does not affect newly deployed servers created from the updated images.&lt;/P&gt;
&lt;H2&gt;How to Identify the Issue&lt;/H2&gt;
&lt;P&gt;You can determine whether your environment is affected if the following warning appears in C:\Windows\Logs\CBS\CBS.log, indicating that the manifest file below cannot be found:&lt;/P&gt;
&lt;P&gt;amd64_microsoft-hyper-v-ram-parser.resources_31bf3856ad364e35_10.0.20348.1_ja-jp_531076835e0e2f98.manifest&lt;/P&gt;
&lt;P&gt;File: C:\Windows\Logs\CBS\CBS.log&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;2026-05-19 12:11:57, Info&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; CBS&amp;nbsp;&amp;nbsp;&amp;nbsp; Exec: Warning - Manifest doesn't exist for: \\?\C:\Windows\CbsTemp\31254331_3251161130\Windows10.0-KB5087545-x64.cab\amd64_microsoft-hyper-v-ram-parser.resources_31bf3856ad364e35_10.0.20348.1_ja-jp_531076835e0e2f98.manifest&amp;nbsp;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;Resolution&lt;/H2&gt;
&lt;P&gt;&lt;STRONG&gt;Important Notes&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Before proceeding, please ensure the following:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Create a backup of the Azure VM using Azure Backup so that the VM can be restored to its pre-maintenance state if necessary.&lt;/LI&gt;
&lt;LI&gt;Stop the SQL Server service before performing the procedure, and set its startup type to Disabled.&lt;/LI&gt;
&lt;LI&gt;After the procedure is completed, restore the SQL Server service startup type to its original setting (such as Automatic or Automatic (Delayed Start)).&lt;/LI&gt;
&lt;LI&gt;Do not restart the operating system after reinstalling the Japanese language pack until the update has been reapplied successfully.&lt;BR /&gt;Immediately after the language resources are reinstalled, the versions of the OS binaries and language resources may not match. Restarting at this stage can cause issues such as the inability to establish Remote Desktop connections.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;1.Remove the Japanese Language Pack&lt;/H3&gt;
&lt;P&gt;1-1. Launch Command Prompt as Administrator.&lt;/P&gt;
&lt;P&gt;1-2. Run the following command to uninstall the Japanese language pack:&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;lpksetup /u ja-JP&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;1-3. Restart the computer.&lt;/P&gt;
&lt;P&gt;1-4. Sign in using an account that previously used the Japanese user interface and confirm that Windows is now displayed in English.&lt;/P&gt;
&lt;P&gt;Notes&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;If Japanese remains listed because it is configured as the preferred language, move Japanese lower in the Preferred languages list and select Remove.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;2.Reinstall the Japanese Language Pack&lt;/H3&gt;
&lt;P&gt;Reinstall the Japanese language pack and then reapply the Windows update.&lt;/P&gt;
&lt;P&gt;The procedure differs depending on whether the server can communicate with Windows Update endpoints.&lt;/P&gt;
&lt;H4&gt;&amp;lt;Environment with Access to Windows Update Endpoints&amp;gt;&lt;/H4&gt;
&lt;P&gt;2-1. Download the Corresponding MSU File&lt;/P&gt;
&lt;P&gt;Download the .msu file that matches the update version currently installed on the server from the Microsoft Update Catalog.&lt;/P&gt;
&lt;P&gt;Microsoft Update Catalog&lt;/P&gt;
&lt;P&gt;&lt;A href="https://www.catalog.update.microsoft.com/home.aspx" target="_blank" rel="noopener"&gt;https://www.catalog.update.microsoft.com/home.aspx&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;How to Identify the Installed Update Version&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Run winver.&lt;/LI&gt;
&lt;LI&gt;Note the current OS build number shown in the About Windows dialog.&lt;/LI&gt;
&lt;LI&gt;Identify the corresponding update from the Windows Server 2022 update history page.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Windows Server 2022 Update History&lt;/P&gt;
&lt;P&gt;&lt;A href="https://support.microsoft.com/en-us/topic/windows-server-2022-update-history-e1caa597-00c5-4ab9-9f3e-8212fe80b2ee" target="_blank" rel="noopener"&gt;https://support.microsoft.com/en-us/topic/windows-server-2022-update-history-e1caa597-00c5-4ab9-9f3e-8212fe80b2ee&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;2-2. Reinstall the Japanese Language Pack (Online)&lt;/P&gt;
&lt;P&gt;Perform the following steps:&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;2-2-1. Go to&amp;nbsp;&lt;STRONG&gt;Settings &amp;gt; Time &amp;amp; Language &amp;gt; Language &amp;amp; Region&lt;/STRONG&gt;, then click &lt;STRONG&gt;Add a language&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;2-2-2. On the&amp;nbsp;&lt;STRONG&gt;Choose a language to install&lt;/STRONG&gt; screen, select 日本語 (&lt;STRONG&gt;Japanese)&amp;nbsp;&lt;/STRONG&gt; and click &lt;STRONG&gt;Next&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;2-2-3. On the&amp;nbsp;&lt;STRONG&gt;Install language features&lt;/STRONG&gt; screen, select &lt;STRONG&gt;Language pack&lt;/STRONG&gt; and click &lt;STRONG&gt;Install&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;2-2-4. Wait for the installation to complete.&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;2-2-5. Once the installation is complete, sign out and then sign back in. After that, you will be able to select Japanese as your display language under [Preferred languages].&lt;/P&gt;
&lt;P&gt;2-3. Launch Command Prompt as Administrator.&lt;/P&gt;
&lt;P&gt;2-4. Extract the Downloaded MSU File&lt;/P&gt;
&lt;P&gt;2-4-1. Create Working Directories&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;md &amp;lt;Folder1&amp;gt;&lt;/P&gt;
&lt;P&gt;md &amp;lt;Folder2&amp;gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;2-4-2. First Extraction&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;expand -f:Windows*.cab &amp;lt;MSU file&amp;gt; &amp;lt;Folder1&amp;gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;2-4-3. Second Extraction&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;expand -f:* &amp;lt;CAB extracted in previous step&amp;gt; &amp;lt;Folder2&amp;gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;2-5. Reapply the Update&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Use DISM to reinstall the extracted update package.&lt;/P&gt;
&lt;P&gt;DISM /Online /Add-Package /PackagePath:&amp;lt;Folder2&amp;gt;\&amp;lt;KB CAB file&amp;gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;2-6. After the language pack has been reinstalled, install the latest update again and verify that the issue is resolved.&lt;/P&gt;
&lt;H4&gt;&amp;lt;Environment Without Access to Windows Update Endpoints&amp;gt;&lt;/H4&gt;
&lt;P&gt;2-1. Download the Corresponding MSU File&lt;/P&gt;
&lt;P&gt;Download the .msu file matching the currently installed update from the Microsoft Update Catalog.&lt;/P&gt;
&lt;P&gt;Microsoft Update Catalog&lt;/P&gt;
&lt;P&gt;&lt;A href="https://www.catalog.update.microsoft.com/home.aspx" target="_blank" rel="noopener"&gt;https://www.catalog.update.microsoft.com/home.aspx&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;Refer to the procedure described above to identify the correct update package.&lt;/P&gt;
&lt;P&gt;2-2. Reinstall the Japanese Language Pack (Offline)&lt;/P&gt;
&lt;P&gt;Perform the following steps:&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;2-2-1. Refer to the link below and download the required media according to your licensing agreement.&lt;BR /&gt;For Windows Server 2022, you will need the &lt;STRONG&gt;"Languages and Optional Features for Windows Server 2022"&lt;/STRONG&gt; media.&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/microsoft-365/commerce/licenses/download-vl-products?view=o365-worldwide&amp;amp;viewFallbackFrom=o365-wo%25E2%2580%25A6" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/microsoft-365/commerce/licenses/download-vl-products?view=o365-worldwide&amp;amp;viewFallbackFrom=o365-wo%25E2%2580%25A6&lt;/A&gt;&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;2-2-2. Mount the downloaded ISO file to the target virtual machine.&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;2-2-3. Right-click the&amp;nbsp;&lt;STRONG&gt;Start&lt;/STRONG&gt; menu and select &lt;STRONG&gt;Run&lt;/STRONG&gt;. Then enter &lt;STRONG&gt;lpksetup.exe&lt;/STRONG&gt; and press &lt;STRONG&gt;Enter&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;2-2-4. Click&amp;nbsp;&lt;STRONG&gt;Install display languages&lt;/STRONG&gt;. On the next screen, click &lt;STRONG&gt;Browse&lt;/STRONG&gt; and select the drive that was created when you mounted the ISO file in Step.&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;2-2-5. Select&amp;nbsp;&lt;STRONG style="color: rgb(30, 30, 30);"&gt;Microsoft-Windows-Server-Language-Pack_x64_ja-jp&lt;/STRONG&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt; and click &lt;/SPAN&gt;&lt;STRONG style="color: rgb(30, 30, 30);"&gt;OK&lt;/STRONG&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;2-2-6. Verify that&amp;nbsp;&lt;STRONG&gt;Japanese (日本語)&lt;/STRONG&gt; is selected, and then click &lt;STRONG&gt;Next&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;2-2-7. Review the license terms. If you agree, select&amp;nbsp;&lt;STRONG&gt;I accept the license terms&lt;/STRONG&gt;, and then click &lt;STRONG&gt;Next&lt;/STRONG&gt; to begin the installation.&lt;/P&gt;
&lt;P class="lia-indent-padding-left-30px"&gt;2-2-8. When the installation is complete and&amp;nbsp;&lt;STRONG&gt;Completed&lt;/STRONG&gt; is displayed, click &lt;STRONG&gt;Close&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P&gt;2-3. Launch Command Prompt as Administrator.&lt;/P&gt;
&lt;P&gt;2-4. Extract the Downloaded MSU File&lt;/P&gt;
&lt;P&gt;2-4-1. Create Working Directories&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;md &amp;lt;Folder1&amp;gt;&lt;/P&gt;
&lt;P&gt;md &amp;lt;Folder2&amp;gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;2-4-2. First Extraction&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;expand -f:Windows*.cab &amp;lt;MSU file&amp;gt; &amp;lt;Folder1&amp;gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;2-4-3. Second Extraction&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;expand -f:* &amp;lt;CAB extracted in previous step&amp;gt; &amp;lt;Folder2&amp;gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;2-5. Reapply the Update&lt;/P&gt;
&lt;P&gt;Use DISM to reinstall the extracted update package.&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;DISM /Online /Add-Package /PackagePath:&amp;lt;Folder2&amp;gt;\&amp;lt;KB CAB file&amp;gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;2-6. After the language pack has been reinstalled, install the latest update again and verify that the issue is resolved.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Thanks.&lt;/P&gt;</description>
      <pubDate>Wed, 26 Aug 2026 07:45:58 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/sql-server-support-blog/windows-update-failures-in-japanese-localized-azure-vm-sql2022/ba-p/4535801</guid>
      <dc:creator>Yuki_Kitazawa</dc:creator>
      <dc:date>2026-08-26T07:45:58Z</dc:date>
    </item>
    <item>
      <title>Cost Management with Azure Resource Manager MCP</title>
      <link>https://techcommunity.microsoft.com/t5/finops-blog/cost-management-with-azure-resource-manager-mcp/ba-p/4550182</link>
      <description>&lt;P&gt;Today, we’re announcing new Cost Management capabilities in the &lt;A class="lia-internal-link lia-internal-url lia-internal-url-content-type-blog" href="https://techcommunity.microsoft.com/blog/azuregovernanceandmanagementblog/introducing-the-azure-resource-manager-mcp-server/4517521" target="_blank" rel="noopener" data-lia-auto-title="Azure Resource Management (ARM) MCP server" data-lia-auto-title-active="0"&gt;Azure Resource Management (ARM) MCP server&lt;/A&gt;. A core set of cost tools is now available by default, making it easier for AI agents to bring cost context into Azure workflows without additional configuration. For more advanced scenarios, you can enable the optional Cost Management toolset to access an expanded set of capabilities.&lt;/P&gt;
&lt;P&gt;Cloud operations and cloud economics are deeply connected, but the data required to make those decisions is often fragmented across different tools and experiences. The default cost capabilities help AI agents combine Azure resource information with essential cost insights in the workflows where cloud decisions are made. When you need broader coverage, the optional Cost Management toolset extends that experience with additional tools for deeper analysis, planning, and optimization.&lt;/P&gt;
&lt;H2&gt;What You Can Do Today&lt;/H2&gt;
&lt;P&gt;&lt;STRONG&gt;Plan Before You Deploy&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Every cloud project starts with a series of tradeoffs. Before resources are deployed, teams need to understand how architectural decisions, resource choices, and regions will impact cost.&lt;/P&gt;
&lt;P&gt;An AI agent can estimate costs, compare deployment options, and factor in negotiated rates where available so you can evaluate cost alongside technical requirements before deployment.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Examples&amp;nbsp;&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;"Estimate the monthly cost of deploying a three-node AKS cluster with my negotiated pricing.”&lt;/LI&gt;
&lt;LI&gt;“How would my estimated monthly cost change if I updated the VM in this ARM template &amp;lt;ARM Template&amp;gt; from Standard_D8s_v5 to Standard_D2s_v5? and is that within our organization’s policy?”&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;&amp;nbsp;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Understand Costs&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Once workloads are running, cost visibility becomes the priority. &amp;nbsp;An AI agent can query your cost data to explain cost trends, changes in spend, and surface cost drivers.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Example&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;"What were my top cost drivers this month?"&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Manage Cloud Spend&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;With visibility in place, budget and alerts help teams avoid surprises. An AI agent can create and review budgets and check alerts so you can see where spend is tracking against your plans.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Example&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;“How am I doing against my budgets this month?”&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;&amp;nbsp;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Understand and Manage Spend&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;As cloud environments grow, identifying the highest-impact savings opportunities becomes increasingly challenging. &lt;BR /&gt;An AI agent can surface appropriate Advisor cost recommendations so you can prioritize actions to improve savings.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Example&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;"Show me any eligible Savings Plans recommendations that would save me at least $200/month.”&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Analyze AKS Workloads&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;If you’re running AKS workloads on Azure, an AI agent can query AKS cost data by cluster and namespace and compare active and idle capacity.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Example&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;“Show me how much of my AKS spend is idle vs actively used in my current subscription”&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Make Cost-Aware Azure Decisions&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Whether you're evaluating a new deployment, investigating spending changes, tracking budgets, identifying savings opportunities, or analyzing AKS workloads, the Cost Management capabilities in the Azure Resource Management (ARM) MCP server help bring financial context directly into Azure workflows. By combining Azure resource information with cost, pricing, budget, and optimization insights, AI agents can help you make more informed decisions without switching between multiple tools and experiences.&lt;/P&gt;
&lt;H2&gt;What's Included&lt;/H2&gt;
&lt;P&gt;The Azure Resource Management (ARM) MCP server supports the following Cost Management scenarios:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Query – Explore historical spend, understand cost trends&lt;/LI&gt;
&lt;LI&gt;Forecast – Project future costs based on current usage&lt;/LI&gt;
&lt;LI&gt;Budgets and alerts – Create, review, and monitor budgets and alerts to help teams stay ahead of overruns.&lt;/LI&gt;
&lt;LI&gt;Reservations and Savings Plans – Review recommendations and utilization data to identify commitment-based savings opportunities.&lt;/LI&gt;
&lt;LI&gt;Pricing – Retrieve retail pricing and negotiated price sheet details where available.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Default Cost Management capabilities&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;A core set of tools is enabled by default, allowing agents to support common cost analysis and pricing scenarios without additional configuration.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 100%; border-width: 1px;"&gt;&lt;colgroup&gt;&lt;col style="width: 50%" /&gt;&lt;col style="width: 50%" /&gt;&lt;/colgroup&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Capabilities&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;&lt;STRONG&gt;Supported Tools&lt;/STRONG&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Query&lt;/td&gt;&lt;td&gt;
&lt;UL&gt;
&lt;LI&gt;query_costs&lt;/LI&gt;
&lt;LI&gt;query_aks_costs&lt;/LI&gt;
&lt;/UL&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Pricing&lt;/td&gt;&lt;td&gt;
&lt;UL&gt;
&lt;LI&gt;get_retail_prices&lt;/LI&gt;
&lt;LI&gt;start_pricesheet_download&lt;/LI&gt;
&lt;LI&gt;get_pricesheet_status&lt;/LI&gt;
&lt;/UL&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Additional Cost Management capabilities&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;For broader cost-management scenarios, you can enable the optional Cost Management toolset. This extends the default experience with additional tools for dimensions, forecasting, budgets, alerts, reservations, and Savings Plans.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 100%; border-width: 1px;"&gt;&lt;colgroup&gt;&lt;col style="width: 50%" /&gt;&lt;col style="width: 50%" /&gt;&lt;/colgroup&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Capabilities&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;&lt;STRONG&gt;Tools&lt;/STRONG&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Query&lt;/td&gt;&lt;td&gt;
&lt;UL&gt;
&lt;LI&gt;list_dimensions&lt;/LI&gt;
&lt;/UL&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Forecast&lt;/td&gt;&lt;td&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;SPAN style="background-color: rgba(0, 0, 0, 0); color: rgb(30, 30, 30);"&gt;forecast_costs&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Budget &amp;amp; Alerts&lt;/td&gt;&lt;td&gt;
&lt;UL&gt;
&lt;LI&gt;list_budgets&lt;/LI&gt;
&lt;LI&gt;get_budget&lt;/LI&gt;
&lt;LI&gt;create_budget&lt;/LI&gt;
&lt;LI&gt;list_alerts&lt;/LI&gt;
&lt;/UL&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Reservations &amp;amp; Savings Plans&lt;/td&gt;&lt;td&gt;
&lt;UL&gt;
&lt;LI&gt;get_benefit_recommendations&lt;/LI&gt;
&lt;LI&gt;list_benefit_utilization&lt;/LI&gt;
&lt;LI&gt;list_reservation_transactions&lt;/LI&gt;
&lt;/UL&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&amp;nbsp;&lt;/DIV&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;Getting Started&lt;/H2&gt;
&lt;P&gt;&lt;STRONG&gt;Pre-Requisites &lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;VS Code installed&lt;/LI&gt;
&lt;LI&gt;Valid Azure account with appropriate permissions&lt;/LI&gt;
&lt;LI&gt;GitHub Copilot subscription&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Installation&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Install ARM MCP&lt;/STRONG&gt; - &lt;A class="lia-external-url" href="https://aka.ms/JoinARMMCP" target="_blank" rel="noopener"&gt;https://aka.ms/JoinARMMCP&lt;/A&gt;
&lt;UL&gt;
&lt;LI&gt;VS Code launches automatically&lt;/LI&gt;
&lt;LI&gt;Click&amp;nbsp;&lt;STRONG style="color: rgb(30, 30, 30);"&gt;Install&lt;/STRONG&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;&amp;nbsp;under&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG style="color: rgb(30, 30, 30);"&gt;Azure Resource Manager MCP Server&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;Sign in with your Azure credentials&lt;/LI&gt;
&lt;LI&gt;If you hit any authentication issues see the&amp;nbsp;&lt;A style="font-style: normal; font-weight: 400; background-color: rgb(255, 255, 255);" href="https://github.com/Azure/Azure-Resource-Manager-MCP/blob/main/docs/Troubleshooting.md" target="_blank" rel="noopener"&gt;Troubleshooting Guide&lt;/A&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;&amp;nbsp;in the ARM MCP repo&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Add cost header to the Azure Resource Manager MCP configuration file&lt;/STRONG&gt; for access to optional cost tools
&lt;UL&gt;
&lt;LI&gt;Opt-in per client with x-mcp-toolset header: CostManagement&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG style="color: rgb(30, 30, 30);"&gt;For GitHub Copilot Chat (in VS Code)&amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp;&lt;/STRONG&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;Configure mcp configuration file for a specific workspace or globally&lt;/SPAN&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;Option 1 (Workspace): configure vscode/mcp.json file in your repo&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;Option 2 (User): Edit mcp.json file in User Configuration:&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;LI&gt;Open the Command Palette (Ctrl+Shift+P / Cmd+Shift+P)&lt;/LI&gt;
&lt;LI&gt;run&amp;nbsp;&lt;STRONG style="color: rgb(30, 30, 30);"&gt;MCP: Open User Configuration &lt;/STRONG&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;and edit the mcp.json it opens&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;LI-CODE lang=""&gt;{
	"servers": {
		"azure-resource-manager-mcp": {
			"type": "http",
			"url": "https://mcp.management.azure.com",
			"headers": {
				"x-mcp-toolset": "CostManagement"
			}
		}
	}
}&lt;/LI-CODE&gt;
&lt;UL&gt;
&lt;LI style="list-style-type: none;"&gt;
&lt;UL&gt;
&lt;LI style="list-style-type: none;"&gt;
&lt;UL&gt;
&lt;LI&gt;Save the file, then restart the server: open the Command Palette and run&amp;nbsp;&lt;STRONG&gt;MCP: List Servers&lt;/STRONG&gt;, pick&amp;nbsp;azure-resource-manager-mcp, and choose&amp;nbsp;&lt;STRONG&gt;Restart&lt;/STRONG&gt;. The new tools will appear under&amp;nbsp;&lt;STRONG&gt;Configure Tools&lt;/STRONG&gt; in Chat.&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI style="list-style-type: none;"&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;For GitHub Copilot CLI&amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp;&amp;nbsp;&lt;/STRONG&gt;
&lt;UL&gt;
&lt;LI&gt;Edit ~/.copilot/mcp-config.json with the relevant CostManagement header&amp;nbsp;&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;LI-CODE lang=""&gt;{
	"mcpServers": {
		"azure-resource-manager-mcp": {
			"type": "http",
			"url": "https://mcp.management.azure.com",
			"headers": {
				"x-mcp-toolset": "CostManagement"
			}
		}
	}
}&lt;/LI-CODE&gt;
&lt;UL&gt;
&lt;LI style="list-style-type: none;"&gt;
&lt;UL&gt;
&lt;LI style="list-style-type: none;"&gt;
&lt;UL&gt;
&lt;LI&gt;Restart the MCP server to pick up the change. In Copilot CLI today this means exiting and relaunching copilot. Verify the server is running with the /mcp show slash command inside copilot.&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;What's Next?&lt;/H2&gt;
&lt;P&gt;Microsoft Cost Management is actively expanding our agentic cost capabilities. Upcoming capabilities will include additional supported tools and skills.&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;Give Feedback&lt;/H2&gt;
&lt;P&gt;We want to hear from you. Please share your feedback here -&amp;nbsp; &lt;A href="https://aka.ms/azurecost/mcptools/give/feedback" target="_blank" rel="noopener"&gt;https://aka.ms/azurecost/mcptools/give/feedback&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&amp;nbsp;&lt;/STRONG&gt;&lt;/P&gt;
&lt;H2&gt;Resources&lt;/H2&gt;
&lt;P&gt;&lt;STRONG&gt;- &lt;/STRONG&gt;&lt;STRONG&gt;📖&amp;nbsp;&lt;/STRONG&gt;&lt;A class="lia-external-url" href="https://github.com/Azure/Azure-Resource-Manager-MCP/blob/main/docs/CostManagementAndPricingTools.md" target="_blank" rel="noopener"&gt;Documentation&lt;/A&gt;&lt;STRONG&gt;&amp;nbsp; – Complete setup and usage guide&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;- &lt;/STRONG&gt;&lt;STRONG&gt;🔗&amp;nbsp;&lt;/STRONG&gt;&lt;A class="lia-external-url" href="https://aka.ms/azurecost/mcptools/give/feedback" target="_blank" rel="noopener"&gt;Feedback&lt;/A&gt;&lt;STRONG&gt;&amp;nbsp; – Share your feedback on Cost tools in ARM MCP&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;- &lt;/STRONG&gt;&lt;STRONG&gt;🐛&amp;nbsp;&lt;/STRONG&gt;&lt;A class="lia-external-url" href="https://github.com/Azure/Azure-Resource-Manager-MCP/issues" target="_blank" rel="noopener"&gt;Report Issues&lt;/A&gt;&lt;STRONG&gt;&amp;nbsp; - Report bugs with ARM MCP&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;- &lt;/STRONG&gt;&lt;STRONG&gt;🛠️&amp;nbsp;&lt;/STRONG&gt;&lt;A class="lia-external-url" href="https://github.com/Azure/Azure-Resource-Manager-MCP/blob/main/docs/Troubleshooting.md" target="_blank" rel="noopener"&gt;Troubleshooting&lt;/A&gt;&lt;STRONG&gt;&amp;nbsp; – Resolve common issues with ARM MCP&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&amp;nbsp;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 26 Aug 2026 02:17:28 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/finops-blog/cost-management-with-azure-resource-manager-mcp/ba-p/4550182</guid>
      <dc:creator>demiajayi</dc:creator>
      <dc:date>2026-08-26T02:17:28Z</dc:date>
    </item>
    <item>
      <title>Maia 200: Software-defined dataflow and all-Ethernet networking for efficient inference on Azure</title>
      <link>https://techcommunity.microsoft.com/t5/azure-infrastructure-blog/maia-200-software-defined-dataflow-and-all-ethernet-networking/ba-p/4548198</link>
      <description>&lt;P&gt;&lt;EM&gt;By Sherry Xu, Prashant Ranjan, Torsten Hoefler&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;As we enter the era of frontier-scale intelligence, the economics of AI are increasingly determined not by peak compute performance but by how efficiently an entire system can generate tokens at scale. Production inference workloads are rapidly expanding beyond interactive chat toward an array of Copilots, coding assistants, deep reasoning, multimodal experiences, agentic workloads, and synthetic data generation. Azure Maia 200, Microsoft’s second-generation AI accelerator, is our latest custom silicon, engineered from the ground up for efficient inference at cloud scale.&lt;/P&gt;
&lt;P&gt;At approximately &lt;A class="lia-external-url" href="https://blogs.microsoft.com/blog/2026/01/26/maia-200-the-ai-accelerator-built-for-inference/" target="_blank" rel="noopener"&gt;30% better performance per dollar&lt;/A&gt; than the latest-generation GPUs in Microsoft’s fleet, Maia 200 is built on one central principle: that co-optimizing the models, application harnesses, reinforcement learning environments, kernels, communication library, systems, and custom silicon will lead to the best customer outcomes – per token, per watt, and per dollar.&amp;nbsp;&lt;/P&gt;
&lt;P&gt;There is no single specification that can achieve this, and no equation that drives faster math. System-level efficiency emerges from co-optimizing models, software, data placement, communication, and silicon together. At &lt;STRONG&gt;Hot Chips 2026&lt;/STRONG&gt;, we are pleased to share the foundation behind that effort: the Software-defined Local Access architecture embodied in Maia 200.&lt;/P&gt;
&lt;H3&gt;&lt;SPAN class="lia-text-color-21"&gt;&lt;STRONG&gt;Inference is becoming a systems challenge&lt;/STRONG&gt;&lt;/SPAN&gt;&lt;/H3&gt;
&lt;P&gt;Inference is no longer a compute problem alone—it is a systems problem that spans memory, networking, and software. As the range of production inference use cases continues to expand, workloads have also become more diverse and demanding.&lt;/P&gt;
&lt;P&gt;Different phases of inference place different stresses on infrastructure: processing an initial prompt often requires substantial compute, while generating tokens over time can become constrained by memory bandwidth and communication. Emerging model architectures, including Mixture-of-Experts designs, sparse execution techniques, and advanced KV-cache strategies, further increase the importance of efficiently moving and managing data. At the same time, the largest models now span many accelerators, making communication between devices a critical factor in overall performance.&lt;/P&gt;
&lt;P&gt;Meeting these demands requires optimizing for more than peak operations per second. An effective inference platform must deliver high throughput, low latency, reliability, and model quality while minimizing cost and energy per token. Maia 200 was designed with this broader objective in mind, and approaches this challenge using the architectural principles of Software-defined Local Access—both inside the accelerator and scaling to an all-Ethernet architecture that extends explicit memory orchestration across the system.&lt;BR /&gt;&lt;BR /&gt;&lt;/P&gt;
&lt;img /&gt;
&lt;H3&gt;&lt;SPAN class="lia-text-color-21"&gt;&lt;STRONG&gt;Software-defined Local Access: making movement explicit&lt;/STRONG&gt;&lt;/SPAN&gt;&lt;/H3&gt;
&lt;P&gt;Software-defined Local Access, or SDLA, is a dataflow architecture that gives software direct control over how data moves between high-bandwidth memory and localized, highly specialized SRAMs. Rather than relying on implicit cache behavior, SDLA separates control from data movement and gives software direct control over both memory movement and placement using Direct Memory Access (DMA) engines, compute engines, network operations, and synchronization. This approach helps the system achieve more predictable execution characteristics, enabling more consistent performance between runs.&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Maia 200 brings the SDLA principles together in a vertically co-designed system-on-chip. Its control, compute, memory, I/O, and synchronization structures are designed to operate concurrently rather than compete for a single control path. The purpose of the architecture is to keep the compute engines productively fed by placing, reshaping, moving, and synchronizing data efficiently.&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;All-Ethernet networking: extending SDLA beyond the chip&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;As models span accelerators, networking becomes part of the execution engine. Maia 200’s two-tier scale-up network applies an all-Ethernet approach, combining a standards-based physical and switching foundation with AI-specific endpoint hardware, transport, topology, congestion management, reliability mechanisms, and collective software.&lt;/P&gt;
&lt;P&gt;Maia 200 delivers scalable, predictable performance for large AI inference deployments through an advanced two-tier, Ethernet-based scale-up networking architecture. Intelligent load balancing distributes traffic across on-chip resources, while built-in resilience mechanisms quickly adapt to network disruptions. This enables customers to run dense inference workloads more efficiently, maximize cluster utilization, reduce stranded capacity, and lower overall infrastructure costs.&lt;/P&gt;
&lt;P&gt;We have contributed this direction to the&amp;nbsp;&lt;SPAN class="lia-text-color-10"&gt;&lt;A class="lia-external-url" href="https://ultraethernet.org/" target="_blank" rel="noopener"&gt;Ultra Ethernet Consortium (UEC)’s AI base transport profile&lt;/A&gt;&lt;/SPAN&gt;, helping establish Ethernet as a common foundation for networking across the industry.&lt;BR /&gt;&lt;BR /&gt;&lt;/P&gt;
&lt;img&gt;
&lt;P&gt;&lt;SPAN class="lia-text-color-21"&gt;&lt;STRONG&gt;Figure 1. &lt;/STRONG&gt;&lt;/SPAN&gt;Maia 200 uses a unified Ethernet hierarchy based on a variant of the HammingMesh topology: four accelerators form a directly connected quad, 48 accelerators form a rack-scale domain, and a two-tier switched network can scale to 6,144 accelerators.&lt;/P&gt;
&lt;/img&gt;
&lt;H3&gt;&lt;STRONG&gt;Sustaining performance across real AI workloads&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;The value of the SDLA architecture appears in sustained kernel performance rather than peak specifications alone. To evaluate Maia 200 under realistic conditions, we measured performance across 6,143 matrix-multiplication shapes representative of production inference workloads. Across these scenarios, Maia 200 kept compute engines highly utilized, minimizing idle time by coordinating computation, memory, and communication as a single workload. Even as AI workloads shift between heavy computation (e.g., prompt processing), memory access (e.g., token generation), and communication (e.g., collective operations such as Allgather), Maia 200 maintains consistently predictable performance—treating the accelerator, memory, network, and software as one coordinated system.&lt;/P&gt;
&lt;img&gt;
&lt;P&gt;&lt;SPAN class="lia-text-color-21"&gt;&lt;STRONG&gt;Figure 2. &lt;/STRONG&gt;&lt;/SPAN&gt;FP8 matrix-multiplication performance across the same set of inference-relevant shapes, showing high utilization across compute- and memory-bound operating points.&lt;/P&gt;
&lt;/img&gt;&lt;img&gt;
&lt;P&gt;&lt;SPAN class="lia-text-color-21"&gt;&lt;STRONG&gt;Figure 3. &lt;/STRONG&gt;&lt;/SPAN&gt;Allgather performance on eight Maia 200 accelerators. Direct exchange minimizes latency for small transfers; ring exchange approaches the network bandwidth limit for larger transfers.&lt;/P&gt;
&lt;/img&gt;
&lt;H3&gt;&lt;SPAN class="lia-text-color-21"&gt;&lt;STRONG&gt;From silicon to useful AI systems&lt;/STRONG&gt;&lt;/SPAN&gt;&lt;/H3&gt;
&lt;P&gt;Maia 200 demonstrates Microsoft’s system co-design philosophy for model performance and efficiency. &lt;SPAN data-teams="true"&gt;Maia 200 delivers over &lt;A class="lia-external-url" href="https://microsoft.ai/pdf/mai-thinking-1.pdf" target="_blank" rel="noopener"&gt;40% higher token generation&lt;/A&gt; under the same rack power budget throughput when running the MAI-Thinking-1 model than other leading accelerators in the Azure fleet&lt;/SPAN&gt;. This efficiency is only possible by co-designing models, application frameworks, reinforcement-learning environments, kernels, communication libraries, systems software, and custom silicon as a unified stack. At the same time, Maia 200 maintains the flexibility to support a broad ecosystem of open models, enabling customers to optimize for both performance and choice.&lt;/P&gt;
&lt;P&gt;As Microsoft continues along a path of co-designing models and infrastructure, Maia 200 provides a foundation for improving quality, performance, cost, and energy efficiency for agentic inference workloads. We look forward to advancing these SDLA architectural principles in future generations of Maia system, and invite you to see the full details and data on Maia 200 operations in our &lt;A class="lia-external-url" href="https://arxiv.org/abs/2608.24664" target="_blank"&gt;latest paper on arXiv:&lt;/A&gt;&lt;A href="https://arxiv.org/abs/2608.24664" target="_blank"&gt;[2608.24664] Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration&lt;/A&gt;.&lt;/P&gt;
&lt;P&gt;To learn more, explore our blogs on Maia 200:&amp;nbsp;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://blogs.microsoft.com/blog/2026/01/26/maia-200-the-ai-accelerator-built-for-inference/" target="_blank" rel="noopener"&gt;Maia 200: The AI accelerator built for inference&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://techcommunity.microsoft.com/blog/azureinfrastructureblog/deep-dive-into-the-maia-200-architecture/4489312" target="_blank" rel="noopener"&gt;Deep dive into the Maia 200 architecture&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;</description>
      <pubDate>Wed, 26 Aug 2026 14:40:52 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-infrastructure-blog/maia-200-software-defined-dataflow-and-all-ethernet-networking/ba-p/4548198</guid>
      <dc:creator>SherryX19</dc:creator>
      <dc:date>2026-08-26T14:40:52Z</dc:date>
    </item>
    <item>
      <title>Instant revocation of service principal bearer tokens with CAE</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-entra-blog/instant-revocation-of-service-principal-bearer-tokens-with-cae/ba-p/4548192</link>
      <description>&lt;H4&gt;&lt;STRONG&gt;Introduction&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P&gt;Workload identities such as service principals are an important part of the attack surface in Entra: they are often used by critical automated processes, and they can carry privileged and sensitive permissions. The authentication flow used by these principals is the Client Credential flow, which results in an access token.&lt;/P&gt;
&lt;P&gt;In this blog post, I'll describe how Continuous Access Evaluation (CAE) can be leveraged to instantly invalidate an access token for a service principal in Microsoft Entra. This is incredibly valuable as part of a kill-switch in incident response scenarios.&lt;/P&gt;
&lt;H4&gt;&lt;STRONG&gt;Access tokens (aka bearer tokens)&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P&gt;Access tokens are &lt;STRONG&gt;bearer tokens&lt;/STRONG&gt;. Instant invalidation of bearer tokens is an important response capability because &lt;EM&gt;normally &lt;/EM&gt;these tokens have a lifetime between 60 and 90 minutes.&lt;/P&gt;
&lt;P&gt;This is a long time in scenarios where a compromised token can be used for triggering malicious automation by a bad actor.&lt;/P&gt;
&lt;H4&gt;&lt;STRONG&gt;How to enable CAE for service principal tokens&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P&gt;CAE is enabled by a “client capability” claim, “xms_cc”, with the value of “cp1”.&lt;/P&gt;
&lt;P&gt;The screenshot below shows an example of a PowerShell function that incorporates this:&lt;/P&gt;
&lt;img /&gt;
&lt;H4&gt;&lt;STRONG&gt;Recognizing a CAE-enabled token&amp;nbsp;&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P&gt;The access token I get back from this request is base 64 encoded. I can echo this base 54 encoded value back in my PowerShell session, then paste it into a decoder to investigate it (jwt.io or jwt.ms).&lt;/P&gt;
&lt;P&gt;The first thing&amp;nbsp;you will notice is the presence of this claim:&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;The second thing you will notice is the lifetime of the token:&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;The expiry value is now (roughly) 24 hours after the “iat” (issued at) time, instead of the 60-90 minutes that we normally see in tokens that are not CAE-enabled.&lt;/P&gt;
&lt;H4&gt;&lt;STRONG&gt;Why does this matter&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P&gt;Although a longer lifetime seems less secure, Microsoft’s philosophy around this is that with Workload Identity Protection enabled, risk levels can dynamically be picked up for the identity.&lt;/P&gt;
&lt;P&gt;If an identity is at a &lt;STRONG&gt;high &lt;/STRONG&gt;risk level, the token is revoked within a couple of minutes &lt;STRONG&gt;automatically.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;High risk is only one of three revocation events. At the time of writing, the revocation events are:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Service Principal at high risk&lt;/LI&gt;
&lt;LI&gt;Service Principal disabled&lt;/LI&gt;
&lt;LI&gt;Service Principal deleted&lt;/LI&gt;
&lt;/UL&gt;
&lt;H4&gt;&lt;STRONG&gt;How to manually test this&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P&gt;In PowerShell, I use the “Connect-And-GetToken" function above to obtain the token and construct a header, which I then use to call a Graph endpoint:&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;(&lt;EM&gt;I can call this &lt;A href="https://graph.microsoft.com/v1.0/users" target="_blank" rel="noopener"&gt;https://graph.microsoft.com/v1.0/users&lt;/A&gt; endpoint, because my app, represented by $cid, has the &lt;SPAN class="lia-text-color-9"&gt;User.Read.All &lt;/SPAN&gt;API permission. This is not relevant for this topic, just an example for demo purposes).&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;Whilst the bearer token (in the header) is still valid, the response that comes back can be something like this:&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;If this were &lt;U&gt;not&lt;/U&gt; a CAE-enabled token, even removing the service principal, changing its password or disabling it, would NOT stop me from doing this, &lt;STRONG&gt;&lt;U&gt;as long as the bearer token lifetime has not expired!&lt;/U&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Now let’s trigger revocation manually.&lt;/P&gt;
&lt;P&gt;I will describe two of the three revocation events: high risk and disablement.&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;Test 1:&amp;nbsp; Simulate service principal at high risk&lt;/STRONG&gt;&lt;/H5&gt;
&lt;P&gt;To simulate putting the service principal at risk level high, we use &lt;EM&gt;another&lt;/EM&gt; service principal that has the API permission &lt;SPAN class="lia-text-color-9"&gt;“IdentityRiskyServicePrincipal.ReadWrite.All”&lt;/SPAN&gt;.&lt;/P&gt;
&lt;P&gt;With this other identity (connect-MgGraph), issue this command:&lt;/P&gt;
&lt;P&gt;&lt;SPAN class="lia-text-color-15"&gt;Confirm-MgRiskyServicePrincipalCompromised -ServicePrincipalIds @("{Object id of our target service principal}”)&amp;nbsp;&amp;nbsp;&amp;nbsp; &lt;/SPAN&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;After a short while, the identity shows as high risk in the portal (under Identity Protection &amp;gt; Risk workload identities).&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;Now, when I try to use the&amp;nbsp;&lt;STRONG&gt;same&lt;/STRONG&gt; header with the &lt;STRONG&gt;same&lt;/STRONG&gt; access token from before, the response I get is:&lt;/P&gt;
&lt;P&gt;&lt;SPAN class="lia-text-color-8"&gt;“Invoke-RestMethod : The remote server returned an error: (401) Unauthorized.”&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;This means that the access token was effectively revoked and the revocation objective was achieved.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;To dismiss (cancel) the risk status on that service principal, click on “Dismiss service principal(s) risk” in the portal (top bar), or use &lt;SPAN class="lia-text-color-15"&gt;Invoke-MgDismissRiskyServicePrincipal&lt;/SPAN&gt; in PowerShell.&lt;/P&gt;
&lt;P&gt;After this risk dismissal, my token and the header containing it are still invalid, so I have to re-authenticate (in my case with the “Connect-And-GetToken” function) if I again want to make the call to Graph.&lt;/P&gt;
&lt;P&gt;‼️This brings me to&amp;nbsp;&lt;U&gt;an important point&lt;/U&gt;: this option &lt;STRONG&gt;must be&lt;/STRONG&gt; paired with Conditional Access for Workload Identities, using a policy that BLOCKS authentication in case the risk level is high.‼️&lt;/P&gt;
&lt;P&gt;Without this, my token may be revoked by the risk event, but a bad actor that also owns the client ID and secret could, in theory, re-authenticate and get a new token that subsequently does not get revoked as the revocation event on the previous token is &lt;U&gt;in the past &lt;/U&gt;(unless a new risk event is picked up).&lt;/P&gt;
&lt;P&gt;Conditional Access for workload identities requires a Workload Identity Premium license.&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;Test 2: Simulate service principal disabled&lt;/STRONG&gt;&lt;/H5&gt;
&lt;P&gt;A simpler and faster option for revoking a CAE-enabled service principal access token is to disable the service principal.&lt;/P&gt;
&lt;P&gt;For this, I use the Graph Explorer in my lab:&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;‼️Notice that this must be done directly on the service principal. Deactivating the registered app also disables the service principal, but it does not trigger the revocation event.‼️&lt;/P&gt;
&lt;P&gt;This instantly gives me a “&lt;SPAN class="lia-text-color-8"&gt;&lt;STRONG&gt;(401) Unauthorized&lt;/STRONG&gt;”&lt;/SPAN&gt; if I try to reuse the header (with the token), and therefore it is &lt;STRONG&gt;a better manual kill switch option&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P&gt;I was able to do this in a tenant without a Workload Identity Premium license.&lt;/P&gt;
&lt;H4&gt;&lt;STRONG&gt;Microsoft Entra Sign-in logs&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P&gt;Service principals have their own section in the Microsoft Entra Sign-in log.&lt;/P&gt;
&lt;P&gt;If the token issued (after the sign-in was successful) was CAE enabled, the “Basic in” tab of the sign-in event will show a line "Continuous access evaluation Yes",&lt;/P&gt;
&lt;P&gt;Under “Additional details”, you will see:&lt;/P&gt;
&lt;img /&gt;
&lt;H4&gt;&lt;STRONG&gt;Warning&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P&gt;The risk level may be reaching "high" without your input, which leads to automatic revocation for CAE-enabled tokens.&lt;/P&gt;
&lt;P&gt;Detections at the time of writing include:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Microsoft Entra threat intelligence&lt;/LI&gt;
&lt;LI&gt;Suspicious Sign-ins&lt;/LI&gt;
&lt;LI&gt;Admin confirmed service principal compromised &lt;EM&gt;(as we have done in this blog)&lt;/EM&gt;&lt;/LI&gt;
&lt;LI&gt;Leaked Credentials&lt;/LI&gt;
&lt;LI&gt;Malicious application&lt;/LI&gt;
&lt;LI&gt;Suspicious application&lt;/LI&gt;
&lt;LI&gt;Anomalous service principal activity&lt;/LI&gt;
&lt;LI&gt;Suspicious API Traffic&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;A href="https://learn.microsoft.com/en-us/entra/id-protection/concept-workload-identity-risk" target="_blank" rel="noopener"&gt;Securing workload identities with Microsoft Entra ID Protection - Microsoft Entra ID Protection | Microsoft Learn&lt;/A&gt;&lt;/P&gt;
&lt;H4&gt;&lt;STRONG&gt;Conclusion&lt;/STRONG&gt;&lt;/H4&gt;
&lt;P&gt;To create the ability to revoke access tokens before the end of their lifetime,&amp;nbsp; enable the "cp1" client capability claim at integration time (when building your solutions and their authentication call).&amp;nbsp;&lt;/P&gt;
&lt;P&gt;You cannot revoke service principal access tokens that were not created with the client capabilities claim of “cp1” in the first place.&lt;/P&gt;
&lt;P&gt;Please refer to&amp;nbsp;&lt;A href="https://learn.microsoft.com/en-us/entra/identity/conditional-access/concept-continuous-access-evaluation-workload" target="_blank" rel="noopener"&gt;Continuous access evaluation for workload identities in Microsoft Entra ID - Microsoft Entra ID | Microsoft Learn&lt;/A&gt; for an up-to-date list of possibilities and limitations for this feature.&lt;/P&gt;
&lt;P&gt;I am currently researching if this can be used in Agent flows, where the token exchange pattern applies, so stay tuned.&lt;/P&gt;
&lt;H5&gt;&lt;STRONG&gt;Learn more&lt;/STRONG&gt;&lt;/H5&gt;
&lt;P&gt;&lt;A href="https://learn.microsoft.com/en-us/entra/identity-platform/access-token-claims-reference" target="_blank" rel="noopener"&gt;Access token claims reference - Microsoft identity platform | Microsoft Learn&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://learn.microsoft.com/en-us/entra/identity/conditional-access/workload-identity" target="_blank" rel="noopener"&gt;Microsoft Entra Conditional Access for workload identities - Microsoft Entra ID | Microsoft Learn&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://learn.microsoft.com/en-us/entra/identity/conditional-access/concept-continuous-access-evaluation-workload" target="_blank" rel="noopener"&gt;Continuous access evaluation for workload identities in Microsoft Entra ID - Microsoft Entra ID | Microsoft Learn&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://developer.microsoft.com/en-us/graph/graph-explorer" target="_blank" rel="noopener"&gt;Graph Explorer | Try Microsoft Graph APIs - Microsoft Graph&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://learn.microsoft.com/en-us/entra/id-protection/concept-workload-identity-risk" target="_blank" rel="noopener"&gt;Securing workload identities with Microsoft Entra ID Protection - Microsoft Entra ID Protection | Microsoft Learn&lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Wed, 26 Aug 2026 14:48:29 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-entra-blog/instant-revocation-of-service-principal-bearer-tokens-with-cae/ba-p/4548192</guid>
      <dc:creator>Simone_Oor</dc:creator>
      <dc:date>2026-08-26T14:48:29Z</dc:date>
    </item>
    <item>
      <title>Retirement of Microsoft HPC Pack</title>
      <link>https://techcommunity.microsoft.com/t5/azure-high-performance-computing/retirement-of-microsoft-hpc-pack/ba-p/4550183</link>
      <description>&lt;H3&gt;&lt;STRONG&gt;Overview&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;Microsoft HPC Pack is a Windows‑based high‑performance computing (HPC) scheduler that enables customers to deploy, manage, and operate HPC workloads in on‑premises and hybrid environments. Since its original release, HPC Pack has supported a range of Windows‑centric HPC scenarios, including job scheduling, workload orchestration, and integration with Windows Server–based technologies.&lt;/P&gt;
&lt;P&gt;Microsoft is announcing the planned retirement of HPC Pack. For customers looking to run HPC workloads going forward, Microsoft’s supported platform for managed HPC and parallel workloads is Azure Batch. Customers may also adopt any other Azure service as appropriate for their workload. These platforms provide modern, cloud‑native capabilities for scheduling, scaling, and operating HPC workloads on Azure.&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;Important Dates&lt;/STRONG&gt;&lt;/H3&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Retirement announcement: &lt;/STRONG&gt;&lt;STRONG&gt;August 27&lt;/STRONG&gt;&lt;STRONG&gt;, 2026&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;End of Support / retirement date: &lt;/STRONG&gt;&lt;STRONG&gt;August&lt;/STRONG&gt;&lt;STRONG&gt; &lt;/STRONG&gt;&lt;STRONG&gt;27&lt;/STRONG&gt;&lt;STRONG&gt;, 2027&lt;/STRONG&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;After August 27, 2027, HPC Pack will no longer receive feature updates, bug fixes, or standard product support. No new versions or enhancements will be released. Existing HPC Pack deployments will not be forcibly disabled; however, they will be considered unsupported.&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;Support During the Retirement Period&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;HPC Pack is now entering a one-year retirement period that ends at End of Support on August 27, 2027. Microsoft will provide limited, retirement-only support during this period. The scope of this support is defined below so customers know what to expect.&lt;/P&gt;
&lt;P&gt;What Microsoft will provide during the retirement period:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Security response, including investigation of applicable security vulnerabilities (CVEs) and any security updates deemed necessary by Microsoft&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Guidance based on existing Microsoft documentation, published best practices, and previously validated HPC Pack configurations&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Assistance with support incidents involving supported HPC Pack components and documented product functionality&lt;/STRONG&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;What is NOT included during the retirement period:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Non-security bug fixes&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Performance, reliability, scalability, or optimization improvements&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;New features or feature enhancements&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Customer-specific hotfixes, custom code changes, or design modifications&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Validation or certification of new operating systems, hardware platforms, drivers, firmware, third-party software, or dependencies&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Validation of new deployment architectures, configurations, or integration scenarios&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Changes to existing product behavior or design&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Creation of new documentation, guidance, or troubleshooting content&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Troubleshooting that requires new product investigation beyond existing product knowledge, documentation, or previously validated scenarios&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;New customer onboarding, solution design, or pre-sales assistance&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Migration planning, migration execution, or migration consulting services&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Support for configurations, integrations, or deployment scenarios that are outside published documentation, established best practices, or previously validated HPC Pack environments&lt;/STRONG&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;After End of Support (August 27, 2027): HPC Pack will receive no further updates, fixes, or technical support, and deployments will be considered unsupported.&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;What Customers Should Do&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;Customers currently using HPC Pack should begin planning migration to a supported alternative as soon as possible. For most scenarios, the following services are recommended:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure Batch&lt;/STRONG&gt; – for managed scheduling and execution of parallel and HPC workloads&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Or any other Azure service as appropriate for your workload&lt;/STRONG&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;&lt;STRONG&gt;Resources&lt;/STRONG&gt;&lt;/H3&gt;
&lt;UL&gt;
&lt;LI&gt;HPC Pack documentation: &lt;A href="https://learn.microsoft.com/powershell/high-performance-computing/overview?view=hpc19-ps" target="_blank"&gt;Microsoft HPC Pack 2019&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Public issue tracker: &lt;A href="https://github.com/azure/hpcpack" target="_blank"&gt;Azure/hpcpack GitHub repository&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;&lt;STRONG&gt;Next Steps&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;Customers who anticipate needing additional planning assistance or migration guidance should engage early with their Microsoft account team or Cloud Solution Architect (CSA) to help ensure a smooth transition and avoid potential disruption. &amp;nbsp;To help us understand customer needs and improve future communications, please complete this Microsoft &lt;A class="lia-external-url" href="https://forms.cloud.microsoft/r/BJZvWHEpg1" target="_blank" rel="noopener"&gt;Form&lt;/A&gt;. Share any technical, operational, or business challenges you anticipate during your transition. Your feedback will help us identify common concerns and refine future guidance.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 26 Aug 2026 16:54:57 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-high-performance-computing/retirement-of-microsoft-hpc-pack/ba-p/4550183</guid>
      <dc:creator>XinXin</dc:creator>
      <dc:date>2026-08-26T16:54:57Z</dc:date>
    </item>
    <item>
      <title>Tracking Batch node state and duration in Log Analytics</title>
      <link>https://techcommunity.microsoft.com/t5/azure-paas-blog/tracking-batch-node-state-and-duration-in-log-analytics/ba-p/4547582</link>
      <description>&lt;H1&gt;&lt;STRONG&gt;1. The problem (a real customer scenario)&lt;/STRONG&gt;&lt;/H1&gt;
&lt;P&gt;Some of our customers described that they have those underlying gaps:&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 929px; height: 258px; border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Symptom&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;What the customer wanted&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Nodes intermittently stayed in the &lt;EM&gt;leavingpool&lt;/EM&gt; state for hours. The customer needed to know &lt;U&gt;which node&lt;/U&gt; was &lt;U&gt;stuck &lt;/U&gt;and for &lt;U&gt;how long&lt;/U&gt;, so an alert could fire only after a threshold (e.g. &amp;gt; 60 minutes).&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;A dashboard/query that isolates one node at a time and calculates its duration in the current state.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;A node was reported unusable because the underlying OS provisioning failed. The customer needed automated detection to remediate (drop the pool to 0 and back up) instead of catching it manually.&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;A per-node feed that carries state + error code + IP, so an alert can identify the offending node and its error.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;In both scenarios we recommended the same workaround which was: Build a small app that polls the Batch REST API for (per-node data) and stores it.&lt;/P&gt;
&lt;P&gt;In this Article we will build this implementation as a reusable component and show the query and alert that would solve similar problems.&lt;/P&gt;
&lt;H1&gt;&lt;STRONG&gt;2. Why the platform cannot show this directly&lt;/STRONG&gt;&lt;/H1&gt;
&lt;P&gt;&lt;SPAN data-path-to-node="2,0"&gt;- The main issue came from the fact that Azure Monitor metrics tools don't provide granular, node-level tracking out of the box (no 'NodeId' or 'PoolId'). Instead, it has a rich metrics that can help monitoring the Batch accounts (see the&amp;nbsp;&lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/azure-monitor/reference/supported-metrics/microsoft-batch-batchaccounts-metrics" target="_blank" rel="noopener"&gt;full metrics&lt;/A&gt;).&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;- All those metrics are single integers for the whole account. The "Split-by" filter in the portal Metrics blade lists no options for these two metrics so the customers can't filter down into a specific node's status.&lt;/P&gt;
&lt;P&gt;- While AzureDiagnostic logs successfully show events at the pool level (example: '&lt;EM&gt;PoolResizeCompleteEvent&lt;/EM&gt;' or '&lt;EM&gt;PoolResizeCompleteEvent&lt;/EM&gt;'), they don't show state-change events for individual nodes. The current schema simply doesn't contain a '&lt;EM&gt;NodeStateChange&lt;/EM&gt;' row, meaning old KQL queries looking for this event will return zero results.&lt;/P&gt;
&lt;P&gt;In summary, using the current Azure monitor metrics can tell you that "3 nodes are leaving," which answers your "&lt;STRONG&gt;What&lt;/STRONG&gt;" question.&lt;/P&gt;
&lt;P&gt;But if you are asking "&lt;STRONG&gt;Who&lt;/STRONG&gt;" are leaving, you won't get a direct answer.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;The good news that this trivial question will be answered in that article ✌️&lt;/P&gt;
&lt;H1&gt;&lt;STRONG&gt;3. The solution&lt;/STRONG&gt;&lt;/H1&gt;
&lt;H4&gt;&lt;STRONG&gt;The solution is mainly divided in 2 parts:&lt;/STRONG&gt;&lt;/H4&gt;
&lt;H6&gt;&lt;STRONG&gt;1- What we can expose from Azure and build on top of:&lt;/STRONG&gt;&lt;/H6&gt;
&lt;P&gt;As mentioned before, the Azure Batch REST API returns rich per-node data via GET /pools/{poolId}/nodes (check the &lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/rest/api/batchservice/nodes/get-node?view=rest-batchservice-2025-06-01&amp;amp;tabs=HTTP#:~:text=tvmps_5d8adec89961dcc011329b38df999a841f6cc815a5710678b741f04b33556ed2_d%3Fapi%2Dversion%3D2025%2D06%2D01-,Sample%20response,-Status%20code%3A" target="_blank" rel="noopener"&gt;sample response&lt;/A&gt;). The most relevant fields are: (&lt;EM&gt;state, schedulingState, stateTransitionTime, errors[], startTaskInfo, ipAddress, &lt;SPAN style="color: rgb(30, 30, 30);"&gt;allocationTime, &lt;/SPAN&gt;lastBootTime&lt;/EM&gt;) that we will use them in the collector (2nd part of the solution).&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;2- A small polling collector&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;This is a python script runs on a schedule.&lt;/P&gt;
&lt;P&gt;1- It calls&amp;nbsp;&lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/rest/api/batchservice/pools/list-pools?view=rest-batchservice-2025-06-01&amp;amp;tabs=HTTP" target="_blank" rel="noopener"&gt;List pools&lt;/A&gt; + &lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/rest/api/batchservice/nodes/list-nodes?view=rest-batchservice-2025-06-01&amp;amp;tabs=HTTP" target="_blank" rel="noopener"&gt;List Nodes&lt;/A&gt;, and posts each snapshot as a row in the Log Analytics custom table "&lt;STRONG&gt;BatchNodeInventory_CL&lt;/STRONG&gt;" which is retained by Log Analytics Workspace "&lt;STRONG&gt;LAW&lt;/STRONG&gt;" standard retention.&lt;/P&gt;
&lt;P&gt;2- It authenticates against the Batch service. In a production environment, this is securely handled using a Managed Identity (MI), meaning no hardcoded passwords or keys are required.&lt;/P&gt;
&lt;P&gt;But in my repro, it simply uses a token generated via the CLI command:&lt;/P&gt;
&lt;LI-CODE lang="powershell"&gt;az account get-access-token&lt;/LI-CODE&gt;
&lt;P&gt;The core loop is:&lt;/P&gt;
&lt;LI-CODE lang="python"&gt;token = aad_token("https://batch.core.windows.net/") for pool in GET /pools: for node in GET /pools/{pool.id}/nodes: row = { TimeGenerated, PoolId, NodeId, State, SchedulingState, StateTransitionTime, VmSize, IpAddress, IsDedicated, RunningTasks, TotalTasksRun, ErrorCount, ErrorCode, ErrorMessage, StartTaskState, StartTaskExitCode, AllocationTime, LastBootTime } rows.append(row) post_to_law(rows)&lt;/LI-CODE&gt;
&lt;P&gt;You can find the full script here:&amp;nbsp;&lt;A class="lia-external-url" href="https://github.com/AhmedKhaledAbdalla/azure-batch-node-state-tracker/blob/main/node-state-collector.py" target="_blank" rel="noopener"&gt;node-state-collector.py&lt;/A&gt;&lt;/P&gt;
&lt;img&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;/img&gt;
&lt;H1&gt;&lt;STRONG&gt;5. The query that answers the customer's question&lt;/STRONG&gt;&lt;/H1&gt;
&lt;P&gt;Now, our customers' questions became just a query 😉&lt;/P&gt;
&lt;P&gt;In my repro, I kept the collector running for few minutes before I start to query:&lt;/P&gt;
&lt;LI-CODE lang="kusto"&gt;BatchNodeInventory_CL
| where TimeGenerated &amp;gt; ago(1h)
| summarize LastSnapshot=max(TimeGenerated),
            State=any(State_s), Pool=any(PoolId_s), IP=any(IPAddress),
            Transition=any(StateTransitionTime_t),
            ErrCode=any(ErrorCode_s)
      by NodeId=NodeId_s
| extend MinutesInState = datetime_diff("minute", LastSnapshot, Transition)
| project Pool, NodeId, State, IP, MinutesInState, ErrorCode
| order by MinutesInState desc&lt;/LI-CODE&gt;&lt;img /&gt;&lt;img&gt;Each row is one physical node with its pool, current state, IP, the number of minutes it has been in that state, and (when applicable) the failure code.&lt;/img&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;For the article's 1st scenario, the query becomes:&lt;/P&gt;
&lt;LI-CODE lang="kusto"&gt;| where State == "leavingpool" and MinutesInState &amp;gt; 60&lt;/LI-CODE&gt;
&lt;P&gt;For the article's 2nd scenario, it becomes:&lt;/P&gt;
&lt;LI-CODE lang="kusto"&gt;| where State == "unusable"&lt;/LI-CODE&gt;
&lt;P&gt;Both are one-line filters on the same table. You can also filter based on any node state, See: &lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/batch/batch-get-resource-counts#node-state-counts" target="_blank" rel="noopener"&gt;Node state counts&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H1&gt;&lt;STRONG&gt;6. The alert&amp;nbsp;&lt;/STRONG&gt;&lt;/H1&gt;
&lt;P&gt;I bet that you won't need a fixed notifications just telling you that the system was checking in (heartbeats). You will need an intelligent alert that would only fire when a specific threshold was crossed.&lt;/P&gt;
&lt;P&gt;By leveraging the custom table (&lt;EM&gt;BatchNodeInventory_CL&lt;/EM&gt;) I created, setting up this alert becomes straightforward using a Log Analytics scheduled query rule:&lt;/P&gt;
&lt;LI-CODE lang="kusto"&gt;BatchNodeInventory_CL
| where TimeGenerated &amp;gt; ago(10m)
| summarize LastSnapshot=max(TimeGenerated),
            State=any(State_s), Pool=any(PoolId_s), IP=any(IPAddress),
            Transition=any(StateTransitionTime_t),
            ErrCode=any(ErrorCode_s), ErrMsg=any(ErrorMessage_s)
      by NodeId=NodeId_s
| extend MinutesInState = datetime_diff("minute", LastSnapshot, Transition)
| where State in ("unusable","leavingpool","starttaskfailed","offline")
     or MinutesInState &amp;gt; 60&lt;/LI-CODE&gt;
&lt;P&gt;The KQL query for the alert is very similar to the diagnostic query we used earlier, but with a few critical modifications designed for automated monitoring:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;
&lt;P&gt;Shorter Time Window (ago(10m)) as this alert rule will run frequently.&lt;/P&gt;
&lt;/LI&gt;
&lt;LI&gt;
&lt;P&gt;We added ErrMsg=any(ErrorMessage_s) to the summarize block. This ensures that when the alert fires, the on-call engineer can read the actual error description without having to dig through logs immediately.&lt;/P&gt;
&lt;/LI&gt;
&lt;LI&gt;
&lt;P&gt;The final line instructs Azure Monitor to trigger only if a node enters a distinctly bad state (like &lt;EM&gt;unusable&lt;/EM&gt;, &lt;EM&gt;starttaskfailed&lt;/EM&gt;, or &lt;EM&gt;offline&lt;/EM&gt;) &lt;STRONG&gt;OR &lt;/STRONG&gt;a&amp;nbsp;node has been stuck in its current state (such as &lt;EM&gt;leavingpool&lt;/EM&gt;) for more than 60 minutes (MinutesInState &amp;gt; 60)&lt;/P&gt;
&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;This rule is typically bound to an &lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/azure-monitor/alerts/action-groups?ocid=AID754288&amp;amp;wt.mc_id=CFID0448" target="_blank" rel="noopener"&gt;Action Group&lt;/A&gt; (which handles notifications via email, Microsoft Teams, or PagerDuty). The generated payload directly includes the NodeId, Pool, State, ErrorCode, and MinutesInState. This means the on-call engineer receives the exact identifier of the failing node and the cause of the failure in a single, actionable message.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H1&gt;&lt;STRONG&gt;7. Bonus &lt;/STRONG&gt;&lt;/H1&gt;
&lt;P&gt;Because we have already done the heavy lifting of continuously ingesting detailed, per-node data into the &lt;STRONG&gt;BatchNodeInventory_CL&lt;/STRONG&gt; table, that exact same dataset can do double duty.&lt;/P&gt;
&lt;P&gt;Without running any additional polling scripts, this single table can act as a live inventory dashboard, providing aggregate counts per pool.&lt;/P&gt;
&lt;LI-CODE lang="kusto"&gt;BatchNodeInventory_CL
| where TimeGenerated &amp;gt; ago(30m)
| summarize Nodes=dcount(NodeId_s),
            Idle=dcountif(NodeId_s, State_s=="idle"),
            Running=dcountif(NodeId_s, State_s=="running"),
            Unusable=dcountif(NodeId_s, State_s=="unusable"),
            StartTaskFailed=dcountif(NodeId_s, State_s=="starttaskfailed")
      by Pool=PoolId_s&lt;/LI-CODE&gt;&lt;img&gt;The KQL results&lt;/img&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H1 data-path-to-node="8"&gt;&lt;STRONG&gt;8. Deploying in Production&lt;/STRONG&gt;&lt;/H1&gt;
&lt;P&gt;To make this solution reliable, the Python script needs to run as a continuous background process.&lt;/P&gt;
&lt;P&gt;I would say that you have 3 infrastructure options for deploying this collector script:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure Function (Timer trigger, Python)&lt;/STRONG&gt;: In my viewpoint, it's ideal because it is serverless, executes on a cron schedule and integrates seamlessly with Managed Identities and Key Vault + it's good cost.&lt;/LI&gt;
&lt;LI&gt;
&lt;P&gt;&lt;STRONG&gt;Container App Job (Schedule trigger)&lt;/STRONG&gt;: I would choose this if the collector needs to run within a private Virtual Network (VNet) to reach a secured Batch account, or if you require a custom base image.&lt;/P&gt;
&lt;/LI&gt;
&lt;LI&gt;
&lt;P&gt;&lt;STRONG&gt;A small VM with cron/Scheduled Task&lt;/STRONG&gt;: I would recommend if you already have a management VM and prefer to utilize the current infra not to provision new Azure resources.&lt;/P&gt;
&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;STRONG&gt;Important Notes&lt;/STRONG&gt;:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;EM&gt;An Azure Monitor Log Analytics Function is not viable here, as this solution requires initiating outbound REST API calls outside of the Log Analytics environment.&lt;/EM&gt;&lt;/LI&gt;
&lt;LI&gt;
&lt;P&gt;&lt;EM&gt;The Log Analytics Shared Key should be stored securely in a Key Vault, or alternatively, the ingestion method should be migrated to Data Collection Endpoints (DCE) / Data Collection Rules (DCR) to eliminate the need for shared keys entirely.&lt;/EM&gt;&lt;/P&gt;
&lt;/LI&gt;
&lt;LI&gt;&lt;EM&gt;As mentioned previously, the local &lt;SPAN class="lia-text-color-6"&gt;az account get-access-token&lt;/SPAN&gt; command should be replaced with the Azure SDK's Managed Identity token flow.&lt;/EM&gt;&lt;/LI&gt;
&lt;LI&gt;
&lt;P&gt;This solution doesn't replace built-in monitoring. Fleet-wide metrics, task events, and pool events are all still worth ingesting into the same LAW so a single dashboard covers everything. This collector just adds one extra dimension that some of our customers asked for.&amp;nbsp;&lt;/P&gt;
&lt;/LI&gt;
&lt;/UL&gt;</description>
      <pubDate>Tue, 25 Aug 2026 18:04:45 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/azure-paas-blog/tracking-batch-node-state-and-duration-in-log-analytics/ba-p/4547582</guid>
      <dc:creator>Ahmed_Khaled</dc:creator>
      <dc:date>2026-08-25T18:04:45Z</dc:date>
    </item>
  </channel>
</rss>

