<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>rss.livelink.threads-in-node</title>
    <link>https://techcommunity.microsoft.com/t5/itops-talk/ct-p/ITOpsTalk</link>
    <description>rss.livelink.threads-in-node</description>
    <pubDate>Sat, 26 Sep 2026 08:13:19 GMT</pubDate>
    <dc:creator>ITOpsTalk</dc:creator>
    <dc:date>2026-09-26T08:13:19Z</dc:date>
    <item>
      <title>Advanced Windows Firewall</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/advanced-windows-firewall/ba-p/4551626</link>
      <description>&lt;P&gt;If you administer Windows, you’re already aware of Windows Firewall. You might use it to allow an application, open a port, or block unwanted inbound traffic. But those familiar tasks only scratch the surface of what it can do.&lt;/P&gt;
&lt;P&gt;The Microsoft Learn module &lt;A href="https://learn.microsoft.com/en-us/training/modules/understand-advanced-windows-firewall/" target="_blank"&gt;Understand advanced Windows Firewall&lt;/A&gt; takes you beyond basic rules and introduces Windows Firewall as a powerful platform for host-based segmentation, authenticated access, traffic protection, operational evidence, and incident response. In the module you'll learn about the following:&lt;/P&gt;
&lt;H3&gt;Rule creation and management&lt;/H3&gt;
&lt;UL&gt;
&lt;LI&gt;View effective, enabled rules in the &lt;CODE&gt;ActiveStore&lt;/CODE&gt;.&lt;/LI&gt;
&lt;LI&gt;Inspect the port, address, application, and service filters associated with a rule.&lt;/LI&gt;
&lt;LI&gt;Create precisely scoped inbound rules for services such as HTTPS, WinRM, Remote Desktop, and WMI.&lt;/LI&gt;
&lt;LI&gt;Enable, disable, modify, and remove existing rules with PowerShell.&lt;/LI&gt;
&lt;LI&gt;Define a traffic contract before creating a rule.&lt;/LI&gt;
&lt;LI&gt;Correctly distinguish local and remote ports and addresses.&lt;/LI&gt;
&lt;LI&gt;Scope rules by protocol, port, address, application, service, profile, user, and computer.&lt;/LI&gt;
&lt;LI&gt;Restrict rules to stable executable paths, Windows services, or packaged application identities.&lt;/LI&gt;
&lt;LI&gt;Limit access to management subnets, jump hosts, privileged workstations, application tiers, and collectors.&lt;/LI&gt;
&lt;LI&gt;Create consistent rule names and groups for ownership and automation.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;Firewall profiles&lt;/H3&gt;
&lt;UL&gt;
&lt;LI&gt;Apply different policies to Domain, Private, and Public networks.&lt;/LI&gt;
&lt;LI&gt;Enable the firewall and block unmatched inbound traffic on every profile.&lt;/LI&gt;
&lt;LI&gt;Restrict administrative exceptions to the profiles that require them.&lt;/LI&gt;
&lt;LI&gt;Inspect active network profiles and their default actions.&lt;/LI&gt;
&lt;LI&gt;Design policies that remain secure during DNS, routing, domain-controller, or network-adapter failures.&lt;/LI&gt;
&lt;LI&gt;Test how network failure states affect profile selection.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;Host segmentation&lt;/H3&gt;
&lt;UL&gt;
&lt;LI&gt;Build a traffic matrix describing permitted communication between device and application tiers.&lt;/LI&gt;
&lt;LI&gt;Implement default-deny inbound segmentation.&lt;/LI&gt;
&lt;LI&gt;Block unnecessary workstation-to-workstation communication.&lt;/LI&gt;
&lt;LI&gt;Preserve approved management, monitoring, application, recovery, and domain-member traffic.&lt;/LI&gt;
&lt;LI&gt;Reduce lateral movement through controlled management paths.&lt;/LI&gt;
&lt;LI&gt;Deliver firewall policy centrally through Group Policy.&lt;/LI&gt;
&lt;LI&gt;Disable local firewall-rule merging.&lt;/LI&gt;
&lt;LI&gt;Disable local connection-security-rule merging.&lt;/LI&gt;
&lt;LI&gt;Stage enforcement through logging, discovery, pilots, and role-based deployment.&lt;/LI&gt;
&lt;LI&gt;Define success criteria and maintain a tested rollback path.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;IPsec and identity-based access&lt;/H3&gt;
&lt;UL&gt;
&lt;LI&gt;Design IPsec connection security rules for peer authentication.&lt;/LI&gt;
&lt;LI&gt;Provide packet integrity, replay protection, and optional encryption.&lt;/LI&gt;
&lt;LI&gt;Select Kerberos or certificate-based authentication for different trust scenarios.&lt;/LI&gt;
&lt;LI&gt;Protect legacy plaintext applications without changing the application.&lt;/LI&gt;
&lt;LI&gt;Coordinate secure firewall rules with compatible connection security rules.&lt;/LI&gt;
&lt;LI&gt;Use request authentication during deployment before enforcing required authentication.&lt;/LI&gt;
&lt;LI&gt;Correctly define IPsec endpoints and traffic selectors.&lt;/LI&gt;
&lt;LI&gt;Require authenticated traffic before allowing access.&lt;/LI&gt;
&lt;LI&gt;Authorize traffic by Active Directory user-group membership.&lt;/LI&gt;
&lt;LI&gt;Require both an authorized user and an authorized managed computer.&lt;/LI&gt;
&lt;LI&gt;Combine identity with network, service, application, and profile restrictions.&lt;/LI&gt;
&lt;LI&gt;Create narrowly scoped authenticated bypass rules.&lt;/LI&gt;
&lt;LI&gt;Design governed identity exceptions.&lt;/LI&gt;
&lt;LI&gt;Validate both successful and denied authorization scenarios.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;Outbound traffic control&lt;/H3&gt;
&lt;UL&gt;
&lt;LI&gt;Understand how stateful inspection permits response traffic.&lt;/LI&gt;
&lt;LI&gt;Avoid unnecessary inbound rules for dynamic client ports.&lt;/LI&gt;
&lt;LI&gt;Identify the dependencies required before introducing outbound default-deny.&lt;/LI&gt;
&lt;LI&gt;Restrict administrative tools, service accounts, high-risk applications, and servers to approved destinations.&lt;/LI&gt;
&lt;LI&gt;Account for dynamic cloud services, proxies, certificate endpoints, and content delivery networks.&lt;/LI&gt;
&lt;LI&gt;Introduce outbound restrictions gradually through discovery, narrow allow rules, pilots, and monitoring.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;Logging and evidence&lt;/H3&gt;
&lt;UL&gt;
&lt;LI&gt;Enable logging for dropped packets and successful connections.&lt;/LI&gt;
&lt;LI&gt;Configure the log location and maximum size for every profile.&lt;/LI&gt;
&lt;LI&gt;Verify effective logging settings.&lt;/LI&gt;
&lt;LI&gt;Interpret firewall log fields such as action, protocol, address, port, interface, and direction.&lt;/LI&gt;
&lt;LI&gt;Use observed traffic to discover application dependencies.&lt;/LI&gt;
&lt;LI&gt;Distinguish observed traffic from authorized traffic.&lt;/LI&gt;
&lt;LI&gt;Use dropped-packet records to confirm that traffic reached the host firewall.&lt;/LI&gt;
&lt;LI&gt;Use successful-connection records to confirm firewall admission.&lt;/LI&gt;
&lt;LI&gt;Forward firewall evidence to protected central storage.&lt;/LI&gt;
&lt;LI&gt;Correlate firewall data with process, authentication, application, and network telemetry.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;Troubleshooting&lt;/H3&gt;
&lt;UL&gt;
&lt;LI&gt;Follow a structured diagnostic sequence from the application listener through firewall and IPsec state.&lt;/LI&gt;
&lt;LI&gt;Inspect the merged runtime policy rather than only an individual policy source.&lt;/LI&gt;
&lt;LI&gt;Trace an effective rule back to Group Policy or another originating store.&lt;/LI&gt;
&lt;LI&gt;Identify conflicting or overriding block rules.&lt;/LI&gt;
&lt;LI&gt;Inspect active IPsec rules and main-mode and quick-mode security associations.&lt;/LI&gt;
&lt;LI&gt;Diagnose authentication, trust, time, name-resolution, selector, and cryptographic mismatches.&lt;/LI&gt;
&lt;LI&gt;Capture and interpret IPsec negotiation traffic on UDP ports 500 and 4500.&lt;/LI&gt;
&lt;LI&gt;Differentiate firewall admission failures from application or identity failures.&lt;/LI&gt;
&lt;LI&gt;Make controlled policy changes without disabling the firewall or creating unrestricted exceptions.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Windows Firewall might already be a familiar part of your Windows environment. This module will help you appreciate just how much security and diagnostic functionality is built into it—and how to apply that functionality with greater precision.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Start learning:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/understand-advanced-windows-firewall/" target="_blank"&gt;Understand advanced Windows Firewall&lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Sun, 30 Aug 2026 22:38:47 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/advanced-windows-firewall/ba-p/4551626</guid>
      <dc:creator>OrinThomas</dc:creator>
      <dc:date>2026-08-30T22:38:47Z</dc:date>
    </item>
    <item>
      <title>Active Directory Domain Services modules on Microsoft Learn</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/active-directory-domain-services-modules-on-microsoft-learn/ba-p/4547604</link>
      <description>&lt;MAIN&gt;
&lt;P&gt;The following is a list of Active Directory Domain Services modules on Microsoft Learn sorted from introductory to advanced. Is there an Active Directory topic that isn't listed here that you'd like to see covered? (AD CS modules are coming).&lt;/P&gt;
&lt;P&gt;If you know all this, remember we have the free 45 minute practical Active Directory administration test that gets you a validated Microsoft credential. You can take the test here &lt;A class="lia-external-url" href="https://aka.ms/ADDSAppliedSkillTest" target="_blank" rel="noopener"&gt;https://aka.ms/ADDSAppliedSkillTest&lt;/A&gt;.&lt;/P&gt;
&lt;H2 id="stage-1-build-the-foundation"&gt;Introductory&lt;/H2&gt;
&lt;H3 id="introduction-to-ad-ds"&gt;1. Introduction to AD DS&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/introduction-to-ad-ds/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/introduction-to-ad-ds/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge beginner"&gt;BEGINNER&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module establishes the essential AD DS vocabulary and mental model. It introduces directory services, users, groups, computers, forests, domains, sites, domain controllers, and organizational units, then shows how directory objects and their properties are managed. It should come first because nearly every later module assumes familiarity with these structures and their relationships.&lt;/P&gt;
&lt;H2 id="stage-2-learn-core-administration"&gt;Intermediate&lt;/H2&gt;
&lt;H3 id="create-and-manage-active-directory-objects"&gt;2. Create and manage Active Directory objects&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/create-manage-active-directory-objects/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/create-manage-active-directory-objects/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge intermediate"&gt;INTERMEDIATE&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module turns the foundational concepts into routine administrative skills. It covers users, groups, computers, organizational units, object properties, object creation and configuration, bulk user-account management, and basic domain controller maintenance. The module is the natural bridge from understanding AD DS to operating it.&lt;/P&gt;
&lt;H3 id="manage-active-directory-domain-services-using-powershell-cmdlets"&gt;3. Manage Active Directory Domain Services using PowerShell cmdlets&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/manage-active-directory-domain-services-use-powershell-cmdlets/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/manage-active-directory-domain-services-use-powershell-cmdlets/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge intermediate"&gt;INTERMEDIATE&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module applies PowerShell to the object-management tasks learned previously. It introduces cmdlets for creating and maintaining users, groups, group memberships, computers, organizational units, and other directory objects, giving learners a repeatable and scalable alternative to graphical administration. Prior PowerShell familiarity is expected.&lt;/P&gt;
&lt;H3 id="deploy-and-manage-active-directory-domain-services-domain-controllers"&gt;4. Deploy and manage Active Directory Domain Services domain controllers&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/deploy-manage-active-directory-domain-services-domain-controllers/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/deploy-manage-active-directory-domain-services-domain-controllers/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge intermediate"&gt;INTERMEDIATE&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module moves from directory objects to the servers that host the directory. It reviews forests and domains, explains domain controller topology, covers domain controller deployment and movement between sites, and introduces operations master roles. It provides the deployment context needed before studying domain controller maintenance and role placement in greater depth.&lt;/P&gt;
&lt;H3 id="manage-ad-ds-domain-controllers-and-fsmo-roles"&gt;5. Manage AD DS domain controllers and FSMO roles&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/manage-active-directory-domain-services-flexible-single-master-operation-roles/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/manage-active-directory-domain-services-flexible-single-master-operation-roles/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge intermediate"&gt;INTERMEDIATE&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module deepens domain controller administration through deployment, maintenance, backup and recovery considerations, global catalog placement, Flexible Single Master Operations (FSMO) role placement and management, and an introduction to schema management. It builds on the preceding deployment module and prepares learners for later architecture, recovery, and design topics.&lt;/P&gt;
&lt;H3 id="implement-group-policy-objects"&gt;6. Implement Group Policy Objects&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/implement-group-policy-objects/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/implement-group-policy-objects/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge intermediate"&gt;INTERMEDIATE&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module introduces domain-based Group Policy Objects (GPOs), including scope, inheritance, creation, configuration, storage, administrative templates, and the Central Store. It should precede the other Group Policy modules because it explains both the logical processing model and the underlying storage components on which later policy and troubleshooting work depends.&lt;/P&gt;
&lt;H3 id="create-and-configure-group-policy-objects-in-active-directory"&gt;7. Create and configure Group Policy Objects in Active Directory&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/create-configure-group-policy-objects-active-directory/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/create-configure-group-policy-objects-active-directory/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge intermediate"&gt;INTERMEDIATE&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module reinforces GPO definition, scope, inheritance, and domain-based configuration, then extends those skills into domain password policy and fine-grained password policy. Although some content overlaps the preceding module, its identity-policy focus makes it a useful second step after learners understand GPO storage, templates, and general processing.&lt;/P&gt;
&lt;H3 id="manage-security-in-active-directory"&gt;8. Manage security in Active Directory&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/manage-security-active-directory/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/manage-security-active-directory/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge intermediate"&gt;INTERMEDIATE&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module addresses practical directory hardening through user rights, access restrictions, delegated permissions, the Protected Users group, Windows Defender Credential Guard, NTLM blocking, and identification of problematic accounts. It connects identity administration and Group Policy knowledge to least-privilege operations and modern credential protection.&lt;/P&gt;
&lt;H3 id="implement-and-manage-active-directory-certificate-services"&gt;9. Implement and manage Active Directory Certificate Services&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/implement-manage-active-directory-certificate-services/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/implement-manage-active-directory-certificate-services/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge intermediate"&gt;INTERMEDIATE&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module introduces public key infrastructure and Active Directory Certificate Services (AD CS), including certification authority types, AD CS design and implementation, certificate enrollment, revocation, and trust. It belongs after core directory security because certificates extend AD-based authentication and authorization into a broader trust infrastructure.&lt;/P&gt;
&lt;H2&gt;Advanced&lt;/H2&gt;
&lt;H3 id="understand-active-directory-group-policy-security-settings"&gt;10. Understand Active Directory Group Policy security settings&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/understand-active-directory-security-policies/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/understand-active-directory-security-policies/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge advanced"&gt;ADVANCED&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module provides an in-depth treatment of security policy design and operation at scale. It covers password, lockout, and Kerberos account policies; fine-grained password policies; user rights; security options; auditing; secured groups, services, registry keys, files, and logs; network and application policies; and Windows Server 2025 OSConfig baselines. Its assumptions about GPO processing, security identifiers, access tokens, ACLs, Kerberos, and NTLM make it the bridge from practical GPO administration to protocol-level authentication hardening.&lt;/P&gt;
&lt;H3 id="active-directory-domain-services-authentication-and-kerberos-hardening"&gt;11. Active Directory Domain Services authentication and Kerberos hardening&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/active-directory-authentication-kerberos/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/active-directory-authentication-kerberos/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge advanced"&gt;ADVANCED&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module examines Windows authentication end to end, explaining how SSPI, Negotiate, NTLM, and Kerberos interact and tracing Kerberos from Ticket Granting Ticket issuance through service-ticket presentation and authorization. It covers diagnosis of Service Principal Name, delegation, Privilege Attribute Certificate, encryption-type, and NTLM fallback problems; staged migration from NTLM to Kerberos; RC4 auditing and remediation for Windows Server 2025; and planning for PKINIT agility, Kerberos encryption policy, SMB NTLM blocking, and password-change hardening. It caps the security stage because it requires learners to combine policy, identity, logging, and authentication-protocol knowledge.&lt;/P&gt;
&lt;H3 id="understand-how-active-directory-domain-services-uses-dns"&gt;12. Understand how Active Directory Domain Services uses DNS&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/understand-active-directory-domain-name-system/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/understand-active-directory-domain-name-system/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge advanced"&gt;ADVANCED&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module explains and diagnoses the DNS mechanisms on which AD DS depends. It traces service discovery and domain controller location through SRV records, LDAP ping, site selection, and caching, and examines AD-integrated zones, application directory partitions, secure dynamic updates, registration, replication convergence, baselining, and failure recovery. This knowledge is essential before designing or troubleshooting site topology and replication.&lt;/P&gt;
&lt;H3 id="active-directory-sites-topology-and-replication"&gt;13. Active Directory sites, topology, and replication&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/active-directory-site-replication/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/active-directory-site-replication/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge advanced"&gt;ADVANCED&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module develops a detailed model of sites, subnets, site links, costs, bridgeheads, connection objects, the Knowledge Consistency Checker (KCC), and the Inter-Site Topology Generator (ISTG). It also covers DC Locator behavior, topology anti-patterns, update sequence numbers, replication vectors, object metadata, replication failure diagnosis, and the Windows Server 2025 replication priority boost. The module supplies the replication foundation needed for the deeper storage and design modules that follow.&lt;/P&gt;
&lt;H3 id="understand-the-active-directory-domain-services-database-and-sysvol"&gt;14. Understand the Active Directory Domain Services database and SYSVOL&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/understand-active-directory-database/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/understand-active-directory-database/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge advanced"&gt;ADVANCED&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module explains how a domain controller stores and replicates directory data and SYSVOL through separate mechanisms. It distinguishes the logical directory from the local ESE database and directory partitions, traces transactional writes and logical AD DS replication, explains SYSVOL access and DFS Replication, follows a GPO across its directory and file-system components, and defines safe backup, recovery, virtualization, and version boundaries. It is especially useful for diagnosing whether a failure lies in directory data, SYSVOL, replication, or client access.&lt;/P&gt;
&lt;H3 id="understand-the-active-directory-domain-services-schema"&gt;15. Understand the Active Directory Domain Services Schema&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/understand-active-directory-schema/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/understand-active-directory-schema/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge advanced"&gt;ADVANCED&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module examines the forest-wide schema that defines directory object classes, attributes, syntax, inheritance, links, identifiers, search behavior, security metadata, and global catalog content. It covers safe inspection, schema versions, governance, extension design and deployment, indexing and replication effects, access control, and schema-related troubleshooting. The topic is advanced because schema changes are forest-wide, effectively permanent, and demand disciplined testing and change control.&lt;/P&gt;
&lt;H3 id="design-a-single-domain-active-directory-forest"&gt;16. Design a single-domain Active Directory forest&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/design-single-domain-active-directory-forest/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/design-single-domain-active-directory-forest/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge advanced"&gt;ADVANCED&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module combines the preceding operational knowledge into a resilient single-domain forest design. It covers durable DNS and UPN namespaces, AD-integrated DNS, sites, subnets, site links, replication convergence and failure behavior, and placement of domain controllers, DNS servers, global catalogs, FSMO roles, and the authoritative time source. Learners also validate designs against objective evidence and failure scenarios.&lt;/P&gt;
&lt;H3 id="design-a-multi-domain-or-multi-forest-active-directory-environment"&gt;17. Design a multi-domain or multi-forest Active Directory environment&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/design-multi-domain-forest-trust/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/design-multi-domain-forest-trust/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge advanced"&gt;ADVANCED&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module extends single-domain design to scenarios requiring security isolation, legal separation, administrative autonomy, mergers, or replication boundaries. It addresses domain trees and partitions, namespace coexistence, delegation, GPO scope, schema governance, trust direction and security, cross-forest name resolution, global catalog and authentication paths, advanced replication and capacity, and read-only domain controller branch designs. It should follow single-domain design because additional domains and forests introduce complexity that must be justified.&lt;/P&gt;
&lt;H3 id="manage-advanced-features-of-ad-ds"&gt;18. Manage advanced features of AD DS&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/manage-advanced-features-of-ad-ds/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/manage-advanced-features-of-ad-ds/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge advanced"&gt;ADVANCED&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module brings together several specialized administration scenarios: creating trust relationships, implementing Enhanced Security Administrative Environment (ESAE) forests, monitoring and troubleshooting replication, and creating custom AD DS partitions. These tasks rely on a mature understanding of forest boundaries, trusts, security, partitions, and replication, so they are best approached after both architecture modules.&lt;/P&gt;
&lt;H3 id="deploy-and-manage-azure-iaas-active-directory-domain-controllers-in-azure"&gt;19. Deploy and manage Azure IaaS Active Directory domain controllers in Azure&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/deploy-manage-azure-iaas-active-directory-domain-controllers-azure/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/deploy-manage-azure-iaas-active-directory-domain-controllers-azure/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge advanced"&gt;ADVANCED&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module applies established AD DS administration to Azure infrastructure. It compares directory and identity service options, prepares Azure virtual networking for domain controllers, deploys and configures AD DS on Azure virtual machines, installs a replica domain controller, and creates a new forest on an Azure virtual network. The broad prerequisites in on-premises AD DS, Azure IaaS, networking, resiliency, security, PowerShell, automation, and monitoring make this advanced in the proposed path.&lt;/P&gt;
&lt;H3 id="active-directory-domain-services-migration"&gt;20. Active Directory Domain Services migration&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/active-directory-domain-services-migration/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/active-directory-domain-services-migration/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge advanced"&gt;ADVANCED&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module evaluates how to move an existing AD DS environment to Windows Server 2025. It compares in-place forest upgrade with migration to a new forest, then outlines the process for each approach. Although concise, migration is placed late because selecting and executing a safe strategy requires solid knowledge of domain controllers, operations roles, DNS, replication, security, architecture, recovery, and change management.&lt;/P&gt;
&lt;H3 id="troubleshoot-active-directory"&gt;21. Troubleshoot Active Directory&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/troubleshoot-active-directory/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/troubleshoot-active-directory/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge advanced"&gt;ADVANCED&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module provides broad recovery and troubleshooting coverage for AD DS failures and degraded performance. It addresses restoring deleted objects with the AD Recycle Bin, recovering the AD DS database and SYSVOL, troubleshooting replication, and diagnosing hybrid authentication issues. The potential impact of recovery operations and the need to correlate several directory subsystems make this an advanced operational module despite its relatively compact scope.&lt;/P&gt;
&lt;H3 id="troubleshoot-active-directory-domain-services-replication"&gt;22. Troubleshoot Active Directory Domain Services replication&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/troubleshoot-active-directory-replication/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/troubleshoot-active-directory-replication/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge advanced"&gt;ADVANCED&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module develops a disciplined process for diagnosing replication failures with Repadmin, DCDiag, PowerShell, and event logs. It teaches learners to distinguish true failures from expected latency or transient topology states, isolate DNS, network, authentication, topology, and domain controller health dependencies, recognize data-consistency risks, apply a targeted correction or escalate safely, and verify convergence and the original business symptom. This diagnostic discipline should precede advanced recovery so learners can distinguish a correctable replication fault from a broader consistency failure that warrants restoration.&lt;/P&gt;
&lt;H3 id="advanced-active-directory-back-up-and-recovery"&gt;23. Advanced Active Directory back up and recovery&lt;/H3&gt;
&lt;P&gt;&lt;STRONG&gt;URL:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/en-us/training/modules/active-directory-backup-recovery/" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/training/modules/active-directory-backup-recovery/&lt;/A&gt;&lt;BR /&gt;&lt;STRONG&gt;Assessment:&lt;/STRONG&gt; &lt;SPAN class="badge advanced"&gt;ADVANCED&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;This module treats AD DS backup and recovery as Tier 0 architecture and incident response rather than a single restore operation. It covers backup design around recovery time and recovery point objectives, topology, retention, and trust boundaries; evidence-based backup validation; selection of object, domain controller, domain, or forest recovery scope; recovery of objects, attributes, hierarchies, DNS, and SYSVOL; domain and forest sequencing involving FSMO roles, RID state, global catalogs, and trusts; safe virtualized domain controller recovery; and comprehensive validation before reconnection. It is the final capstone because safe recovery requires the learner to integrate nearly every preceding topic while making high-impact decisions under controlled change and security processes.&lt;/P&gt;
&lt;/MAIN&gt;</description>
      <pubDate>Tue, 01 Sep 2026 22:51:08 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/active-directory-domain-services-modules-on-microsoft-learn/ba-p/4547604</guid>
      <dc:creator>OrinThomas</dc:creator>
      <dc:date>2026-09-01T22:51:08Z</dc:date>
    </item>
    <item>
      <title>Resilient Azure Platforms: Durable Functions, Cosmos DB, and DR by Design</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/resilient-azure-platforms-durable-functions-cosmos-db-and-dr-by/ba-p/4546001</link>
      <description>&lt;P&gt;Hello Folks!&lt;/P&gt;
&lt;P&gt;Operating at Azure scale means managing change across multiple interconnected systems. As applications, services, and dependencies evolve, resilience becomes a foundational design principle rather than an afterthought. In this&lt;A class="lia-external-url" href="https://www.youtube.com/playlist?list=PLjt5SKzX1iI8con7FJDB56G6hHqxGm7ki" target="_blank"&gt; Microsoft Azure Infra Summit 2026 session&lt;/A&gt;, Bhavana Konchada, Principal Software Engineer at Microsoft and lead architect of the Resilience Control Platform, takes us behind the scenes of a production-grade resilience platform built on Azure and explains the engineering choices that helped bring it to life.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;DIV class="lia-embeded-content" contenteditable="false"&gt;&lt;IFRAME src="https://www.youtube.com/embed/zJLeW84pky4?si=GAc2Cifdpr8RORJ4" width="100%" title="YouTube video player" allowfullscreen="allowfullscreen" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" frameborder="0" style="aspect-ratio: 16/9; height: auto;" sandbox="allow-scripts allow-same-origin allow-forms"&gt;
&lt;/IFRAME&gt;&lt;/DIV&gt;
&lt;H2&gt;Why IT Pros Should Care&lt;/H2&gt;
&lt;P&gt;Most of us have shipped a system that worked beautifully on day one and then quietly fell apart the first time something downstream blinked. Bhavana’s session is a brutally honest tour of the decisions you make early that determine whether your platform survives reality. Here’s what you walk away with:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;A pragmatic blueprint for service boundaries that you can actually operate at 2 a.m.&lt;/LI&gt;
&lt;LI&gt;Concrete Durable Functions patterns (Monitor, continue-as-new, idempotency) that keep long-running workflows healthy.&lt;/LI&gt;
&lt;LI&gt;A Cosmos DB partitioning strategy grounded in real access patterns, not gut feel.&lt;/LI&gt;
&lt;LI&gt;A multi-region, fail-and-continue mindset (instead of fail-and-recover) that holds up when a region disappears.&lt;/LI&gt;
&lt;LI&gt;Real lessons from production, including the “non-deterministic orchestration” outages nobody warns you about.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;In short, if you build, run, or modernize platforms on Azure, this session reshapes how you think about reliability.&lt;/P&gt;
&lt;H2&gt;What DR by Design Actually Means, a Technical Overview&lt;/H2&gt;
&lt;P&gt;Bhavana frames the Resilience Control Platform as five chapters: architecture and service boundaries, orchestration with Durable Functions, the Cosmos DB data layer, identity across multiple user realms, and the resilience playbook itself.&lt;/P&gt;
&lt;P&gt;The platform has four moving parts:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;A portal where operators define and monitor scenarios.&lt;/LI&gt;
&lt;LI&gt;An orchestration engine that executes long-running workflows.&lt;/LI&gt;
&lt;LI&gt;Cosmos DB as a shared persistence layer.&lt;/LI&gt;
&lt;LI&gt;Downstream infrastructure APIs the engine acts on.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The big “DR by Design” idea is that resiliency isn’t bolted on later. It’s a property of every choice from boundaries upward. As Bhavana puts it, you stop designing for “fail and recover” and start designing for “fail and continue.” Users don’t know (or care) which region runs their workflow; they just need it to run reliably, consistently, and without interruption.&lt;/P&gt;
&lt;H2&gt;How It Works, Under the Hood&lt;/H2&gt;
&lt;P&gt;Bhavana’s team made several deliberate design moves worth borrowing.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Arm’s-length service boundaries.&lt;/STRONG&gt; Version one had the portal and orchestration engine tightly coupled with shared dependency injection and a shared database context. It felt clean until they tried to operate it. Now the two services talk over REST contracts, each with its own dependencies. Yes, that means a bit of duplicated code. What they gained, independent deployments, isolated failures, and clear ownership, more than paid for it.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;The right runtime for the workload.&lt;/STRONG&gt; The portal is a session-driven web app, so it lives on App Service. The orchestration engine bursts on demand and runs workflows for minutes (sometimes hours), so it’s built on Durable Functions. Forcing both into one model would have looked simpler on paper and been worse in practice.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Accept Fast, Process Asynchronously.&lt;/STRONG&gt; Clicking Execute returns a 202 immediately. The orchestrator does the heavy lifting in the background and updates status in Cosmos DB. The portal just reflects progress. Users never wait on long workflows.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Durable Functions patterns that actually scale.&lt;/STRONG&gt; Three lessons stood out:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;The Monitor pattern replaces busy polling with durable timers. The orchestrator wakes up, checks status, and goes back to sleep without holding compute.&lt;/LI&gt;
&lt;LI&gt;Orchestrators are state machines, not scripts. Calling DateTime.UtcNow inside an orchestrator produces non-deterministic replay and random production failures. The fix is to use the orchestration context for time and IDs.&lt;/LI&gt;
&lt;LI&gt;Continue-as-new keeps replay history bounded. Long-running orchestrations otherwise spend more time replaying history than doing real work.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;Cosmos DB designed around access, not org charts.&lt;/STRONG&gt; Partitioning by tenant feels logical and creates hotspots the moment one tenant gets busy. The team partitions by entity (each plan owns its partition) and uses hierarchical keys combining plan ID and execution ID. They also lean on TTL for data lifecycle so completed records expire automatically, no cleanup jobs required.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Identity as an execution boundary.&lt;/STRONG&gt; Corporate users authenticate through Microsoft Entra with OpenID Connect. Operations users come in through a federated WS-Federation system. Instead of forking the app, the team built home realm discovery at the front door, normalized everything into a single identity model behind it, and added custom middleware in the Azure Functions isolated worker model to extract, validate, enrich, and fail-fast on every token. Authorization is config-driven so every endpoint gets the same treatment.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Multi-region from day one.&lt;/STRONG&gt; The full stack (portal, engine, APIs, supporting services) runs in parallel across regions, fronted by Azure Front Door as the global entry point. Health probes drive automatic regional failover with no human in the loop.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Cosmos DB single-write with automatic failover.&lt;/STRONG&gt; Multi-write looks attractive on a slide and introduces real conflict-resolution complexity. The team chose one primary write region plus a replica with automatic failover. The Cosmos SDK detects region unavailability and routes requests to the promoted region without application code changes.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Idempotency from day zero.&lt;/STRONG&gt; Once you have retries (and Front Door, the SDK, and your clients all retry), every operation has to be safe to run more than once. Client-provided IDs, Cosmos conflict detection (a 409 means “already succeeded”), and idempotent orchestration events make sure the same outcome lands no matter how many times a signal arrives.&lt;/P&gt;
&lt;H2&gt;Real-World Value, Use Cases, ROI, Scenarios&lt;/H2&gt;
&lt;P&gt;What does this buy you in practice?&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Scenario validation under stress&lt;/STRONG&gt; without compromising production. The platform is built to proactively validate and govern system behavior at scale.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Long-running workflows that survive everything.&lt;/STRONG&gt; Host restarts, transient downstream errors, regional failovers, none of them lose work in flight.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Predictable cost.&lt;/STRONG&gt; Durable timers and continue-as-new mean you stop paying for compute that’s only waiting.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Operability at scale.&lt;/STRONG&gt; Independent services, clean contracts, and centralized identity all mean a smaller cognitive load when something breaks at 2 a.m.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Honest tradeoffs.&lt;/STRONG&gt; Single-write Cosmos loses theoretical write latency in the second region and gains predictable behavior, no conflict ambiguity, and far easier debugging during failovers. That’s usually the right trade.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;In short, the platform behaves the same on a quiet Tuesday and during a regional outage. That’s the whole point.&lt;/P&gt;
&lt;H2&gt;Getting Started&lt;/H2&gt;
&lt;P&gt;You don’t need to build the Resilience Control Platform tomorrow. You can start applying these patterns this week.&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Map your service boundaries honestly. If two services share a DI container or database context, decouple them behind a REST contract.&lt;/LI&gt;
&lt;LI&gt;Pick runtimes by workload, not by consistency. Interactive UI on App Service; long-running orchestrations on Durable Functions.&lt;/LI&gt;
&lt;LI&gt;Adopt the 202-Accepted pattern for anything that could take more than a couple of seconds.&lt;/LI&gt;
&lt;LI&gt;Audit your Durable orchestrators for DateTime.UtcNow, Guid.NewGuid, and direct HTTP calls. Move them into activities, use the orchestration context for time and IDs, and apply continue-as-new on long loops.&lt;/LI&gt;
&lt;LI&gt;Revisit your Cosmos partition keys against actual access patterns and enable TTL for transient data.&lt;/LI&gt;
&lt;LI&gt;Stand up a second region behind Azure Front Door, enable Cosmos DB automatic failover, and make every write operation idempotent with client-provided IDs.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2&gt;Resources&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/well-architected/reliability/principles" target="_blank"&gt;Reliability design principles, Azure Well-Architected Framework&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/durable-task/common/durable-task-orchestrations" target="_blank"&gt;Durable Orchestrations overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/azure-functions/durable/" target="_blank"&gt;Azure Durable Functions documentation&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/azure-functions/" target="_blank"&gt;Azure Functions documentation&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/cosmos-db/hierarchical-partition-keys" target="_blank"&gt;Hierarchical partition keys in Azure Cosmos DB&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/cosmos-db/" target="_blank"&gt;Azure Cosmos DB documentation&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/frontdoor/" target="_blank"&gt;Azure Front Door documentation&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/entra/identity/" target="_blank"&gt;Microsoft Entra ID documentation&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Keep Learning at the Summit&lt;/H2&gt;
&lt;P&gt;Catch the full &lt;A class="lia-external-url" href="https://www.youtube.com/playlist?list=PLjt5SKzX1iI8con7FJDB56G6hHqxGm7ki" target="_blank"&gt;Microsoft Azure Infra Summit 2026 session playlist here&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;Cheers!&lt;/P&gt;
&lt;P&gt;Pierre Roman&lt;/P&gt;</description>
      <pubDate>Thu, 13 Aug 2026 07:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/resilient-azure-platforms-durable-functions-cosmos-db-and-dr-by/ba-p/4546001</guid>
      <dc:creator>Pierre_Roman</dc:creator>
      <dc:date>2026-08-13T07:00:00Z</dc:date>
    </item>
    <item>
      <title>Container Network Insights Agent (CNIA): Your AI Teammate for AKS Networking Incidents</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/container-network-insights-agent-cnia-your-ai-teammate-for-aks/ba-p/4545971</link>
      <description>&lt;P&gt;Hello Folks!&lt;/P&gt;
&lt;P&gt;If you run AKS in production, you already know the script. A pod cannot reach an external service, every dashboard says the cluster is healthy, and somebody is SSHing into a node with five browser tabs open trying to piece the story together.&amp;nbsp; This session from the Microsoft Azure Infra Summit 2026 tackles that exact pain.&lt;/P&gt;
&lt;P&gt;Shaifali Garg (PM for Azure Container Networking on AKS) sits down with Jonathan Wang, an AKS operator running 30 clusters across two regions on Cilium, and they walk through what a real networking incident feels like, then introduce the Container Network Insights Agent (CNIA) live in the cluster.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;DIV class="lia-embeded-content" contenteditable="false"&gt;&lt;IFRAME src="https://www.youtube.com/embed/WNFRZYewolA?si=9uddwWP5y-EawHwj" width="100%" title="YouTube video player" allowfullscreen="allowfullscreen" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" frameborder="0" style="aspect-ratio: 16/9; height: auto;" sandbox="allow-scripts allow-same-origin allow-forms"&gt;
&lt;/IFRAME&gt;&lt;/DIV&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;Why IT Pros Should Care&lt;/H2&gt;
&lt;P&gt;In Jonathan’s environment, about 40% of incidents end up being networking problems. The tools all exist (kubectl, dashboards, detectors, Hubble), but the time sink is figuring out which layer the problem lives in and what to check next. CNIA goes after that gap. Here is what you actually get back:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;A symptom-to-classification jump in seconds, so you skip the first 30 minutes of “is this DNS, policy, node, or app?”&lt;/LI&gt;
&lt;LI&gt;One chat window with one evidence table, one root cause, and one copy-paste fix command, instead of jumping across five tabs&lt;/LI&gt;
&lt;LI&gt;Senior SRE tribal knowledge baked into the workflow, so anyone on the team can run the same investigation a principal engineer would&lt;/LI&gt;
&lt;LI&gt;Read-only by design, so the agent never changes anything on your cluster. You stay the human in the loop&lt;/LI&gt;
&lt;LI&gt;Installs as an AKS extension (no Helm chart, no YAML to babysit), and Azure handles the lifecycle&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;In short, CNIA is not trying to replace your SRE team. It hands them back 20 or 30 minutes on every networking ticket, which adds up fast across a fleet.&lt;/P&gt;
&lt;H2&gt;What CNIA Is, A Technical Overview&lt;/H2&gt;
&lt;P&gt;Think of CNIA as an AI teammate that lives inside your AKS cluster as a pod. You describe what is broken in plain English, the way you would ping a senior engineer on Slack, and behind the scenes the agent does four things in order. It classifies the kind of problem (DNS, egress, policy, node, app), it pulls live evidence from your cluster, it analyzes that evidence, and it hands you back a clean report with evidence, root cause, and a copy-paste exec command.&lt;/P&gt;
&lt;P&gt;Two architectural choices stand out. First, the agent uses your own Azure OpenAI resource (bring your own), so prompts and diagnostic content stay in your tenant and your region. Microsoft does not see your diagnostic data, and nothing gets persisted externally. Second, the answer is grounded in evidence pulled from your cluster, not from the internet. Your pods, your policies, your CoreDNS, your host-level NIC and kernel counters. If the evidence is inconclusive, CNIA says so rather than fabricating a root cause. That last bit is what earns trust with senior SREs.&lt;/P&gt;
&lt;P&gt;CNIA fits inside the broader Advanced Container Networking Services (ACNS) story on AKS. ACNS gives you metrics in Azure Managed Prometheus and Grafana, stored and on-demand network logs with Hubble, and FQDN-based filtering with Cilium. CNIA sits on top, automating the triage loop across those signals so you do not have to walk through the playbook by hand every time.&lt;/P&gt;
&lt;H2&gt;How It Works, Under the Hood&lt;/H2&gt;
&lt;P&gt;The install is an AKS extension. Roughly 5 to 7 minutes from “az aks extension” to “you have an SRE buddy in your cluster.” One small pod runs continuously. A second helper only spins up on the node during a deep packet-drop investigation, reads host-level network counters, and is cleaned up right after. Nothing left behind.&lt;/P&gt;
&lt;P&gt;Permissions are deliberately narrow:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Read-only RBAC on the cluster. The agent looks, it never changes anything&lt;/LI&gt;
&lt;LI&gt;A workload identity tied to your Azure OpenAI resource. No shared credentials&lt;/LI&gt;
&lt;LI&gt;Outbound traffic is HTTPS to your OpenAI endpoint on port 443, and nothing else. If you want to log that further through an NSG or firewall, that is supported&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;On the safety side, CNIA layers two protections against prompt injection. The agent is scope-restricted by design, so off-topic requests get rejected straight away. In one of Jonathan’s demos, Shefali asks the agent to “delete core-dns” and to “write a script to scrape LinkedIn profiles.” Both are refused on the spot. The second layer is the read-only RBAC at the cluster level. Even if someone tricked the prompt into emitting a destructive command, the cluster itself would refuse. The pod’s execution is scoped to specific diagnostic commands. It is not an open shell.&lt;/P&gt;
&lt;P&gt;Honest tradeoffs, because you will ask:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;It is one cluster at a time. Multi-cluster correlation is not in scope yet&lt;/LI&gt;
&lt;LI&gt;It does not auto-remediate. It tells you the fix, you verify and run it&lt;/LI&gt;
&lt;LI&gt;It is AKS only. EKS and GKE are not supported today&lt;/LI&gt;
&lt;LI&gt;Session state lives in the pod in memory. If the pod restarts, you start a fresh chat (past sessions are still available in history)&lt;/LI&gt;
&lt;LI&gt;Heavy packet-drop investigations have been validated up to around 7 concurrent users on smaller clusters. The team is actively scaling that up&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Real-World Value&lt;/H2&gt;
&lt;P&gt;The session includes two demos that map directly to incidents you have probably lived through.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Demo 1, egress that silently dies.&lt;/STRONG&gt; Pods cannot reach google.com. CoreDNS resolves it fine, example.com works from the same pod, every dashboard says healthy. CNIA classifies it as an egress connectivity problem (not DNS) and surfaces the actual culprit: a Cilium network policy named “restrict external FQDN” with a toFQDN rule that only allows example.com. Everything else gets silently dropped at the egress gate. DNS was allowed, the TCP connection was not. The fix command (a kubectl patch to add google.com to the allow list) is right there in the report. End-to-end fix in under a minute.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Demo 2, the target port typo.&lt;/STRONG&gt; A service is down with connection refused. Pods running, service exists, endpoints populated, no network policies. The agent goes inside the pod, looks at the actual listening sockets, and proves the mismatch: target port 8080, but nginx listens on port 80. One-digit typo in YAML that no kubectl get would surface on its own.&lt;/P&gt;
&lt;P&gt;The ROI math is straightforward. If your team handles networking incidents weekly and each one costs 20 to 30 minutes of “where do I even start,” that capacity adds up across the org. And critically, the win is not just speed. When the one engineer who knows where to look goes on leave, the rest of the team is no longer stuck calling them at home.&lt;/P&gt;
&lt;H2&gt;Getting Started&lt;/H2&gt;
&lt;P&gt;Three steps. That is it.&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Read the public docs, get an overview, scan the use cases, and understand what CNIA does and does not cover&lt;/LI&gt;
&lt;LI&gt;Pick a cluster (dev or staging is a great place to start) and install the AKS extension. Give it 5 to 7 minutes&lt;/LI&gt;
&lt;LI&gt;Run a few real network tickets through it. Compare your time-to-answer before and after. Hit thumbs-up or thumbs-down in the chat so the product team sees real signal&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Pricing in preview: no license fee. You pay for the Azure OpenAI tokens it uses (your tenant, your resource), plus the tiny bit of cluster compute for the pod. If you already have Azure OpenAI in your tenant, just point CNIA at it.&lt;/P&gt;
&lt;H2&gt;Resources&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/aks/container-network-observability-guide" target="_blank"&gt;Diagnose and resolve AKS network issues with Advanced Container Networking Services&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/aks/advanced-container-networking-services-overview" target="_blank"&gt;Advanced Container Networking Services overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/aks/azure-cni-powered-by-cilium" target="_blank"&gt;Configure Azure CNI Powered by Cilium in AKS&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/aks/cluster-extensions" target="_blank"&gt;AKS cluster extensions&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/aks/workload-identity-deploy-cluster" target="_blank"&gt;Deploy and configure Microsoft Entra Workload ID on an AKS cluster&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/ai-services/openai/overview" target="_blank"&gt;What is Azure OpenAI Service?&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Keep Learning at the Summit&lt;/H2&gt;
&lt;P&gt;Catch the full Microsoft Azure Infra Summit 2026 session playlist here: https://www.youtube.com/playlist?list=PLjt5SKzX1iI8con7FJDB56G6hHqxGm7ki&lt;/P&gt;
&lt;P&gt;Cheers!&lt;/P&gt;
&lt;P&gt;Pierre Roman&lt;/P&gt;</description>
      <pubDate>Wed, 12 Aug 2026 07:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/container-network-insights-agent-cnia-your-ai-teammate-for-aks/ba-p/4545971</guid>
      <dc:creator>Pierre_Roman</dc:creator>
      <dc:date>2026-08-12T07:00:00Z</dc:date>
    </item>
    <item>
      <title>Operating Azure Backup at Scale: Day-2 Excellence for IaaS, PaaS, and Storage Workloads</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/operating-azure-backup-at-scale-day-2-excellence-for-iaas-paas/ba-p/4545638</link>
      <description>&lt;P&gt;Hello Folks!&lt;/P&gt;
&lt;P&gt;If you have ever inherited a sprawling Azure environment and quietly wondered whether every VM, database, AKS cluster, and storage account in it is actually being backed up the way the business thinks it is, you are in good company.&lt;/P&gt;
&lt;P&gt;In session this session of the Microsoft Azure Infra Summit 2026, Bhavya Tadikonda and Shobhit Garg from the Azure Resiliency product team walked us through how Azure Backup is evolving into a unified, application-centric service that protects IaaS, PaaS, AKS, PostgreSQL, and unstructured storage from a single pane of glass.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;DIV class="lia-embeded-content" contenteditable="false"&gt;&lt;IFRAME src="https://www.youtube.com/embed/y_QLEQTI_mk?si=JcPkIlYf2z6tc1f5" width="100%" title="YouTube video player" allowfullscreen="allowfullscreen" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" frameborder="0" style="aspect-ratio: 16/9; height: auto;" sandbox="allow-scripts allow-same-origin allow-forms"&gt;
&lt;/IFRAME&gt;&lt;/DIV&gt;
&lt;H2&gt;Why IT Pros Should Care&lt;/H2&gt;
&lt;P&gt;Backup is one of those topics nobody talks about until the day it really matters. Then it is the only topic. The session framed Azure Resiliency around three pillars (infrastructure resiliency, data resiliency, and cyber recovery), and Azure Backup sits squarely in the middle of the last two. The reason this session lands hard for ops teams is that the surface area we are expected to protect keeps growing: VMs, SQL on Azure VMs, SAP HANA, Sybase, AKS, PostgreSQL flexible servers, Azure Files, blobs, ADLS, and on it goes.&lt;/P&gt;
&lt;P&gt;Here is why this should matter to you:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;One vault model now protects IaaS, PaaS, AKS, PostgreSQL flexible server, and storage workloads, with consistent policies and reporting.&lt;/LI&gt;
&lt;LI&gt;Cyber resiliency is built into the vault layer with immutability, soft delete, and multi-user authorization, so backups themselves can survive a ransomware event.&lt;/LI&gt;
&lt;LI&gt;A new threat detection preview (powered by Microsoft Defender for Cloud) scans restore points and tags them healthy or suspicious before you recover.&lt;/LI&gt;
&lt;LI&gt;Azure Backup for AKS protects cluster resources and persistent volumes with granular restores and immutable recovery points.&lt;/LI&gt;
&lt;LI&gt;You can configure backups from VS Code through the Azure MCP server using natural language prompts, which is genuinely useful when you are protecting dozens of resources.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;In short, fewer point tools, fewer scripts, and a much better chance of actually meeting your RPO and RTO targets when the day comes.&lt;/P&gt;
&lt;H2&gt;What Operating Azure Backup at Scale Means, a Technical Overview&lt;/H2&gt;
&lt;P&gt;The session opened with a quick reminder that resiliency in Azure stands on three pillars working together. Infrastructure resiliency keeps the underlying VMs, zones, and networks alive. Data resiliency keeps your data intact, available, and recoverable. Cyber recovery assumes the worst (a ransomware attack or insider event) and gives you air-gapped, immutable backups plus isolated recovery to restore safely.&lt;/P&gt;
&lt;P&gt;Azure Backup is the connective tissue across data resiliency and cyber recovery. At the data layer, it offers snapshot tier backups for instant operational recovery (with up to a four-hour RPO), vault tier backups for long-term retention, and an archive tier for cold compliance storage. For databases, you get database-aware protection for SQL Server in Azure VMs, SAP HANA, and SAP ASE (Sybase), with point-in-time restore and log backups as frequent as every 15 minutes. That gets you to an RPO as low as 15 minutes for SQL, which is a number most IT pros will recognise as good enough for the vast majority of business apps.&lt;/P&gt;
&lt;P&gt;At the vault layer, three security primitives stack together: soft delete (deleted backups are kept for an additional retention window), immutability (no operation can shorten retention or destroy recovery points before expiry), and multi-user authorization (critical operations need approval from a second admin via a Resource Guard). These are not bolt-ons. They are baked into Recovery Services vaults and Backup vaults.&lt;/P&gt;
&lt;H2&gt;How It Works, Under the Hood&lt;/H2&gt;
&lt;P&gt;The session followed a Contoso scenario where John, a cloud architect, configures backup for an application VM and a database VM. He picks a Recovery Services vault, creates a backup policy, and defines frequency and retention based on his RTO and RPO requirements. For the Linux application tier, John enables the new &lt;STRONG&gt;agentless, crash-consistent backup&lt;/STRONG&gt;, which is non-invasive and protects performance-sensitive workloads without an in-guest agent.&lt;/P&gt;
&lt;P&gt;For the database tier, John enables Azure Backup for SQL in Azure VMs. The service auto-discovers all databases inside the VM, removes the manual config dance, and lets him layer log backups, differential backups, and archival retention. For SQL Always On, HANA HSR, and Sybase HA clusters, snapshot-based acceleration gives him faster backups and instant restores.&lt;/P&gt;
&lt;P&gt;Then John turns to cyber resiliency. From vault properties he reviews soft delete, immutability, and multi-user authorization, then enables the new &lt;STRONG&gt;threat detection preview&lt;/STRONG&gt;. This integration with Microsoft Defender for Cloud scans restore points for malware so you can confirm a recovery point is clean before you roll back. Inside the protected items view, each restore point is marked healthy or suspicious, which is exactly the signal you want during an incident response.&lt;/P&gt;
&lt;P&gt;For PaaS and cloud-native, Shobhit took over and walked through Azure Backup for AKS and Azure Backup for PostgreSQL flexible server. AKS protection covers the cluster resources, the persistent volumes, and the namespaces, with automated scheduled backups, granular restores, immutable recovery points, and flexible retention. PostgreSQL flexible server gets vaulted backups with long-term retention plus a unified view for monitoring and alerts.&lt;/P&gt;
&lt;P&gt;The piece that made the room sit up was the demo of configuring backup from VS Code using the &lt;STRONG&gt;Azure MCP server&lt;/STRONG&gt;. John installs the Azure MCP extension, validates mcp.json, opens the chat window, and starts the MCP server. He prompts it to list unprotected AKS clusters in his subscription, then asks it to configure backup for a specific cluster. The MCP server reuses an existing vault and policy, creates the protected item, and applies the enterprise security defaults. That is the kind of conversational ops experience that scales nicely when you have hundreds of resources.&lt;/P&gt;
&lt;P&gt;For unstructured data, Azure Backup brings file shares, ADLS data, application artifacts, and large object stores into the same vault-based model, with off-site protection, long-term retention, immutability, soft delete, and MUA applied consistently.&lt;/P&gt;
&lt;H2&gt;Real-World Value&lt;/H2&gt;
&lt;P&gt;So where does the ROI show up? A few honest scenarios:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Ransomware attack on production VMs.&lt;/STRONG&gt; With immutability and MUA, even a compromised admin account cannot destroy your recovery points. With threat detection, you avoid restoring an infected snapshot.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Accidental deletion of an AKS namespace.&lt;/STRONG&gt; Granular AKS backup gets you a controlled, application-aware restore without redeploying the whole cluster.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Compliance audit on a regulated workload.&lt;/STRONG&gt; Vault tier plus archive tier gives you the retention you need without inflating hot storage costs.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;A cloud architect onboarding 30 new VMs and 10 PostgreSQL servers.&lt;/STRONG&gt; Using Azure MCP from VS Code, they can configure backup conversationally instead of click-clicking through portal blades.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;A BCDR drill.&lt;/STRONG&gt; The resiliency agent (powered by Azure Copilot) can recommend enabling Azure Site Recovery on top of Azure Backup for stricter RTO and RPO, then guide you through enabling it.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Honest tradeoff: threat detection is in preview, agentless crash-consistent backup is newer than the in-guest variant, and multi-user authorization requires a Resource Guard that lives in a separate subscription (ideally a separate tenant). That is extra setup work, but it is the right design for separation of duties.&lt;/P&gt;
&lt;H2&gt;Getting Started&lt;/H2&gt;
&lt;P&gt;Concrete first steps you can take this week:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Open Backup Center (or the new Resiliency in Azure experience) and inventory what is already protected versus exposed.&lt;/LI&gt;
&lt;LI&gt;Pick one Recovery Services vault and turn on enhanced soft delete with a meaningful retention period, then make it AlwaysOn for production.&lt;/LI&gt;
&lt;LI&gt;Stand up a Resource Guard in a separate subscription or tenant and wire up MUA on your most critical vault.&lt;/LI&gt;
&lt;LI&gt;For a non-production AKS cluster, install the Backup extension and protect a namespace end to end, including a test restore.&lt;/LI&gt;
&lt;LI&gt;Try the Azure MCP server from VS Code to list unprotected resources and configure backup with a prompt.&lt;/LI&gt;
&lt;LI&gt;If you run SQL on Azure VMs, enable log backups every 15 minutes on one database and validate a point-in-time restore.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2&gt;Resources&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/backup/" target="_blank"&gt;Azure Backup documentation&lt;/A&gt; (official docs for vaults, policies, and workload protection)&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/backup/multi-user-authorization" target="_blank"&gt;Configure Multi-user authorization using Resource Guard&lt;/A&gt; (separation of duties for critical backup operations)&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/backup/threat-detection-overview" target="_blank"&gt;Threat detection in Azure Backup with Microsoft Defender for Cloud (preview)&lt;/A&gt; (healthy or suspicious tagging for VM restore points)&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/backup/azure-kubernetes-service-cluster-backup" target="_blank"&gt;Back up Azure Kubernetes Service by using Azure Backup&lt;/A&gt; (cluster resources, namespaces, and persistent volumes)&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/backup/backup-azure-database-postgresql" target="_blank"&gt;Azure Backup for PostgreSQL flexible server&lt;/A&gt; (vaulted backups with long-term retention)&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/site-recovery" target="_blank"&gt;Azure Site Recovery documentation&lt;/A&gt; (DR replication on top of Azure Backup)&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Keep Learning...&lt;/H2&gt;
&lt;P&gt;Catch the full Microsoft Azure Infra Summit 2026 session &lt;A class="lia-external-url" href="http://%20https://www.youtube.com/playlist?list=PLjt5SKzX1iI8con7FJDB56G6hHqxGm7ki" target="_blank"&gt;playlist here&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;Cheers!&lt;/P&gt;
&lt;P&gt;Pierre&lt;/P&gt;</description>
      <pubDate>Tue, 11 Aug 2026 07:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/operating-azure-backup-at-scale-day-2-excellence-for-iaas-paas/ba-p/4545638</guid>
      <dc:creator>Pierre_Roman</dc:creator>
      <dc:date>2026-08-11T07:00:00Z</dc:date>
    </item>
    <item>
      <title>Zonal Resiliency in Azure: Application-Centric Goals, Recovery Plans, and Drills</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/zonal-resiliency-in-azure-application-centric-goals-recovery/ba-p/4542514</link>
      <description>&lt;P&gt;Hello Folks&lt;/P&gt;
&lt;P&gt;&amp;nbsp;If you have ever stared at a multi-tier app in Azure and asked yourself, “Is this actually going to survive a zone outage?”, you are not alone. In session MAIS23 of the Microsoft Azure Infra Summit 2026, Bhavya, Aditya, and Chaya from the Azure Resiliency product team walked us through the new Resiliency in Azure experiences (formerly Azure Business Continuity Center) and showed how to stop treating resiliency as a per-resource checkbox and start treating it as an application-level outcome.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;DIV class="lia-embeded-content" contenteditable="false"&gt;&lt;IFRAME src="https://www.youtube.com/embed/dpayrI09zUA?si=dsdcR-aIgd-NSyo0" width="100%" title="YouTube video player" allowfullscreen="allowfullscreen" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" frameborder="0" style="aspect-ratio: 16/9; height: auto;" sandbox="allow-scripts allow-same-origin allow-forms"&gt;
&lt;/IFRAME&gt;&lt;/DIV&gt;
&lt;H2&gt;Why IT Pros Should Care&lt;/H2&gt;
&lt;P&gt;Most of us have lived this story. An app is “in the cloud”, spread across IaaS VMs, PaaS databases, an app service plan, and a shared Azure Firewall managed by some other team. Then a zonal blip hits, and suddenly nobody can answer the simple question: was this app supposed to be zone resilient or not?&lt;/P&gt;
&lt;P&gt;The session opened with a customer scenario called Zava, a fast-growing insurance company running a claims app at 99.9 percent availability that just lost more than $40,000 in revenue in one week because of zonal outages. That is the price tag the speakers put on the problem, and it lines up with the patterns I see every week.&lt;/P&gt;
&lt;P&gt;Here is why this matters to IT pros:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;You finally get a single pane to see zonal resiliency posture across IaaS, PaaS, and shared services.&lt;/LI&gt;
&lt;LI&gt;Resiliency goals are set at the application level, not buried inside each resource blade.&lt;/LI&gt;
&lt;LI&gt;You get tailored Azure Advisor recommendations plus an Azure Copilot guided flow that emits remediation scripts.&lt;/LI&gt;
&lt;LI&gt;You can run zone-down drills powered by Azure Chaos Studio without stitching together five different tools.&lt;/LI&gt;
&lt;LI&gt;Recovery plans orchestrate failover in a defined order, with on-demand readiness checks before the next real outage.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;In short, less guessing, less spreadsheet bookkeeping, and a lot more confidence that the app will behave the way you told the business it would.&lt;/P&gt;
&lt;H2&gt;What Resiliency in Azure Does, a Technical Overview&lt;/H2&gt;
&lt;P&gt;The team has rebranded Azure Business Continuity Center to &lt;STRONG&gt;Resiliency in Azure&lt;/STRONG&gt;. It is a unified solution that covers infra, data, and cyber resiliency in one place. Today the focus is &lt;STRONG&gt;zonal resiliency&lt;/STRONG&gt;, with regional disaster recovery (and proper RPO/RTO goals) on the roadmap.&lt;/P&gt;
&lt;P&gt;The central concept is the &lt;STRONG&gt;service group&lt;/STRONG&gt;. A service group is a logical application unit that can span subscriptions and resource groups. You add the VMs, databases, app service plans, Redis caches, and other Azure resources that make up an application, and from that point on, resiliency operations work against the whole app, not one resource at a time.&lt;/P&gt;
&lt;P&gt;There are two views you will spend most of your time in:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Resource resiliency&lt;/STRONG&gt;, a zonal configuration summary across the (roughly 20) resource types supported today.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Service group resiliency&lt;/STRONG&gt;, the same summary but pivoted to the application level, so you can prioritize the apps that need attention first.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The speakers were honest about scope. Goals today are a simple intent (“this service group should be evaluated for zonal resilience”). Once additional pillars like regional DR ship, goals will expand to include RPO and RTO targets. I appreciate that they did not oversell it.&lt;/P&gt;
&lt;H2&gt;How It Works, Under the Hood&lt;/H2&gt;
&lt;P&gt;Once a service group exists, the workflow has three big building blocks. Each one solves a problem I bet you have hit.&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt; Goals and recommendations.&lt;/STRONG&gt; You assign a zonal resiliency goal to the service group, and Azure Advisor surfaces tailored recommendations for the resources inside it. Two details I liked:&lt;/LI&gt;
&lt;/OL&gt;
&lt;UL&gt;
&lt;LI&gt;The view shows &lt;STRONG&gt;cost implications&lt;/STRONG&gt; before you flip the switch. Some Azure services have no cost delta for zone redundancy. Others do. You see it inline, not in a separate calculator tab.&lt;/LI&gt;
&lt;LI&gt;There is an &lt;STRONG&gt;Azure Copilot guided remediation&lt;/STRONG&gt; flow that walks you through the recommendation and, at the end, emits a script. That script accounts for resource-type corner cases (SKU changes, redeploys, and so on) and is meant to be run through your automation pipeline.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;You can also &lt;STRONG&gt;exclude&lt;/STRONG&gt; a resource with a reason (“not critical, zonal redundancy not required”) or &lt;STRONG&gt;manually attest&lt;/STRONG&gt; a resource when your own custom solution already provides resiliency that the platform cannot auto-detect. That escape hatch is important, because real environments always have a few weird cases.&lt;/P&gt;
&lt;OL start="2"&gt;
&lt;LI&gt;&lt;STRONG&gt; Application-centric recovery plans.&lt;/STRONG&gt; Instead of failing over one resource at a time, a recovery plan orchestrates the entire app. It auto-detects existing solutions (Azure Site Recovery for VMs, for example), lets you group and order the resources for failover, and excludes resources that are already configured for high availability (no point failing them over if they did not go down). You can run an &lt;STRONG&gt;on-demand readiness check&lt;/STRONG&gt; any time the app structure changes, so you find configuration drift before an outage finds it for you.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt; Zone-down drills powered by Azure Chaos Studio.&lt;/STRONG&gt; A zone-down drill template identifies the service group resources, pre-populates the right native faults per resource type (think a Redis cache fault, a VM scale set shutdown, and so on), bundles in identity and permission checks, monitoring, and the recovery plan you already built. When you execute, you pick the region and the target zone, the drill runs a pre-validation check, injects the fault, runs failover, then reprotection and failback, and tracks all of it as a single job in the execution report. Per-resource metrics let you visualize the actual downtime each component experienced. If a native fault is not what you want, you can override with a custom runbook.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;That last point is the part I think a lot of folks miss. &lt;STRONG&gt;A drill is not just fault injection.&lt;/STRONG&gt; It is fault injection plus failover plus reprotection plus failback, all measured and attestable in one place.&lt;/P&gt;
&lt;H2&gt;Real-World Value&lt;/H2&gt;
&lt;P&gt;Back to Zava. They needed to answer three questions: what is our current zonal resiliency posture across these Azure services, what should we prioritize against our 99.9 percent target, and how do we validate that we will actually perform during an outage? Resiliency in Azure answers all three without forcing the platform team to write a 200-line PowerShell script.&lt;/P&gt;
&lt;P&gt;Use cases that should be on your shortlist:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Regulated workloads&lt;/STRONG&gt; (insurance, healthcare, financial services) that need to evidence drills for compliance. The notes and manual attestation features were clearly designed with auditors in mind.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Apps with mixed estates&lt;/STRONG&gt;, where a central platform team owns shared services (firewalls, identity) and app teams own everything else. Service groups can be parented to mirror that org structure.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Apps with custom resiliency solutions&lt;/STRONG&gt; that the platform cannot detect. Manual attestation keeps the dashboard honest without forcing you to refactor.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Game-day rehearsals.&lt;/STRONG&gt; The pre-built zone-down template means you can run a meaningful drill in an afternoon instead of standing up a custom Chaos Studio experiment from scratch.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The honest tradeoff: zone redundancy is not free for every service, and not every resource type is in scope yet (around 20 today). Plan accordingly, exclude what is not critical, and attest what is covered by something else.&lt;/P&gt;
&lt;H2&gt;Getting Started&lt;/H2&gt;
&lt;P&gt;Here is the path I would take on a Monday morning:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Open the Azure portal and search for &lt;STRONG&gt;Resiliency&lt;/STRONG&gt;. You will land on the Resiliency in Azure page that replaces the old Business Continuity Center.&lt;/LI&gt;
&lt;LI&gt;Create a &lt;STRONG&gt;service group&lt;/STRONG&gt;. Add resources directly, or add resource groups if each resource group is already an application boundary in your environment.&lt;/LI&gt;
&lt;LI&gt;Assign the &lt;STRONG&gt;zonal resiliency goal&lt;/STRONG&gt; to the service group.&lt;/LI&gt;
&lt;LI&gt;Review the summary tiles. Exclude or manually attest the resources that need it.&lt;/LI&gt;
&lt;LI&gt;Walk the &lt;STRONG&gt;Advisor recommendations&lt;/STRONG&gt;. Use the Copilot guided flow to generate a remediation script and run it through your automation.&lt;/LI&gt;
&lt;LI&gt;Build an &lt;STRONG&gt;application-centric recovery plan&lt;/STRONG&gt;, group and order the resources, run an on-demand readiness check.&lt;/LI&gt;
&lt;LI&gt;Create a &lt;STRONG&gt;zone-down drill&lt;/STRONG&gt; from the template, validate identity, monitoring, and faults, then execute the drill in a non-production zone first.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2&gt;Resources&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/reliability/" target="_blank"&gt;Resiliency in Azure documentation&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/reliability/availability-zones-zonal-resource-resiliency" target="_blank"&gt;Zonal resources and zone resiliency&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/governance/service-groups/overview" target="_blank"&gt;Azure service groups overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/advisor/advisor-reference-reliability-recommendations" target="_blank"&gt;Azure Advisor reliability recommendations&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/chaos-studio/" target="_blank"&gt;Azure Chaos Studio documentation&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/site-recovery/site-recovery-overview" target="_blank"&gt;Azure Site Recovery overview&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Keep Learning at the Summit&lt;/H2&gt;
&lt;P&gt;Catch the full Microsoft Azure Infra Summit 2026 session playlist here: https://www.youtube.com/playlist?list=PLjt5SKzX1iI8con7FJDB56G6hHqxGm7ki&lt;/P&gt;
&lt;P&gt;Cheers!&lt;/P&gt;
&lt;P&gt;Pierre Roman&lt;/P&gt;</description>
      <pubDate>Thu, 06 Aug 2026 07:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/zonal-resiliency-in-azure-application-centric-goals-recovery/ba-p/4542514</guid>
      <dc:creator>Pierre_Roman</dc:creator>
      <dc:date>2026-08-06T07:00:00Z</dc:date>
    </item>
    <item>
      <title>Agentic Migrations and Modernization: How the Azure Migrate Agent Keeps Your Intent Alive End to End</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/agentic-migrations-and-modernization-how-the-azure-migrate-agent/ba-p/4542513</link>
      <description>&lt;P&gt;Hello Folks!&lt;/P&gt;
&lt;P&gt;If you have ever tried to move a few hundred VMs, a pile of databases, and a couple of web apps from on-prem to Azure, you already know the hard part is not the tooling.&lt;/P&gt;
&lt;P&gt;The hard part is keeping context, intent, and momentum alive across weeks of planning, hand-offs, and decisions. In session MAIS15 at the Microsoft Azure Infra Summit 2026, Ankur Gupta (Senior Product Manager on the Azure Migrate team) walked us through the new Azure Migrate agent, an AI layer that sits on top of Azure Migrate and carries your intent from “I have an idea” all the way to “the landing zone is deployed.”&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;DIV class="lia-embeded-content" contenteditable="false"&gt;&lt;IFRAME src="https://www.youtube.com/embed/oa8CHk33AuQ?si=WvlU89BjysZsTJLw" width="100%" title="YouTube video player" allowfullscreen="allowfullscreen" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" frameborder="0" style="aspect-ratio: 16/9; height: auto;" sandbox="allow-scripts allow-same-origin allow-forms"&gt;
&lt;/IFRAME&gt;&lt;/DIV&gt;
&lt;H2&gt;Why IT Pros Should Care&lt;/H2&gt;
&lt;P&gt;Ankur opened with a line that stuck with me. Infrastructure complexity has far outpaced human scale. We have MySQL here, PostgreSQL there, web apps, storage devices, networking gear, multiple dashboards, multiple alerts, and we are all expected to move faster than ever, with fewer mistakes. In short, migrations rarely fail because someone picked the wrong tool. They fail because the system between the stages breaks.&lt;/P&gt;
&lt;P&gt;Here is why the agentic approach matters for the folks in the trenches:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;It keeps context across the entire lifecycle, so the intent you set on day one is still the intent at execution.&lt;/LI&gt;
&lt;LI&gt;It guides you when you are stuck, instead of leaving you to figure out which of three discovery methods is the right one.&lt;/LI&gt;
&lt;LI&gt;It compresses tasks that used to take days of analysis (think side-by-side business cases) into a few hours.&lt;/LI&gt;
&lt;LI&gt;It connects IT ops, architects, and developers through a single thread of information, including a clean handoff to GitHub Copilot for code work.&lt;/LI&gt;
&lt;LI&gt;It builds on the Azure Migrate portal you already know, so nothing you have learned goes to waste.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;That last point is important. The portal does not go away. The agent is a layer on top. You can still do everything you do today.&lt;/P&gt;
&lt;H2&gt;What the Azure Migrate Agent Is, technical overview&lt;/H2&gt;
&lt;P&gt;Azure Migrate has always been Microsoft’s hub for discover, assess, and migrate. What Ankur showed at &lt;A class="lia-external-url" href="https://www.youtube.com/live/oa8CHk33AuQ" target="_blank"&gt;MAIS15 &lt;/A&gt;is the next evolution. Azure Migrate is becoming a migration control plane that spans the whole lifecycle (Decide, Plan, Execute), and the Azure Migrate agent is the conversational, guidance-oriented layer that ties it all together.&lt;/P&gt;
&lt;P&gt;In Ankur’s words, the agent is educational and guidance-oriented. You ask one natural-language question, like “how should I plan moving my VMware workloads to Azure,” and you get the next steps you actually need to take. Behind the scenes, the agent is doing three things very well.&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;It maintains state across the entire lifecycle. Preferences you set early stay with you.&lt;/LI&gt;
&lt;LI&gt;It carries context across discovery, plan, and execute. You can jump around, repeat steps, change your mind, and the agent remembers what your goal was.&lt;/LI&gt;
&lt;LI&gt;It recommends the right next move based on what it has learned about your intent.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;This is the heart of the “agentic” part. The agent is not a chatbot grafted onto a portal. It is a stateful workflow runner that remembers you.&lt;/P&gt;
&lt;H2&gt;How It Works, under the hood&lt;/H2&gt;
&lt;P&gt;The session walked through a full VMware-to-Azure scenario, and the flow is worth seeing because it shows how the pieces snap together.&lt;/P&gt;
&lt;P&gt;The agent supports three discovery methods today: appliance-based discovery, RV Tools, and the new Azure Migrate collector. The collector is the new lightweight option. It ships as a set of PowerShell scripts you run on a machine that can reach your vCenter, it produces a zip file, and you upload that zip to your Azure Migrate project. No appliance to deploy, no inbound network plumbing.&lt;/P&gt;
&lt;P&gt;Once the inventory lands, the agent reads it. Ankur asked for a summary and got a card showing 207 VMs, 177 SQL databases, one PostgreSQL instance, and some web apps. He then asked for all servers with an out-of-support OS, got a list of around 50, and tagged them right inside the conversation so he could refer back to them later.&lt;/P&gt;
&lt;P&gt;Next came the business case. Ankur asked the agent to generate one based on a modernize preference. A few minutes later, he had Azure cost, on-prem cost, and projected savings. Then he asked for a second business case for lift and shift, and a side-by-side comparison. The agent ran it, showed that the on-prem cost in the lift-and-shift comparison was higher and that lift-and-shift TCO savings were actually higher in that specific scenario, and gave him the data points he needed to bring the decision to leadership.&lt;/P&gt;
&lt;P&gt;From there, Ankur moved to application assessment, this time back in the portal. He created an assessment for two apps (Airsonic and Parts Unlimited), let the high-confidence plan run, and got a modernize recommendation with 100% readiness and a target cost of about $580 per month, plus an emissions estimate of 32 kgs of CO2. App Service for the web tier, Azure Database for PostgreSQL for the data tier. Both were flagged “ready with conditions,” with clickable links into why.&lt;/P&gt;
&lt;P&gt;Then came one of my favorite parts. Ankur connected GitHub for a Copilot Assessment, which adds code-level insights on top of the infrastructure readiness assessment. The system recalculated, and the migration effort estimate sharpened up.&lt;/P&gt;
&lt;P&gt;Finally, the agent built a wave plan from the assessment, then generated a platform landing zone aligned with Azure best practices. He could ask the agent about chosen defaults, request changes to the deployment mechanism, swap in a third-party firewall, or apply naming conventions. The agent produced a downloadable Infrastructure-as-Code template and handed it to a cloud architect, who refined it in their IDE using GitHub Copilot.&lt;/P&gt;
&lt;P&gt;That last handoff is the bridge between Azure Migrate’s planning world and the developer world.&lt;/P&gt;
&lt;H2&gt;Real-World Value (use cases, ROI, scenarios)&lt;/H2&gt;
&lt;P&gt;So where does this actually pay off? A few scenarios stood out.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Pitching the business case to leadership. Ankur framed the demo around “I need to pitch a migration proposal to the planning committee.” Generating modernize, lift-and-shift, and Azure VMware Solution business cases used to be days of spreadsheet work. With the agent, it is hours.&lt;/LI&gt;
&lt;LI&gt;Cleaning up legacy debt. Tagging out-of-support servers in one conversational step lets you plan upgrades without exporting CSVs and slicing them by hand.&lt;/LI&gt;
&lt;LI&gt;Mixed estates with web apps and databases. The agent surfaces App Service and Azure Database for PostgreSQL targets, gives SKU recommendations, and flags the warnings worth investigating.&lt;/LI&gt;
&lt;LI&gt;Closing the IT-to-developer gap. The GitHub Copilot Assessment and the IaC handoff to the IDE means developers and architects work from the same context.&lt;/LI&gt;
&lt;LI&gt;Reducing intent drift on long migrations. Multi-week journeys lose their original intent. The agent remembers.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;In short, the ROI here is measured in calendar time, not just dollars. And honestly, in fewer late-night calls when something goes sideways because nobody remembered the original decision.&lt;/P&gt;
&lt;P&gt;Tradeoffs worth flagging: the agentic capabilities are landing in preview, and outputs are advisory. You still need human review, testing, and governance on every recommendation. That is by design.&lt;/P&gt;
&lt;H2&gt;Getting Started (concrete first steps)&lt;/H2&gt;
&lt;P&gt;Here is a practical onramp.&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Stand up an Azure Migrate project in the Azure portal if you do not already have one.&lt;/LI&gt;
&lt;LI&gt;Pick a discovery method that fits your environment. If you cannot deploy an appliance, try the new collector. Download the PowerShell scripts, run them from a host that can reach vCenter, and upload the zip.&lt;/LI&gt;
&lt;LI&gt;Bring in the Azure Migrate agent from inside the portal and ask it to summarize your discovered inventory.&lt;/LI&gt;
&lt;LI&gt;Generate at least two business cases (modernize and lift-and-shift). Compare them.&lt;/LI&gt;
&lt;LI&gt;Run an application assessment on a small, representative set of apps. Connect GitHub and add a Copilot Assessment for code-level insight.&lt;/LI&gt;
&lt;LI&gt;Ask the agent to build a wave plan and a platform landing zone template, then push the IaC to your repo for the architects.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Start small, build the muscle, and scale out.&lt;/P&gt;
&lt;H2&gt;Resources&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/migrate/?view=migrate" target="_blank"&gt;Azure Migrate documentation (Microsoft Learn)&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://github.com/MicrosoftDocs/azure-docs/blob/main/articles/migrate/migrate-services-overview.md" target="_blank"&gt;About Azure Migrate, including the Azure Copilot migration agent (Microsoft Learn)&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/developer/github-copilot-app-modernization/overview" target="_blank"&gt;GitHub Copilot modernization overview (Microsoft Learn)&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/developer/github-copilot-app-modernization/modernization-agent/overview" target="_blank"&gt;GitHub Copilot modernization agent overview (Microsoft Learn)&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/developer/github-copilot-app-modernization/" target="_blank"&gt;GitHub Copilot modernization documentation (Microsoft Learn)&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/dotnet/azure/migration/appmod/quickstart" target="_blank"&gt;Assess and migrate a .NET project with GitHub Copilot modernization (Microsoft Learn)&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/dotnet/core/porting/github-copilot-app-modernization/overview" target="_blank"&gt;GitHub Copilot modernization overview for .NET (Microsoft Learn)&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Keep Learning at the Summit&lt;/P&gt;
&lt;P&gt;Catch the full Microsoft Azure Infra Summit 2026 session playlist here &lt;A class="lia-external-url" href="http://%20https://www.youtube.com/playlist?list=PLjt5SKzX1iI8con7FJDB56G6hHqxGm7ki" target="_blank"&gt;Microsoft Azure Infra Summit 2026&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;Cheers!&lt;/P&gt;
&lt;P&gt;Pierre Roman&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 05 Aug 2026 07:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/agentic-migrations-and-modernization-how-the-azure-migrate-agent/ba-p/4542513</guid>
      <dc:creator>Pierre_Roman</dc:creator>
      <dc:date>2026-08-05T07:00:00Z</dc:date>
    </item>
    <item>
      <title>From Alert to Resolved: Building a Self-Healing Azure Platform with SRE Agent</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/from-alert-to-resolved-building-a-self-healing-azure-platform/ba-p/4542491</link>
      <description>&lt;P&gt;Hello Folks!&lt;/P&gt;
&lt;P&gt;It’s 3 a.m. Your phone lights up. A critical workload that spans multiple clouds is on fire, ownership is fuzzy, the alert routed to the wrong team first, and now it’s your problem. You sit up in bed, cold and groggy, and start the ritual. Open the runbook. Pull logs from one place. Pull metrics from another. Stare at three dashboards. None of them tell the whole story. So you build a theory. The theory is wrong. The clock keeps ticking. The customer impact keeps climbing. Every wrong turn costs you time, context, and confidence.&lt;/P&gt;
&lt;P&gt;That is the scene Lee Oommen opened with at MAIS14, and it is the reason Azure SRE Agent exists. In this session, Lee walks through the four classic SRE pain points and shows how an agentic operations platform compresses MTTR from hours to minutes. I am going to unpack what he showed, why it matters for IT pros, and how to get your hands on it.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;DIV class="lia-embeded-content" contenteditable="false"&gt;&lt;IFRAME src="https://www.youtube.com/embed/IGpxUe8iEMQ?si=_WqmEA4R-6W2k4tR" width="100%" title="YouTube video player" allowfullscreen="allowfullscreen" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" frameborder="0" style="aspect-ratio: 16/9; height: auto;" sandbox="allow-scripts allow-same-origin allow-forms"&gt;
&lt;/IFRAME&gt;&lt;/DIV&gt;
&lt;H2&gt;Why IT Pros Should Care&lt;/H2&gt;
&lt;P&gt;If you carry a pager, write runbooks, or get pulled into post-mortems, this one is for you.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;The clock is the enemy, not the incident.&lt;/STRONG&gt; Lee said it plainly: there is almost always an expert who can fix the problem. The real damage comes from the minutes spent finding that expert and reconstructing context.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Dashboards lie in both directions.&lt;/STRONG&gt; False positives create alert fatigue. False negatives let the customer call you before your monitors do. Neither outcome is acceptable.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;RCAs take weeks because the answer never lives in one layer.&lt;/STRONG&gt; Infrastructure, network, deployments, dependencies, databases, app code. You need someone, or something, that can correlate across all of them in one pass.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;You did not become an SRE to be a dashboard watcher.&lt;/STRONG&gt; Toil is the work that holds back the people who should be designing reliability into the next generation of services.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;What is the SRE Agent&lt;/H2&gt;
&lt;P&gt;Let’s start with what it is not. It is not a dashboard. It is not a monitoring tool. It is not a chatbot.&lt;/P&gt;
&lt;P&gt;Azure SRE Agent is an end-to-end agentic operations platform. Think of it as a senior SRE who sits inside your team, works 24 by 7, never gets tired, never misses a signal, and is fluent in your stack.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Reasons over telemetry, not just text.&lt;/STRONG&gt; It pulls metrics, logs, traces, deployment history, and activity logs, then correlates across them.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Takes governed actions.&lt;/STRONG&gt; Every action runs inside the permission boundary you define. You decide whether the agent proposes a fix, asks for approval, or acts autonomously.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Authors RCAs in minutes, not weeks.&lt;/STRONG&gt; It traces the root cause in a single flow across the entire stack and produces the report immediately after remediation.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Remembers.&lt;/STRONG&gt; It captures organizational memory from every incident, every chat, and every scheduled task, then applies it to the next investigation.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Lee called it the operations half of DevOps, and that framing stuck with me. We have automated build and deploy. The operations side has been stuck in toil. SRE Agent closes that loop.&lt;/P&gt;
&lt;H2&gt;From Alert to Resolved (the workflow)&lt;/H2&gt;
&lt;P&gt;Lee demonstrated the full loop live. Here is what it looks like end to end.&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;Detect.&lt;/STRONG&gt; The agent integrates with Azure Monitor, PagerDuty, or ServiceNow. When an alert fires, the agent acknowledges it within seconds. The human does not have to wake up cold.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Investigate.&lt;/STRONG&gt; The agent runs diagnostics in parallel across the connected resources. It queries Log Analytics, App Insights, Azure Monitor metrics, and any third-party observability tools you wired in through MCP connectors.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Correlate.&lt;/STRONG&gt; It uses distributed tracing, cross-workspace KQL queries, and time-based signal alignment to connect dots across services that do not even share trace IDs. It also checks past incidents in memory to see if this looks familiar.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Diagnose.&lt;/STRONG&gt; It produces a root cause analysis with the relevant evidence linked inline. No more reconstruction exercise across multiple teams.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Propose or Act.&lt;/STRONG&gt; Based on your run mode and the permissions granted, the agent either proposes a fix and waits for approval, or executes the remediation autonomously. Lee demonstrated both. He set up a bad slot swap on Azure App Service, generated HTTP 500 errors, watched the agent acknowledge the alert, investigate, identify the bad slot, ask for permission, and then perform the slot swap to restore the service.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Close the loop.&lt;/STRONG&gt; The agent files a GitHub or Azure DevOps issue with full context, opens a pull request with proposed code changes when appropriate, and writes a session insights summary you can review.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Three modes to interact with it:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Interactive.&lt;/STRONG&gt; Chat with the agent like a copilot. Most customers start here to build trust.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Reactive.&lt;/STRONG&gt; Event-driven. The agent reacts to incidents from Azure Monitor, PagerDuty, or ServiceNow. One agent per incident platform.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Proactive.&lt;/STRONG&gt; Scheduled tasks that run every five minutes, every hour, daily, or weekly. Certificate health audits, well-architected framework assessments, cost optimization sweeps, compliance checks. Lee showed a scheduled task that flagged a certificate expiring in 50 days before it could ever fire an alert.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Real-World Value&lt;/H2&gt;
&lt;P&gt;This is where the conversation gets practical. A few things from Lee’s demo and the live Q&amp;amp;A that I want to call out.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Multi-cloud reality.&lt;/STRONG&gt; SRE Agent lives in Azure but is not limited to Azure. Custom runbooks, Python execution, MCP servers, and connectors let it orchestrate across AWS, GCP, and on-premises. Treat it as the central SRE brain.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Your data stays yours.&lt;/STRONG&gt; Each agent gets a dedicated data store in your subscription and resource group. Memory, knowledge, threads, and session insights live in your chosen region. Nothing is used to train the model provider. Encryption at rest, TLS 1.2 in transit, Azure RBAC, managed identity, customer-managed keys all apply.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Identity boundary you already know.&lt;/STRONG&gt; The agent uses standard Azure managed identity. Grant the identity RBAC on any cross-subscription resource it needs to reach. Least privilege still applies.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Region availability.&lt;/STRONG&gt; At session time, agents can be deployed in EastUS2, Sweden Central, and Australia East. The list is updating roughly monthly. Canada is coming. An agent in one region can act on resources globally, but if you have data residency rules, deploy the agent inside the same jurisdiction.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Private endpoints today.&lt;/STRONG&gt; If your Log Analytics Workspace or databases are fully locked down behind private endpoints with public access disabled, the agent currently needs a VNet-integrated Azure Function as a proxy. Microsoft is actively working on injecting agents directly into private networks.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Memory is the multiplier.&lt;/STRONG&gt; A principal engineer is more valuable than a junior engineer because of pattern recognition. SRE Agent captures that pattern recognition for the whole team, every time it investigates.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Getting Started&lt;/H2&gt;
&lt;P&gt;The pattern is simple, and Lee summarized it cleanly: you teach the tool, you make the connections, and it works for you.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Provision the agent.&lt;/STRONG&gt; Go to sre.azure.com or the Azure portal, pick a subscription and resource group, pick a region, and stand it up. Takes a few minutes.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Onboard it like a new engineer.&lt;/STRONG&gt; Tell it about your team, your workloads, and your procedures. Upload runbooks, troubleshooting guides, wikis, and architecture docs to the knowledge base. If you do not have these documents, ask the agent to draft them for you.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Connect your observability stack.&lt;/STRONG&gt; Azure Monitor, Log Analytics, App Insights are wired in by default. Add third-party tools through MCP connectors.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Wire in your incident platform.&lt;/STRONG&gt; Azure Monitor, PagerDuty, or ServiceNow. One agent per platform.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Grant code access.&lt;/STRONG&gt; Connect your GitHub or Azure DevOps repositories so the agent can reason over application code, propose fixes, and open pull requests.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Pick your run mode.&lt;/STRONG&gt; Start in interactive mode while you build trust. Move to approval-gated reactive mode. Graduate to autonomous mode on safe operations once you have the audit trail you trust.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Resources&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/sre-agent/" target="_blank"&gt;Azure SRE Agent documentation on Microsoft Learn&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="http://%20https://sre.azure.com/docs/" target="_blank"&gt;Azure SRE Agent product docs&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="http://%20https://sre.azure.com/docs/get-started/" target="_blank"&gt;Get Started guide&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/sre-agent/incident-response" target="_blank"&gt;Automate incident response&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="http://%20https://github.com/microsoft/sre-agent" target="_blank"&gt;Official Microsoft SRE Agent GitHub repository (issues, labs, resources)&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Watch the Rest of the Summit&lt;/H2&gt;
&lt;P&gt;If you found this useful, the rest of the Azure Infra Summit 2026 is packed with sessions on identity, AKS, deployment, storage, networking, and resiliency. Grab the full playlist here and binge what is relevant to your stack:&lt;A class="lia-external-url" href="http:// https://www.youtube.com/playlist?list=PLjt5SKzX1iI8con7FJDB56G6hHqxGm7ki&amp;nbsp;" target="_blank"&gt;Microsoft Azure Infra Summit 2026&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;Big thanks to Lee Oommen for walking us through this. The 3 a.m. pager scenario is something every one of us has lived, and seeing an agent take the first hour of that incident off your plate is a tangible win.&lt;/P&gt;
&lt;P&gt;Cheers!&lt;/P&gt;
&lt;P&gt;Pierre Roman&lt;/P&gt;</description>
      <pubDate>Tue, 04 Aug 2026 07:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/from-alert-to-resolved-building-a-self-healing-azure-platform/ba-p/4542491</guid>
      <dc:creator>Pierre_Roman</dc:creator>
      <dc:date>2026-08-04T07:00:00Z</dc:date>
    </item>
    <item>
      <title>Designing Azure Networks That Scale: From Small Deployments to Enterprise-Grade</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/designing-azure-networks-that-scale-from-small-deployments-to/ba-p/4542489</link>
      <description>&lt;P&gt;Hello Folks!&lt;/P&gt;
&lt;P&gt;If you have ever spent a long afternoon untangling overlapping CIDR ranges, chasing down a broken VNet peering, or trying to remember which UDR points to which firewall, this MAIS 2026 session is going to feel uncomfortably familiar. Jon Ormond (Principal PM, Azure Networking) brought along Jay Li and Jeff Lovett from the Azure Networking team to walk through what actually happens when an Azure network grows from a handful of VNets into a real enterprise estate, and where most teams hit the wall.&lt;/P&gt;
&lt;P&gt;The headline they kept coming back to is simple. Azure networks do not usually fail because they were built wrong on day one. They fail because they did not evolve fast enough. Scale is not a smooth ramp. It is a step function, and every step adds an order of magnitude of complexity.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;DIV class="lia-embeded-content" contenteditable="false"&gt;&lt;IFRAME src="https://www.youtube.com/embed/ZTcuDbRPtjA?si=GUB4JlRs4qySlCTb" width="100%" title="YouTube video player" allowfullscreen="allowfullscreen" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" frameborder="0" style="aspect-ratio: 16/9; height: auto;" sandbox="allow-scripts allow-same-origin allow-forms"&gt;
&lt;/IFRAME&gt;&lt;/DIV&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;Why IT Pros Should Care&lt;/H2&gt;
&lt;P&gt;You may be running three VNets today. That is fine. But the day a second team shows up, or you cross into a second region, or somebody asks for hybrid connectivity to the datacenter, your operating model changes whether you planned for it or not. The session is built around two pivots every growing Azure environment hits:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Management and control inside Azure (VNets, peerings, routes, security rules).&lt;/LI&gt;
&lt;LI&gt;Connectivity and hybrid (VPN, ExpressRoute, Virtual WAN, reliability).&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Both of those break quietly. By the time you notice, you are already firefighting drift, broken peerings, or unpredictable latency from on-prem.&lt;/P&gt;
&lt;P&gt;Bottom line, here is what you take away:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Design for the next stage, not the one you are in.&lt;/LI&gt;
&lt;LI&gt;Put the management layer in before complexity outpaces manual effort.&lt;/LI&gt;
&lt;LI&gt;Treat reliability as a design choice, not an afterthought.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Start Small, Plan to Grow&lt;/H2&gt;
&lt;P&gt;One VNet, one subnet, one workload. Nothing wrong with that. You can manage it with the portal, a spreadsheet for CIDR tracking, and a calm heart.&lt;/P&gt;
&lt;P&gt;The problem is that the jump from “one VNet” to “a few VNets across teams” is not gradual. As soon as you have a second team that needs isolation, you are into hub and spoke territory. Ten spokes feels manageable. Fifty spokes across multiple subscriptions does not. And by the time you hit a hundred, the spreadsheet is a liability.&lt;/P&gt;
&lt;P&gt;Jay made the case that the smartest move at small scale is not to stay manual until it hurts. It is to put Azure Virtual Network Manager (AVNM) in early, even if you only have three VNets. AVNM lets you declare intent once and let the platform handle the rest:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;IP address management (IPAM) so new spokes get non-overlapping CIDRs automatically.&lt;/LI&gt;
&lt;LI&gt;Network groups with tag-based dynamic membership so VNets land in the right group the moment they exist.&lt;/LI&gt;
&lt;LI&gt;Connectivity (hub and spoke or mesh) without hand-built peerings.&lt;/LI&gt;
&lt;LI&gt;Security admin rules pushed centrally across the estate.&lt;/LI&gt;
&lt;LI&gt;Routing intent so traffic flows through the right firewall by default.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The honest tradeoff: AVNM is one more thing to learn and operate, and it adds cost. The counter-question Jay kept asking is, “What is the cost of drift?” One overlapping CIDR or one missing UDR at 100 VNets can cascade into an outage that takes days to unwind. That is the real tradeoff.&lt;/P&gt;
&lt;H2&gt;Mid-Stage Patterns: Hub and Spoke, Peering, and the First Cracks&lt;/H2&gt;
&lt;P&gt;The hub and spoke topology is the workhorse of Azure networking and the pattern the Cloud Adoption Framework recommends for most enterprises. It centralises shared services (firewall, DNS, ExpressRoute and VPN gateways, Private DNS zones) in a hub VNet, and connects spoke VNets through peerings.&lt;/P&gt;
&lt;P&gt;Where teams get into trouble at this stage:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Peering sprawl.&lt;/STRONG&gt; Every new spoke needs a peering, sometimes two if you want transitive paths. Doing this by hand across subscriptions is where human error lives.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Route table drift.&lt;/STRONG&gt; UDRs copied from spoke to spoke get out of sync. One spoke routes through the firewall, another bypasses it. Now you have a compliance problem.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Security rule drift.&lt;/STRONG&gt; NSGs and security policies start as a copy paste exercise and end as a forensic exercise.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;CIDR collisions.&lt;/STRONG&gt; “Just give me a /24” turns into a multi day investigation when the new spoke overlaps with on-prem.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Jay’s point on this was sharp. The mistake is not the topology. Hub and spoke is the right pattern. The mistake is staying manual on top of it. AVNM network groups let you say, “any VNet tagged environment=production joins the production group, gets the production security baseline, peers to the production hub, and inherits the routing intent that sends east-west traffic through the firewall.” No tickets, no copy paste, no drift.&lt;/P&gt;
&lt;P&gt;If you are already deployed via Azure Landing Zones (ALZ) with Bicep or Terraform, AVNM is not a replacement, it is another construct in your template. As Jon put it in the chat, it is “just another object” in your ALZ, and the two layers work together rather than competing.&lt;/P&gt;
&lt;H2&gt;Enterprise Scale: Virtual WAN, Segmentation, and Governance&lt;/H2&gt;
&lt;P&gt;At some point hub and spoke stops scaling cleanly. You start adding regions. Branch offices show up. You need SD-WAN integration, more than 30 IPsec tunnels, or transitive routing between VPN and ExpressRoute. That is when Microsoft pushes you toward Azure Virtual WAN.&lt;/P&gt;
&lt;P&gt;Virtual WAN is a Microsoft managed global transit network. You deploy regional virtual hubs and connect everything (Azure VNets, branches, remote users, ExpressRoute circuits) into them with consistent routing and security. The trade up is real:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Any to any connectivity by default.&lt;/STRONG&gt; Hub to hub mesh is built in.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Routing intent and policies&lt;/STRONG&gt; for centralised internet egress and east-west inspection through Azure Firewall or a partner NVA in a secured hub.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Branch scale.&lt;/STRONG&gt; Tens or hundreds of sites stop being a custom integration project.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Operational simplification.&lt;/STRONG&gt; Microsoft owns the hub control plane so you stop babysitting peerings.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;For hybrid connectivity itself, Jeff walked the curve every customer travels:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;VPN Gateway&lt;/STRONG&gt; is the on-ramp. Cheap, fast to stand up, good enough until public internet latency, throughput, or regulatory requirements force a change.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;ExpressRoute circuits&lt;/STRONG&gt; give you dedicated bandwidth from 50 Mbps to 100+ Gbps, with predictable performance and over 200 service providers worldwide.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Scalable ExpressRoute virtual network gateways&lt;/STRONG&gt; grow and shrink with usage, so you deploy once and stop re-architecting every time traffic changes.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;ExpressRoute Metro&lt;/STRONG&gt; is the headliner. Same price as a standard circuit, but the redundant device lives in a second, physically distinct co-location facility across town. Building fire, flood, or power outage in one site, and your traffic keeps flowing.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Multiple circuits&lt;/STRONG&gt; are still on the table when “this cannot fail” actually means it cannot fail.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Honest tradeoff on Virtual WAN: it is opinionated, Microsoft managed, and you give up some of the granular control you have in a customer managed hub. For most enterprises that is a win. For the few with very specific routing requirements or heavy NVA investments, traditional hub and spoke with Azure Route Server can still be the right call. The CAF guidance lays this out in detail.&lt;/P&gt;
&lt;H2&gt;Getting Started&lt;/H2&gt;
&lt;P&gt;If you take one thing from this session, take this. Design for the next stage. Three concrete moves:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Stand up AVNM now, even at small scale.&lt;/STRONG&gt; Declare your intent for IPAM, connectivity, security, and routing once. Let new VNets inherit it.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Pick your topology with eyes open.&lt;/STRONG&gt; Hub and spoke for customer managed control, Virtual WAN for Microsoft managed global transit at scale. The CAF decision tree is the right starting point.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Plan hybrid for failure, not for the sunny day.&lt;/STRONG&gt; ExpressRoute with Metro by default. Multiple circuits for the workloads that genuinely cannot go down. Test the failover.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Resources&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/virtual-network/network-manager-overview" target="_blank"&gt;Azure Virtual Network Manager overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/expressroute/expressroute-introduction" target="_blank"&gt;Azure ExpressRoute introduction&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/expressroute/expressroute-about-virtual-network-gateways" target="_blank"&gt;About ExpressRoute virtual network gateways&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/vpn-gateway/vpn-gateway-about-vpngateways" target="_blank"&gt;About Azure VPN Gateway&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/virtual-wan/virtual-wan-about" target="_blank"&gt;About Azure Virtual WAN&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/architecture/networking/architecture/hub-spoke" target="_blank"&gt;Hub-spoke network topology in Azure&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/cloud-adoption-framework/ready/azure-best-practices/define-an-azure-network-topology" target="_blank"&gt;Define an Azure network topology (Cloud Adoption Framework)&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/cloud-adoption-framework/ready/azure-best-practices/virtual-wan-network-topology" target="_blank"&gt;Virtual WAN network topology in an Azure landing zone&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Watch the Rest of the Summit&lt;/H2&gt;
&lt;P&gt;This was one of many great sessions at the Microsoft Azure Infra Summit 2026. If you want to catch the keynotes, the deep dives on storage and AKS, and everything in between, the full playlist is here:&lt;/P&gt;
&lt;P&gt;&lt;A href="https://www.youtube.com/playlist?list=PLjt5SKzX1iI8con7FJDB56G6hHqxGm7ki" target="_blank"&gt;Microsoft Azure Infra Summit 2026 Playlist&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;Big thanks to Jon Ormond for moderating, and to Jay Li and Jeff Lovett for the practical, no-fluff walk through what actually breaks at scale and how to design ahead of it.&lt;/P&gt;
&lt;P&gt;Cheers!&lt;/P&gt;
&lt;P&gt;Pierre Roman&lt;/P&gt;</description>
      <pubDate>Mon, 03 Aug 2026 07:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/designing-azure-networks-that-scale-from-small-deployments-to/ba-p/4542489</guid>
      <dc:creator>Pierre_Roman</dc:creator>
      <dc:date>2026-08-03T07:00:00Z</dc:date>
    </item>
    <item>
      <title>Digital Takeover/Lockdown</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk/digital-takeover-lockdown/m-p/4543071#M2690</link>
      <description>&lt;P&gt;My web developer has taken over my M365 &amp;amp; Copilot (along with my domain, google workspace, GitHub, Manus.im account, Stripe, Shopify, my financial accounts, and so on) and made has made himself Super admin and/or created an enterprise hierarchy that then turns myself, THE OWNER, into a USER; if he allows me access at all. I have been in an active digital lockout for going on 5 months now. I am at the local library currently trying to find a solution to regaining control. I have been dealing with redirected and limited browsers, blocked accounts, email accounts forwarded and new emails created, and then, even my calls and texts are being forwarded and/or blocked. From taking over my account's admin, a malicious DNS change, using rootkit malware, harmful scripts, injected code, utilizing my API's and tokens to gain access and then lock me out, to deploying autonomous ai agents onto my desktop...It only gets worse! He is also a third-party service provider and has led this persistent attack on my business by hacking EVERY mobile device I have purchased by using my location, Bluetooth, and sharing apps to hack my network and take control. No, this is not a joke. This is my VERY REAL AND VERY CURRENT NIGHTMARE! PLEASE, HELP ME!&lt;/P&gt;&lt;P&gt;&amp;nbsp;-Brittany Hamm&lt;/P&gt;&lt;P&gt;Diary of a Momtrepreneur/Cullmanspaces.com/brittanyhamm.base44.app&lt;/P&gt;</description>
      <pubDate>Sat, 01 Aug 2026 15:30:04 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk/digital-takeover-lockdown/m-p/4543071#M2690</guid>
      <dc:creator>BHamm</dc:creator>
      <dc:date>2026-08-01T15:30:04Z</dc:date>
    </item>
    <item>
      <title>Azure Files, Reimagined: Top-Level Shares with Per-Share Networking, Billing, and Scale</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/azure-files-reimagined-top-level-shares-with-per-share/ba-p/4535079</link>
      <description>&lt;P&gt;Hello Folks!&lt;/P&gt;
&lt;P&gt;If you have ever wrestled with Azure Files inside a storage account, juggling shared RBAC, shared networking, and shared IOPS across a pile of shares that really should not live together, this session is going to address all that. During Microsoft Azure Infra Summit 2026, Vincent Du and Will Gries (both Product Managers on the Azure Files team) walked us through the new Microsoft.FileShares resource provider, a management model that promotes the file share itself to a top-level Azure resource.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;DIV class="lia-embeded-content" contenteditable="false"&gt;&lt;IFRAME src="https://www.youtube.com/embed/IWktcxpru7c?si=STWrMOtHBHxWtFl3" width="100%" title="YouTube video player" allowfullscreen="allowfullscreen" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" frameborder="0" style="aspect-ratio: 16/9; height: auto;" sandbox="allow-scripts allow-same-origin allow-forms"&gt;
&lt;/IFRAME&gt;&lt;/DIV&gt;
&lt;H2&gt;Why IT Pros Should Care&lt;/H2&gt;
&lt;P&gt;For years, file shares lived inside a storage account, and that storage account dictated a lot of decisions for you. If one team needed a private endpoint and another needed a service endpoint, you either compromised or you created another storage account. If one share got hot and consumed all the IOPS, the other shares felt it too. Vincent and Will are on the team that built the new model to remove that compromise.&lt;/P&gt;
&lt;P&gt;Here is what changes for you as an IT pro:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Each file share is its own Azure resource with its own RBAC, networking, billing, IOPS, and throughput.&lt;/LI&gt;
&lt;LI&gt;Per-share cost shows up directly in Azure Cost Management’s per-resource view, no more Excel guesswork.&lt;/LI&gt;
&lt;LI&gt;Encryption in transit is on by default for NFS shares, at no extra cost.&lt;/LI&gt;
&lt;LI&gt;Provisioning is dramatically faster. In their head-to-head demo, 200 shares finished in about 50 seconds on the new model versus about 720 seconds with the classic flow.&lt;/LI&gt;
&lt;LI&gt;A new MCP server lets you create and manage shares from GitHub Copilot in VS Code with natural language.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;In short, the new model trades the storage-account-as-gatekeeper pattern for something that feels a lot more like the rest of Azure (think VMs and disks, where the resource you care about is the resource you actually manage).&lt;/P&gt;
&lt;H2&gt;What Microsoft.FileShares Does, a Technical Overview&lt;/H2&gt;
&lt;P&gt;The new Microsoft.FileShares resource provider lets you deploy a file share without first standing up a storage account. When you go into the Azure portal, search for “File share,” and click create, you fill out a single create blade with the things that actually matter for that share: name, region, redundancy (LRS or ZRS), provisioned capacity, IOPS and throughput, networking, and tags. Microsoft Learn confirms the provisioned capacity range is 32 GiB to 262,144 GiB, and only LRS and ZRS redundancy are available at launch (see the Create a file share doc linked below).&lt;/P&gt;
&lt;P&gt;At GA, the new experience supports NFS 4.1 on the SSD media tier. SMB support, HDD support, customer-managed key encryption at rest, soft delete, and the AKS CSI driver integration are all on the roadmap and called out as the most-requested follow-ups. If you need those features today, the classic file share inside a storage account is still there for you.&lt;/P&gt;
&lt;P&gt;In the portal, Vincent showed off a small but meaningful detail: the icon color changed from blue (classic) to purple (new). It is a small thing, but when you are scanning a resource group, that visual cue saves you a click.&lt;/P&gt;
&lt;H2&gt;How It Works Under the Hood&lt;/H2&gt;
&lt;P&gt;The new model is built on the provisioned v2 billing structure. Microsoft Learn describes provisioned v2 as a billing model where you independently provision storage, IOPS, and throughput, and you pay for what you provision regardless of how much you actually use. This is a real shift from the older provisioned v1 model, where IOPS and throughput were a function of how much storage you provisioned.&lt;/P&gt;
&lt;P&gt;Will walked through the math. In his example, provisioning 14 TiB of storage on v1 gave 17,000 IOPS, about 1.5 GB/s throughput, and a bill of roughly $2,297. Moving to v2 with the exact same numbers was already noticeably cheaper. Then, because v2 lets you tune storage, IOPS, and throughput separately, he provisioned the exact storage he needed with slightly less IOPS and throughput, dropping the bill to roughly a third. For database-hot workloads you can dial IOPS up; for hot archive scenarios you can dial them down to the minimum. That kind of flexibility is genuinely useful.&lt;/P&gt;
&lt;P&gt;Encryption in transit deserves its own callout. The new shares default to encrypted NFS mounts using the AZNFS mount helper. Microsoft Learn explains that AZNFS wraps the NFS connection in a Stunnel-based TLS tunnel using AES-GCM, so you get TLS protection without needing Kerberos or external authentication. The helper installs cleanly on Ubuntu, RHEL, SUSE, Rocky, Oracle Linux, Alma Linux, and Azure Linux. If a workload genuinely cannot use the encrypted mount, you can uncheck the box and fall back to a traditional NFS mount.&lt;/P&gt;
&lt;P&gt;Networking is per share. You can attach a service endpoint or a private endpoint to each individual share, which means you can put a strict private-endpoint-only share next to a service-endpoint share for dev/test, all in the same resource group, without compromise.&lt;/P&gt;
&lt;P&gt;On the request side, classic shares throttle with a fixed window (you can burst, then you are locked out for the rest of the window). The new model uses a token-bucket algorithm (the same one Azure Resource Manager itself uses), which means you get a sustained refill rate. The team also gave you a separate delete bucket, so a big cleanup operation does not starve writes. That detail matters more than it sounds: batch cleanups against the classic model regularly crowd out new share creation.&lt;/P&gt;
&lt;H2&gt;Real-World Value&lt;/H2&gt;
&lt;P&gt;Where does this actually pay off? A few honest scenarios:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Mission-critical and regulated workloads.&lt;/STRONG&gt; A healthcare org with workloads at different sensitivity levels can put strict private-endpoint-only shares next to less sensitive service-endpoint shares without the storage-account ceiling.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Chargeback and showback.&lt;/STRONG&gt; With per-share resources, finance can pull a cost report that lines up to the team or project that owns each share. No more saying “we cannot itemize, the storage account is shared.”&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;High-density tenants.&lt;/STRONG&gt; The classic model effectively caps you at 34 file shares on an SSD provisioned v2 storage account (because of IOPS minimums) and 50 absolute. The new model goes up to 10,000 shares per subscription per region. That is a different game.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Tuned database and analytics shares.&lt;/STRONG&gt; Provisioned v2 lets you right-size IOPS to the workload. As Will showed, that can drop the bill to roughly a third for the right shape of workload.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Faster deployment automation.&lt;/STRONG&gt; A 14x improvement on a 200-share deployment is not a micro-optimization. If you spin up environments for CI, training, or per-customer tenants, that adds up quickly.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The honest tradeoff: today, the new model is NFS-only on SSD. If you need SMB, HDD, customer-managed keys for NFS, or AKS CSI driver support, stay on the classic model for now. The team was upfront about that, and the GA-and-then-iterate roadmap is clear.&lt;/P&gt;
&lt;H2&gt;Getting Started&lt;/H2&gt;
&lt;P&gt;Here is the concrete path:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Register the Microsoft.FileShares and Microsoft.Storage resource providers on your subscription (Subscriptions, Resource providers, Register).&lt;/LI&gt;
&lt;LI&gt;From the Azure portal, search for “File share” in the marketplace and click Create. Pick LRS or ZRS, set the capacity between 32 GiB and 262 TiB, and either accept the recommended IOPS/throughput or set them manually.&lt;/LI&gt;
&lt;LI&gt;On the Advanced tab, leave “Require encryption in transit” enabled (it is on by default) and pick a custom mount name if you want one distinct from the resource name.&lt;/LI&gt;
&lt;LI&gt;On the Networking tab, attach a service endpoint or a private endpoint, per share.&lt;/LI&gt;
&lt;LI&gt;Mount it on your Linux VM with the AZNFS mount helper. The portal generates the exact command for your distribution.&lt;/LI&gt;
&lt;LI&gt;If you live in IaC land, the Microsoft.FileShares ARM and Bicep types are available, and Terraform support is coming.&lt;/LI&gt;
&lt;LI&gt;If you live in AI-assisted dev land, install the Azure MCP server and ask Copilot in VS Code to create a share for you, pointing at an existing VNet.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2&gt;Resources&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/storage/files/create-file-share" target="_blank" rel="noopener"&gt;Create an Azure file share with Microsoft.FileShares&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/storage/files/understanding-billing" target="_blank" rel="noopener"&gt;Understand Azure Files billing (provisioned v1 and v2)&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/storage/files/encryption-in-transit-for-nfs-shares" target="_blank" rel="noopener"&gt;Encryption in Transit for NFS Azure file shares&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/storage/files/files-nfs-protocol" target="_blank" rel="noopener"&gt;NFS file shares in Azure Files (protocol overview)&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/storage/files/" target="_blank" rel="noopener"&gt;Azure Files documentation home&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Keep Learning at the Summit&lt;/H2&gt;
&lt;P&gt;Catch the full &lt;A class="lia-external-url" href="https://www.youtube.com/playlist?list=PLjt5SKzX1iI8con7FJDB56G6hHqxGm7ki" target="_blank" rel="noopener"&gt;Microsoft Azure Infra Summit 2026 session playlist&lt;/A&gt; here.&lt;/P&gt;
&lt;P&gt;Cheers!&lt;/P&gt;
&lt;P&gt;Pierre Roman&lt;/P&gt;</description>
      <pubDate>Thu, 30 Jul 2026 19:57:34 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/azure-files-reimagined-top-level-shares-with-per-share/ba-p/4535079</guid>
      <dc:creator>Pierre_Roman</dc:creator>
      <dc:date>2026-07-30T19:57:34Z</dc:date>
    </item>
    <item>
      <title>Some tools and techniques for hardening Windows Server</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/some-tools-and-techniques-for-hardening-windows-server/ba-p/4539840</link>
      <description>&lt;P&gt;In this post I go over some tools and techniques exist for hardening Windows Server. As always, apply controls according to the server's role, test them against representative workloads before production rollout, document approved exceptions, and maintain tested console and recovery access in case a security control affects management or application compatibility.&lt;/P&gt;
&lt;H2&gt;Apply the role-specific Windows Server 2025 baseline with OSConfig&lt;/H2&gt;
&lt;P&gt;Security baselines turn hundreds of individual security decisions into a consistent, role-aware desired state. Using OSConfig reduces exposure caused by insecure defaults, legacy protocols, inconsistent administrator choices, and configuration drift that attackers can exploit for credential theft, lateral movement, or persistence.&lt;/P&gt;
&lt;P&gt;You can use use OSConfig at build time to apply the Microsoft security baseline that matches the server role: &lt;CODE&gt;SecurityBaseline/WindowsServer/2025/MemberServer&lt;/CODE&gt;, &lt;CODE&gt;SecurityBaseline/WindowsServer/2025/DomainController&lt;/CODE&gt;, or &lt;CODE&gt;SecurityBaseline/WindowsServer/2025/WorkgroupMember&lt;/CODE&gt;. The baseline contains more than 300 settings covering network exposure, credentials, lateral movement, persistence resistance, and auditing. It can be managed through PowerShell, Windows Admin Center, or Azure Policy for Azure Arc-enabled servers.&lt;/P&gt;
&lt;P&gt;You can keep OSConfig drift control enabled so unauthorized or accidental changes are detected and corrected. Pilot the baseline with each workload, record required exceptions, and manage those exceptions centrally rather than weakening the baseline broadly.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Identify the server's role and the management authority that will own its settings, install the OSConfig PowerShell module, review the matching scenario, and apply it first to a representative test server. Validate application and management access, deploy the scenario in controlled rings, schedule and complete the restart required after applying the baseline, verify the desired configuration and compliance results, enable drift control, and record any approved exceptions and recovery procedures.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; A baseline can disrupt legacy applications, authentication methods, network flows, or management tools that depend on weaker settings. Drift control can also reverse intentional emergency changes if they aren't recorded through the correct authority, so staged testing, documented exceptions, and tested recovery access are essential.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation on Learn:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/windows-server/security/osconfig/osconfig-overview" target="_blank" rel="noopener noreferrer"&gt;OSConfig security configuration for Windows Server&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows-server/security/osconfig/osconfig-how-to-configure-security-baselines" target="_blank" rel="noopener noreferrer"&gt;Deploy Windows Server 2025 security baselines with OSConfig&lt;/A&gt; | &lt;A href="https://github.com/microsoft/osconfig/tree/main/security" target="_blank" rel="noopener noreferrer"&gt;OSConfig security settings repository&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;Use Secured-core hardware and enable platform security&lt;/H2&gt;
&lt;P&gt;Windows Server secured-core combines hardware, firmware, virtualization, and operating-system protections to establish trust before Windows starts and preserve that trust while it runs. These controls mitigate bootkits, malicious or vulnerable kernel drivers, direct memory access attacks, firmware tampering, and attempts to extract credentials from the operating system.&lt;/P&gt;
&lt;P&gt;You deploy on hardware or virtual machines that support TPM 2.0, UEFI Secure Boot, virtualization-based security, DMA protection, and the other Secured-core requirements. Enable the OSConfig &lt;CODE&gt;SecuredCore&lt;/CODE&gt; scenario and verify that Credential Guard, hypervisor-protected code integrity, kernel protections, and the signed boot chain are active.&lt;/P&gt;
&lt;P&gt;You need to keep system firmware, TPM firmware, hypervisor components, and hardware drivers current. Test older drivers before enabling enforcement because incompatible kernel drivers can prevent security features from activating or can affect boot reliability.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Confirm that the physical server or virtual-machine platform meets the Secured-core requirements, update firmware and drivers, and enable TPM 2.0, Secure Boot, virtualization extensions, and DMA or IOMMU protection in the platform configuration. Apply the OSConfig &lt;CODE&gt;SecuredCore&lt;/CODE&gt; scenario or configure the features through Windows Admin Center, restart as required, verify that each protection is active, record evidence of the hardware capabilities and running Windows protections rather than only the assigned policy, and monitor for driver or workload compatibility issues.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; Secured-core features require compatible hardware, firmware, hypervisors, and signed drivers, which can increase procurement costs or rule out older systems. Virtualization-based protections can introduce a workload-dependent performance impact, and incompatible drivers or firmware can cause application failures, feature activation problems, or difficult boot recovery.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/windows-server/security/secured-core-server" target="_blank" rel="noopener noreferrer"&gt;What is Secured-core server?&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows-server/security/configure-secured-core-server" target="_blank" rel="noopener noreferrer"&gt;Configure Secured-core server&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows-server/get-started/hardware-requirements#secured-core-server-requirements" target="_blank" rel="noopener noreferrer"&gt;Windows Server 2025 secured-core hardware requirements&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;Deploy Server Core and minimize installed components&lt;/H2&gt;
&lt;P&gt;Attack-surface reduction removes code, services, interfaces, and utilities that an attacker could exploit or misuse after gaining access. A minimal Server Core deployment lowers the number of vulnerabilities that require patching and reduces opportunities for interactive attacks, malicious browsing, persistence, and abuse of unnecessary administrative tools.&lt;/P&gt;
&lt;P&gt;Install Server Core unless a supported workload specifically requires Desktop Experience. Server Core has a smaller local interface and component footprint, reducing exposed code, maintenance requirements, and opportunities for interactive misuse. Windows Server 2025 can't convert between Server Core and Server with Desktop Experience after installation, so changing this choice later requires a clean installation.&lt;/P&gt;
&lt;P&gt;Install only the roles, features, management agents, and application components required for the server's purpose. Remove obsolete utilities and unused software, avoid browsing the web from servers, and disable unnecessary services only after confirming role and application dependencies. Where practical, dedicate each server to a single security or workload role.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Confirm that the workload and vendor support Server Core, formally record the installation-option decision before deployment, select Server Core during installation, and define the minimum roles, features, agents, and software required for the server's purpose. Install only those components, configure remote management and recovery access, remove or disable unused components after dependency testing, and verify that the application, monitoring, backup, patching, and support processes still function.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; Server Core can make local troubleshooting less familiar and increases reliance on remote management, automation, and command-line skills. Some vendor applications, support tools, or administrators require Desktop Experience, and removing roles or disabling services without dependency testing can break workloads, monitoring, backup, or recovery operations.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/windows-server/administration/server-core/what-is-server-core" target="_blank" rel="noopener noreferrer"&gt;What is the Server Core installation option?&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows-server/get-started/install-options-server-core-desktop-experience" target="_blank" rel="noopener noreferrer"&gt;Server Core and Desktop Experience installation options&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;Ensure rapid patching and continuous vulnerability management&lt;/H2&gt;
&lt;P&gt;Patching and vulnerability management identify and close known weaknesses before attackers can reliably exploit them. This practice reduces exposure to remote-code execution, privilege escalation, ransomware, vulnerable drivers, compromised third-party components, and attacks that target publicly documented vulnerabilities soon after disclosure. Make sure you are aware if any of your server workloads have not got the latest security updates deployed&lt;/P&gt;
&lt;P&gt;Maintain an inventory of operating-system, application, driver, firmware, and management-agent versions. Use deployment rings to test updates quickly, meet defined remediation deadlines, install out-of-band security updates when required, and monitor update compliance and pending restarts. Azure Update Manager can provide centralized assessment and orchestration for Azure and Azure Arc-enabled servers.&lt;/P&gt;
&lt;P&gt;Use Microsoft Defender Vulnerability Management or an equivalent platform to discover exposures, prioritize remediation by exploitability and business impact, and verify that fixes actually remove the vulnerability. Patching Windows while leaving internet-facing applications, drivers, or firmware obsolete does not adequately harden the server.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Inventory servers and every supported update source, define remediation deadlines and deployment rings, and configure Azure Update Manager or another actively developed orchestration platform. Windows Server Update Services remains supported and available but is deprecated and should be treated as a legacy option rather than the preferred platform for a new long-term design. Run vulnerability assessments, prioritize exposed and actively exploited weaknesses, test updates, deploy them with coordinated reboots, verify compliance after installation, and maintain rollback and exception procedures.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; Updates can require reboots, consume maintenance windows, or introduce application, driver, and performance regressions. Vulnerability scanners and management agents also consume resources and can generate false positives, so organizations need test rings, rollback procedures, maintenance coordination, and a risk-based process for temporary deferrals.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/azure/update-manager/overview" target="_blank" rel="noopener noreferrer"&gt;Azure Update Manager overview&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/azure/azure-arc/servers/cloud-native/patch-management" target="_blank" rel="noopener noreferrer"&gt;Cloud-native patch management for Azure Arc-enabled servers&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/defender-vulnerability-management/defender-vulnerability-management" target="_blank" rel="noopener noreferrer"&gt;Microsoft Defender Vulnerability Management&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows-server/get-started/removed-deprecated-features-windows-server#features-no-longer-in-development" target="_blank" rel="noopener noreferrer"&gt;Deprecated Windows Server features&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;Implement Microsoft Defender Antivirus and endpoint detection and response&lt;/H2&gt;
&lt;P&gt;Antivirus and endpoint detection and response combine prevention with behavioral monitoring and investigation. They mitigate malicious files, ransomware, web and network-delivered payloads, suspicious process activity, persistence mechanisms, credential theft, and attacks that evade simple signature-based detection.&lt;/P&gt;
&lt;P&gt;Run Microsoft Defender Antivirus in active mode unless a documented and tested security architecture requires another antimalware product. Enable real-time protection, behavior monitoring, cloud-delivered protection, automatic sample submission, and frequent security-intelligence updates. Use the OSConfig &lt;CODE&gt;Defender/Antivirus/WindowsServer/2025&lt;/CODE&gt; scenario as the recommended Server 2025 configuration starting point.&lt;/P&gt;
&lt;P&gt;Onboard servers to Microsoft Defender for Endpoint or Microsoft Defender for Servers for endpoint detection and response, investigation, and centralized visibility. Enable tamper protection and keep exclusions narrow, workload-specific, and regularly reviewed; broad path, process, or extension exclusions create useful hiding places for attackers.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Confirm licensing, connectivity, proxy, and third-party antivirus requirements, then apply the OSConfig Defender Antivirus scenario or an equivalent centrally managed policy. Enable real-time, behavior, cloud-delivered, sample-submission, and tamper protections; onboard the server to Defender for Endpoint or Defender for Servers; review Microsoft's built-in, automatic server-role, and workload-specific exclusions before adding any manual exclusion; verify sensor health, signature currency, alert delivery, and investigation access; and periodically confirm that every manual exclusion remains necessary.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; Real-time scanning and endpoint telemetry can add CPU, memory, disk I/O, network, and licensing costs, particularly on high-throughput workloads. False positives or quarantine actions can interrupt services, while cloud-delivered capabilities can raise connectivity, privacy, or data-residency considerations; performance exclusions must therefore be tested and kept narrowly scoped.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/defender-endpoint/microsoft-defender-antivirus-windows" target="_blank" rel="noopener noreferrer"&gt;Microsoft Defender Antivirus in Windows&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/defender-endpoint/microsoft-defender-endpoint-windows" target="_blank" rel="noopener noreferrer"&gt;Microsoft Defender for Endpoint on Windows&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/defender-endpoint/microsoft-defender-antivirus-exclusions-overview" target="_blank" rel="noopener noreferrer"&gt;Defender Antivirus exclusions&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/defender-endpoint/tamper-resiliency" target="_blank" rel="noopener noreferrer"&gt;Protect against security-setting tampering&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;Deploy attack surface reduction and network protection on your Windows Server workloads&lt;/H2&gt;
&lt;P&gt;Attack surface reduction rules prevent high-risk behaviors rather than waiting for a specific malicious file to be identified, while network protection blocks access to known or suspicious destinations. Together they mitigate ransomware, credential theft, malicious scripts, abuse of trusted tools, vulnerable drivers, command-and-control traffic, and payload delivery.&lt;/P&gt;
&lt;P&gt;Configure Microsoft Defender attack surface reduction rules to block common behaviors used by ransomware, credential theft, malicious scripts, vulnerable signed drivers, and executable content. Begin with audit or warning mode, review telemetry for legitimate workload dependencies, create narrowly scoped exclusions, and then move suitable rules to block mode on a defined schedule.&lt;/P&gt;
&lt;P&gt;Enable network protection where supported to prevent processes from reaching malicious or untrusted destinations. Manage these controls centrally through Group Policy, Microsoft Defender for Endpoint security settings management, or another supported policy platform.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Inventory server workloads and make a rule-by-rule applicability decision for each server role rather than reusing a generic workstation ASR profile unchanged. Create a centrally managed ASR and network-protection policy that initially uses audit or warning mode, collect and review events, confirm business-critical dependencies, create narrowly scoped exclusions, move applicable rules to block mode through deployment rings, verify that protected applications remain functional, and continuously review detections and exception use.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; ASR rules can block legitimate automation, administrative tools, installers, scripts, or line-of-business applications that exhibit high-risk behavior. Audit mode can produce substantial telemetry, and broad exclusions can undermine the protection, so successful deployment requires workload testing, event review, careful exception design, and ongoing tuning.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/defender-endpoint/attack-surface-reduction-overview" target="_blank" rel="noopener noreferrer"&gt;Attack surface reduction capabilities&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/defender-endpoint/evaluate-mdav-using-gp" target="_blank" rel="noopener noreferrer"&gt;Evaluate Microsoft Defender Antivirus and ASR rules using Group Policy&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;Allow only trusted code with App Control for Business&lt;/H2&gt;
&lt;P&gt;Application control changes the execution model from allowing everything except known malware to allowing only code that satisfies an approved policy. This technique mitigates unknown malware, ransomware, unauthorized administrative utilities, malicious scripts, unapproved drivers, and opportunistic payloads that antivirus has not yet classified.&lt;/P&gt;
&lt;P&gt;Use App Control for Business to define which executables, scripts, installers, libraries, and drivers may run. Windows Server 2025 includes OSConfig scenarios for Microsoft's default policy and application blocklist. Start in audit mode, collect Code Integrity event ID 3076, create required supplemental allow policies, and move to enforcement only after representative workload testing. There are some good GUI tools written by MVPs published on GitHub that make this very easy. &lt;A class="lia-external-url" href="https://github.com/HotCakeX/Harden-Windows-Security" target="_blank"&gt;https://github.com/HotCakeX/Harden-Windows-Security&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;Monitor blocked-code event ID 3077 after enforcement and maintain a controlled process for policy updates and emergency recovery. Application allowlisting is substantially stronger than relying only on malware signatures because unapproved code is prevented from running even when it has not yet been classified as malicious.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Verify that the device is running a production-signed Windows Server 2025 build because the OSConfig default policy doesn't permit flight-signed binaries. Inventory approved applications, scripts, drivers, publishers, and update mechanisms, then deploy the default policy and application blocklist through OSConfig in audit mode. Collect event ID 3076, build and deploy required supplemental policies, and sign policies only when the additional tamper resistance is required and certificate lifecycle, policy servicing, removal, and offline recovery have been tested. Test application updates and recovery, move the policy to enforcement in stages, and monitor event ID 3077 and policy health after deployment.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; Poorly designed policies can block legitimate applications, updates, scripts, drivers, or boot-critical components and can cause a severe service outage. Maintaining allow policies creates operational overhead, especially for frequently changing software, and approved tools can still be abused, so audit-mode deployment, controlled updates, and offline recovery procedures are necessary. Signed policies provide stronger tamper resistance but are intentionally harder to remove, including during recovery.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/windows-server/security/osconfig/osconfig-how-to-configure-app-control-for-business" target="_blank" rel="noopener noreferrer"&gt;Configure App Control for Business by using OSConfig&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows/security/application-security/application-control/app-control-for-business/appcontrol" target="_blank" rel="noopener noreferrer"&gt;App Control for Business&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;Keep Windows Defender Firewall enabled with restrictive rules&lt;/H2&gt;
&lt;P&gt;A host firewall limits which systems and applications can communicate with the server, even when upstream network controls are absent or bypassed. Restrictive rules reduce exposure to service exploitation, scanning, lateral movement, remote administration abuse, command-and-control traffic, and accidental publication of listening services.&lt;/P&gt;
&lt;P&gt;Enable Windows Defender Firewall on Domain, Private, and Public profiles. Retain the default block for unsolicited inbound traffic and create only the rules required by the server role. Scope rules by program or service, protocol, local port, remote address, interface, and profile rather than creating broad port-based or any-source exceptions.&lt;/P&gt;
&lt;P&gt;Log dropped packets and successful connections where operationally appropriate, centrally monitor policy changes, and review stale rules. Apply explicit outbound restrictions to high-value or tightly controlled servers where feasible, especially when they should communicate with only a small set of update, identity, management, and application endpoints.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Inventory listening services and required inbound and outbound flows, enable the firewall on all profiles, and create narrowly scoped rules for the server role. Decide whether locally created rules may merge with centrally deployed rules for each profile, verify the effective policy on representative servers, remove obsolete or duplicate rules, test application, domain, cluster, backup, and management traffic, enable appropriate logging, deploy the policy centrally, and monitor rule changes and blocked connections before introducing selective outbound restrictions.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; Incorrect firewall rules can interrupt application traffic, clustering, domain operations, monitoring, backup, or remote management and can make diagnosis difficult. Detailed connection logging consumes storage, while restrictive outbound policies require continuous maintenance as service endpoints change, so rules should be documented, tested, and deployed with recovery access.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/windows/security/operating-system-security/network-security/windows-firewall/rules#firewall-rules-recommendations" target="_blank" rel="noopener noreferrer"&gt;Windows Firewall rule recommendations&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows-server/security/osconfig/osconfig-how-to-configure-security-baselines#why-security-baselines-matter" target="_blank" rel="noopener noreferrer"&gt;OSConfig baseline network protections&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;Harden Remote Desktop and remote administration&lt;/H2&gt;
&lt;P&gt;Remote administration exposes privileged authentication and interactive control paths that are attractive targets for brute-force attacks, credential theft, session hijacking, and exploitation of internet-facing services. Gateways, multifactor authentication, encrypted sessions, restricted source networks, and credential isolation reduce the likelihood that a stolen password or exposed management port leads directly to server compromise.&lt;/P&gt;
&lt;P&gt;Disable Remote Desktop Services when it is not required. When it is required, use a VPN or Remote Desktop Gateway, require multifactor authentication and Network Level Authentication, restrict source networks and authorized groups, use trusted TLS certificates, and configure sensible idle and disconnected-session limits. Disable clipboard, drive, printer, port, and device redirection unless the operational need outweighs the data-transfer risk.&lt;/P&gt;
&lt;P&gt;Use Remote Credential Guard only for compatible direct RDP administration of Active Directory-joined targets using Kerberos so credentials aren't sent to the remote host. Remote Credential Guard isn't supported through Remote Desktop Gateway or Remote Desktop Connection Broker. For helpdesk access to a potentially compromised host, use Restricted Admin mode instead of Remote Credential Guard. Never expose TCP port 3389 directly to the internet, and avoid using saved privileged credentials on ordinary administrator workstations.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Disable RDP on servers that don't require it. For brokered or externally initiated access, place RDP behind a VPN or MFA-protected Remote Desktop Gateway and restrict permitted users and source networks. For compatible direct RDP administration of Active Directory-joined targets, configure Remote Credential Guard separately; use Restricted Admin mode for appropriate helpdesk scenarios. Configure Network Level Authentication, trusted TLS certificates, session limits, and required redirection controls, test routine and emergency access, and monitor remote logons and gateway activity.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; Gateways, VPNs, MFA services, and secure administrative hosts add licensing, infrastructure, and support dependencies, and their outage can block legitimate administration. Device-redirection restrictions can hinder support workflows, while Network Level Authentication and Remote Credential Guard have compatibility and delegation limitations; a separately secured emergency access path is therefore required.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/windows-server/remote/remote-desktop-services/rds-plan-mfa" target="_blank" rel="noopener noreferrer"&gt;Plan multifactor authentication for Remote Desktop Services&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows/security/identity-protection/remote-credential-guard" target="_blank" rel="noopener noreferrer"&gt;Remote Credential Guard&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;Enforce least privilege and separate administrative identities&lt;/H2&gt;
&lt;P&gt;Least privilege limits each identity and session to the minimum actions required for its task. Separating standard and privileged accounts constrains the damage caused by phishing, token or password theft, malicious insiders, vulnerable administrative tools, and compromised lower-trust devices, while reducing opportunities for privilege escalation and persistence.&lt;/P&gt;
&lt;P&gt;Give administrators standard user accounts for routine work and separate privileged accounts for administrative duties. Minimize membership of local Administrators, Domain Admins, Enterprise Admins, and other powerful groups; review membership and assigned user rights regularly; and prevent highly privileged identities from signing in to lower-trust servers and workstations.&lt;/P&gt;
&lt;P&gt;Use Just Enough Administration endpoints, Windows Admin Center role-based access control, and time-limited elevation where possible. Delegate specific tasks rather than granting unrestricted interactive or PowerShell access, and maintain separately protected emergency accounts for identity-service outages.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Inventory privileged human accounts, service principals, managed identities, service and automation credentials, scheduled tasks, local group membership, duties, and logon locations. Create separate standard and administrative identities, remove unnecessary standing memberships, delegate tasks through role groups, JEA endpoints, Windows Admin Center RBAC, or time-limited elevation, restrict high-tier logons to secured administrative hosts, test that each role can perform its approved duties, and monitor privileged-group, role, automation, and emergency-account use.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; Designing roles, JEA endpoints, approval processes, and time-limited access requires ongoing engineering and governance. Excessively narrow delegation can delay troubleshooting or incident response, while separate accounts add friction for administrators, so permissions should be tested against real duties and emergency access should remain tightly controlled but usable.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/powershell/scripting/security/remoting/jea/overview?view=powershell-7.6" target="_blank" rel="noopener noreferrer"&gt;Just Enough Administration&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows-server/manage/windows-admin-center/plan/user-access-options" target="_blank" rel="noopener noreferrer"&gt;Windows Admin Center user access options&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/security/privileged-access-workstations/privileged-access-access-model" target="_blank" rel="noopener noreferrer"&gt;Enterprise access model&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;Deploy Windows LAPS for local administrator credentials&lt;/H2&gt;
&lt;P&gt;Windows LAPS replaces shared or manually maintained local administrator passwords with unique, random, automatically rotated credentials. It mitigates password reuse, pass-the-hash attacks, credential dumping, and broad lateral movement in which compromise of one server's local administrator credential grants access to many others.&lt;/P&gt;
&lt;P&gt;Use Windows Local Administrator Password Solution to assign a unique, random, automatically rotated local administrator password to every server. The backup destination depends on join state: Active Directory-only devices can use only Active Directory, Microsoft Entra-only devices can use only Microsoft Entra ID, and hybrid-joined devices can use either destination but not both simultaneously. Tightly restrict and audit password retrieval, configure password history and post-authentication rotation, and monitor policy-processing failures.&lt;/P&gt;
&lt;P&gt;Never reuse a common local administrator password across servers because one compromised password or hash can enable broad lateral movement. OSConfig provides the &lt;CODE&gt;LAPS/WindowsServer/2025/MemberServer&lt;/CODE&gt; scenario for member servers. Workgroup systems can be managed through LAPS for Azure Arc, which Microsoft currently documents as a preview feature. On domain controllers, use Windows LAPS to manage the Directory Services Restore Mode password where appropriate.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Select the password-backup destination permitted by the device's join state, prepare Active Directory or Microsoft Entra ID, and identify the local account to manage. For workgroup systems, evaluate the operational and support implications of the preview LAPS for Azure Arc service before adoption. Configure password length, complexity, age, history, and post-authentication actions through policy or OSConfig; configure DSRM password management for applicable domain controllers; delegate password read and reset permissions to a small approved group; pilot the policy; verify password backup and rotation; test authorized recovery; and monitor LAPS processing and retrieval events.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; LAPS introduces directory, policy, permissions, auditing, and recovery dependencies that must be designed correctly. Scripts or applications that rely on a fixed local password can fail, password rotation can disrupt active sessions or automation, and overly broad rights to retrieve stored passwords can create a new privileged credential repository for attackers to target.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/windows-server/identity/laps/laps-overview" target="_blank" rel="noopener noreferrer"&gt;What is Windows LAPS?&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows-server/security/osconfig/osconfig-how-to-configure-security-baselines#available-security-baseline-scenarios" target="_blank" rel="noopener noreferrer"&gt;OSConfig Windows LAPS scenario&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/azure/osconfig/overview-laps-azure-arc" target="_blank" rel="noopener noreferrer"&gt;LAPS for Azure Arc&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;Replace static service-account passwords with managed service accounts&lt;/H2&gt;
&lt;P&gt;Managed service accounts replace human-managed, long-lived service passwords with complex credentials that Active Directory changes automatically. This reduces exposure to password theft, reuse, weak password selection, expired credentials, secrets embedded in scripts, and persistence based on service accounts whose passwords are rarely rotated.&lt;/P&gt;
&lt;P&gt;Use group managed service accounts for supported Windows services, scheduled tasks, and application pools in Active Directory environments. gMSAs provide automatic password management and reduce the need to store or manually rotate long-lived service credentials.&lt;/P&gt;
&lt;P&gt;Windows Server 2025 also introduces delegated Managed Service Accounts for supported migrations from traditional service accounts. A dMSA binds authentication to approved machine identities, uses managed randomized keys, and disables use of the original service-account password. dMSA deployment requires a discoverable Windows Server 2025 domain controller, and an existing gMSA can't be migrated to a dMSA.&lt;/P&gt;
&lt;P&gt;Grant each gMSA only the logon rights, resource permissions, and password-retrieval scope it requires. Do not make service accounts members of privileged groups unless unavoidable, prohibit interactive sign-in, remove obsolete accounts promptly, and monitor changes to the hosts permitted to retrieve each managed password.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Inventory service identities and application dependencies, then select a gMSA for supported services that can directly use a managed account or evaluate a dMSA for a supported Windows Server 2025 migration from a traditional service account. For a gMSA, confirm Active Directory and key-distribution prerequisites, limit which hosts may retrieve its password, assign only required logon rights, permissions, and service principal names, install and test the account on approved hosts, migrate the service or task, and disable or remove the former static-password account. For a dMSA, confirm a discoverable Windows Server 2025 domain controller and follow the documented migration and rollback process.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; Managed service accounts depend on Active Directory and are not supported by every application, installer, or cross-platform workload. Migration can involve service-principal-name, delegation, permission, and clustering changes, while an overly broad password-retrieval scope allows additional hosts to use the identity. dMSA also requires Windows Server 2025 domain-controller availability and has migration rules that differ from gMSA, so compatibility, rollback, and access boundaries require careful testing.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/entra/architecture/service-accounts-group-managed" target="_blank" rel="noopener noreferrer"&gt;Secure group managed service accounts&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows-server/identity/ad-ds/manage/delegated-managed-service-accounts/delegated-managed-service-accounts-overview" target="_blank" rel="noopener noreferrer"&gt;Delegated Managed Service Accounts overview&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows-server/identity/ad-ds/manage/delegated-managed-service-accounts/delegated-managed-service-accounts-faq" target="_blank" rel="noopener noreferrer"&gt;Delegated Managed Service Accounts FAQ&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;Protect credentials and phase out legacy authentication&lt;/H2&gt;
&lt;P&gt;Credential isolation and modern authentication reduce the value of secrets that an attacker can extract or relay. Credential Guard, LSA protection, Kerberos AES, and retirement of weak authentication mitigate memory scraping, pass-the-hash, pass-the-ticket, NTLM relay, downgrade attacks, and cracking of obsolete password representations.&lt;/P&gt;
&lt;P&gt;Verify that Credential Guard and Local Security Authority protection are active where hardware and workload compatibility permit. Windows Server 2025 enables Credential Guard by default on eligible domain-joined systems that aren't domain controllers, but the state should still be verified and centrally enforced where required. Use Negotiate with Kerberos and modern AES encryption for domain authentication, prevent storage of LM hashes or reversibly encrypted passwords, and keep delegated credentials non-exportable. NTLMv1 is removed in Windows Server 2025, and deprecated NTLMv2 should be treated only as a temporary compatibility fallback rather than an end state.&lt;/P&gt;
&lt;P&gt;Audit NTLM and other legacy authentication dependencies before restricting or disabling them, then remove those dependencies in a controlled sequence. Do not disable legacy protocols blindly on production servers, but do not leave them enabled indefinitely solely because their consumers have not been inventoried.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Confirm hardware and driver support, verify the default Credential Guard state, and use OSConfig or centrally managed policy to enforce Credential Guard and LSA protection where required. Enable NTLM auditing, inventory clients and services using legacy authentication, configure Negotiate and Kerberos AES, update affected service accounts, remediate dependencies, assign an owner and retirement date to every NTLMv2 exception, introduce NTLM restrictions in stages, and monitor authentication failures before broader enforcement.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; Virtualization-based credential protection requires compatible hardware and can have a workload-dependent performance or compatibility impact. Legacy devices, applications, trusts, or service configurations may still depend on NTLM or weaker cryptography, and disabling them without complete auditing can cause widespread authentication failures or outages.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/windows/security/identity-protection/credential-guard/" target="_blank" rel="noopener noreferrer"&gt;Credential Guard overview&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows-server/security/osconfig/osconfig-how-to-configure-security-baselines#why-security-baselines-matter" target="_blank" rel="noopener noreferrer"&gt;OSConfig baseline credential protections&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows-server/get-started/removed-deprecated-features-windows-server#features-no-longer-in-development" target="_blank" rel="noopener noreferrer"&gt;Deprecated Windows Server features&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;Harden SMB and file-server access&lt;/H2&gt;
&lt;P&gt;SMB hardening protects Windows file sharing and related management traffic against protocol downgrade, relay attacks, on-path tampering, brute-force authentication, guest access, data disclosure, and exploitation of obsolete implementations such as SMBv1. Signing verifies message integrity, while encryption protects sensitive content in transit.&lt;/P&gt;
&lt;P&gt;Remove SMBv1, prevent insecure guest logons, retain and verify the Windows Server 2025 default requirement for inbound and outbound SMB signing, use SMB encryption for sensitive or untrusted network paths, and use SMB 3.x for modern file services. Treat any relaxation of signing for an incompatible third-party device as a documented, isolated, and time-bound exception. Windows Server 2025 also provides SMB authentication rate limiting and stronger signing and encryption capabilities that should be retained unless a documented compatibility requirement exists.&lt;/P&gt;
&lt;P&gt;Restrict TCP port 445 to approved clients and servers, apply share and NTFS permissions according to least privilege, enable access-based enumeration where appropriate, and audit access to sensitive shares. Do not publish traditional SMB directly to the internet; use a supported secure access design such as SMB over QUIC when its requirements and threat model fit.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Inventory SMB clients, servers, protocol versions, shares, and access requirements, then remove SMBv1 and insecure guest access. Verify that inbound and outbound signing remain required, document and isolate any temporary third-party compatibility exception, configure encryption, authentication rate limiting, and firewall scope according to the workload, review share and NTFS permissions, pilot changes with older clients and high-throughput workloads, and monitor SMB security, authentication, and performance events after enforcement.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; Mandatory SMB signing and encryption consume processor resources and can reduce throughput or increase latency on demanding file workloads. Older storage appliances, scanners, applications, or clients might not support modern SMB requirements, and overly restrictive port or permission changes can disrupt file access, administration, Group Policy, or backup operations.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/windows-server/storage/file-server/smb-security-hardening" target="_blank" rel="noopener noreferrer"&gt;SMB security hardening&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows-server/storage/file-server/smb-secure-traffic" target="_blank" rel="noopener noreferrer"&gt;Secure SMB traffic in Windows Server&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;Require modern TLS and manage certificates securely&lt;/H2&gt;
&lt;P&gt;Modern TLS protects application and management traffic by authenticating endpoints and encrypting data in transit. Requiring current protocol versions, strong cipher suites, and trusted certificates mitigates eavesdropping, man-in-the-middle attacks, protocol downgrade, weak-cryptography attacks, and impersonation using invalid or compromised certificates.&lt;/P&gt;
&lt;P&gt;Require TLS 1.2 or later and prefer TLS 1.3 where the application stack supports it. Windows Server 2025 disables TLS 1.0 and TLS 1.1 by default; verify that these protocols and obsolete SSL versions remain disabled and prevent unauthorized re-enablement. Disable weak cipher suites and obsolete hashes through a tested baseline rather than ad hoc registry changes. Inventory old agents, middleware, and network appliances first so incompatible dependencies can be upgraded instead of becoming permanent exceptions.&lt;/P&gt;
&lt;P&gt;Use certificates from a trusted public or enterprise certification authority, protect private keys with restrictive access control, select appropriate key sizes and algorithms, monitor expiration, and automate renewal. After dependency review, remove expired, untrusted, orphaned, or unnecessary certificates from server stores.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Inventory listening services, clients, protocol versions, cipher dependencies, and installed certificates, then replace weak or expiring certificates and confirm application support for modern TLS. Verify that TLS 1.0 and TLS 1.1 remain disabled, apply tested Schannel or OSConfig settings in stages, disable other legacy protocols and weak ciphers, validate every client and integration, rescan the endpoints, remove only certificates confirmed to be unnecessary, and implement automated certificate enrollment, renewal, expiration alerting, and private-key access review.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; Disabling old protocols and ciphers can break legacy clients, middleware, monitoring agents, or network devices with no modern TLS support. Certificate issuance, private-key protection, renewal automation, and revocation checking add operational complexity, and an expired or incorrectly deployed certificate can cause a complete service outage.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/windows-server/security/tls/tls-ssl-schannel-ssp-overview" target="_blank" rel="noopener noreferrer"&gt;TLS/SSL and Schannel overview&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows-server/security/osconfig/osconfig-how-to-configure-security-baselines#why-security-baselines-matter" target="_blank" rel="noopener noreferrer"&gt;OSConfig baseline protocol protections&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows-server/get-started/removed-deprecated-features-windows-server#features-no-longer-in-development" target="_blank" rel="noopener noreferrer"&gt;Deprecated Windows Server features&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;Encrypt operating-system and data volumes with BitLocker&lt;/H2&gt;
&lt;P&gt;BitLocker encrypts data at rest so possession of a disk or offline copy does not provide immediate access to its contents. It mitigates data theft from lost or stolen servers, removed drives, improperly decommissioned hardware, offline password-reset attacks, and attempts to read files by booting an alternate operating system.&lt;/P&gt;
&lt;P&gt;Enable BitLocker on operating-system and data volumes, using TPM-backed protectors and additional startup authentication where the physical threat model and availability requirements justify it. Use virtual TPMs and supported host or cloud protections for virtual machines. Encryption protects data on removed drives, decommissioned hardware, stolen systems, and offline copies.&lt;/P&gt;
&lt;P&gt;Escrow recovery information in a protected, recoverable directory or management service before enforcement. Limit access to recovery keys, audit retrieval, include key recovery in incident procedures, and test recovery on representative systems so encryption does not become an availability risk.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Inventory operating-system and data volumes, confirm TPM or virtual TPM readiness, select protectors that meet the physical and availability threat model, and configure a protected recovery-key escrow location. Enable BitLocker in controlled stages, verify encryption and key backup, test normal reboot and recovery scenarios, document break-glass procedures, and continuously monitor encryption and protector compliance. Suspend protection for firmware or boot-chain maintenance only through an approved procedure, then verify that BitLocker protection resumes afterward.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; Lost recovery material can make encrypted data permanently inaccessible, while firmware, TPM, boot, or hardware changes can unexpectedly trigger recovery. Encryption can add some performance and operational overhead, and startup PINs can conflict with unattended reboot requirements, so protector selection, key escrow, and recovery testing must reflect the server's availability needs.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/windows/security/operating-system-security/data-protection/bitlocker/planning-guide" target="_blank" rel="noopener noreferrer"&gt;BitLocker planning guide&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows/security/operating-system-security/data-protection/bitlocker/operations-guide" target="_blank" rel="noopener noreferrer"&gt;BitLocker operations guide&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows/security/operating-system-security/data-protection/bitlocker/recovery-overview" target="_blank" rel="noopener noreferrer"&gt;BitLocker recovery overview&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;Configure detailed auditing and protect local logs&lt;/H2&gt;
&lt;P&gt;Detailed auditing records security-relevant activity so suspicious behavior can be detected, investigated, and attributed. Authentication, privilege, process, PowerShell, policy, and firewall logs help expose brute-force attempts, credential misuse, privilege escalation, persistence, defense evasion, and attacker efforts to alter system configuration.&lt;/P&gt;
&lt;P&gt;Enable advanced audit policy for successful and failed logons, credential validation, account and group changes, sensitive privilege use, process creation with command-line capture, policy changes, removable storage, file shares, firewall activity, and other events relevant to the server role. The OSConfig baseline enables a broad audit configuration and increases important log sizes to improve forensic coverage.&lt;/P&gt;
&lt;P&gt;Enable PowerShell module and script block logging, and use protected event logging where appropriate because command content can contain sensitive data. Increase log capacity and retention for the expected event volume, restrict permissions to clear or modify logs, monitor audit-policy changes, and synchronize time with trusted sources.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Define the activities and events required for detection, investigation, and compliance, then apply advanced audit policy through OSConfig or Group Policy. Enable process command-line and PowerShell logging, measure event volume during a representative pilot, size and protect each log and forwarding path from the observed rates, configure trusted time synchronization, generate representative test events to confirm collection, and review event volume, retention, and policy health regularly.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; Detailed auditing can generate large volumes of events, consume storage and processing resources, and overwhelm analysts with noise if collection isn't tuned. Command-line and PowerShell logs can contain credentials or other sensitive data, while undersized logs may overwrite useful evidence, so access, retention, filtering, and capacity require deliberate design.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/windows-server/security/osconfig/osconfig-how-to-configure-security-baselines#why-security-baselines-matter" target="_blank" rel="noopener noreferrer"&gt;OSConfig baseline auditing and visibility&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/windows/security/operating-system-security/device-management/use-windows-event-forwarding-to-assist-in-intrusion-detection#appendix-a---minimum-recommended-minimum-audit-policy" target="_blank" rel="noopener noreferrer"&gt;Recommended audit policy for Windows Event Forwarding&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/powershell/module/microsoft.powershell.core/about/about_logging_windows?view=powershell-7.6" target="_blank" rel="noopener noreferrer"&gt;PowerShell logging on Windows&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;Centralize security telemetry and alert on suspicious activity&lt;/H2&gt;
&lt;P&gt;Centralized telemetry moves evidence away from the system that generated it and correlates activity across servers, identities, and networks. This improves detection of distributed attacks, limits an intruder's ability to erase local evidence, and shortens response time for credential attacks, lateral movement, persistence, defense evasion, and destructive actions.&lt;/P&gt;
&lt;P&gt;Forward security-relevant logs away from each server using Windows Event Forwarding, Azure Monitor, Microsoft Defender, a SIEM such as Microsoft Sentinel, or another protected collection platform. Include Security, System, Windows Defender, PowerShell, Code Integrity, Windows Firewall, Windows LAPS, and role-specific operational logs.&lt;/P&gt;
&lt;P&gt;Create actionable alerts for repeated authentication failures, new or changed administrators, unexpected service or scheduled-task creation, security-control changes, Defender detections, App Control blocks, log clearing, unusual remote administration, and backup deletion. Restrict access to collectors and retention systems so an attacker who compromises a server cannot erase the centralized evidence.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Select Windows Event Forwarding, Azure Monitor, Microsoft Defender, a SIEM, or a combination; define prioritized detection use cases and role-specific retention before selecting log channels and verbosity; design resilient collectors, access control, and capacity; and deploy the required agents or subscriptions. Onboard the prioritized channels, verify end-to-end ingestion and timestamps, create and test high-value detections and notifications, restrict access to the monitoring platform, and continuously monitor collection health and tune noisy rules.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; Central collection introduces bandwidth, storage, ingestion, licensing, retention, and analyst costs and can expose sensitive operational data if the monitoring platform is poorly secured. Collector failures create visibility gaps, while poorly tuned rules produce false positives and alert fatigue, so the design needs resilience, health monitoring, access controls, and continuous tuning.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/windows/security/operating-system-security/device-management/use-windows-event-forwarding-to-assist-in-intrusion-detection" target="_blank" rel="noopener noreferrer"&gt;Use Windows Event Forwarding for intrusion detection&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/defender-endpoint/microsoft-defender-endpoint-windows#security-capabilities-for-windows-environments" target="_blank" rel="noopener noreferrer"&gt;Microsoft Defender for Endpoint security capabilities&lt;/A&gt;&lt;/P&gt;
&lt;H2&gt;Maintain ransomware-resilient backups and test recovery&lt;/H2&gt;
&lt;P&gt;Ransomware-resilient backups preserve a trustworthy recovery path when production data, operating systems, or identity services are encrypted, deleted, or corrupted. Isolated and immutable copies mitigate ransomware, destructive administrators, compromised backup credentials, accidental deletion, hardware failure, and attacks intended to eliminate both systems and their recovery data.&lt;/P&gt;
&lt;P&gt;Keep multiple protected backup copies, including a copy that is offline, immutable, or otherwise isolated from normal server and domain administrator credentials. Use separate backup administration identities, multifactor authorization for destructive operations, encryption, soft delete or immutability controls, and alerts for policy changes or mass deletion. Hypervisor snapshots alone are not an adequate backup strategy.&lt;/P&gt;
&lt;P&gt;Back up application data and configuration as well as system state and bare-metal recovery data where required by the server role. Define recovery-point and recovery-time objectives, test file, application, system-state, and full-server restoration regularly, and record the evidence. Domain controllers, certificate authorities, and other identity infrastructure require workload-aware recovery procedures.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Implementation steps:&lt;/STRONG&gt; Classify workloads and define recovery-point and recovery-time objectives, then select local, offsite, offline, and immutable backup targets appropriate to the risk. Use separate backup identities and MFA, schedule application data, configuration, system-state, and bare-metal backups as required, enable encryption and deletion protections, and monitor every job and policy change. Perform regular isolated restore tests that verify application consistency and role-specific recovery semantics for identity systems, not only successful restoration of files or virtual disks, and maintain documented recovery runbooks.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Possible drawbacks:&lt;/STRONG&gt; Multiple isolated copies, immutable storage, long retention, and regular restore exercises increase storage, network, licensing, staffing, and operational costs. Backups can create false confidence when they are incomplete, stale, infected, or untested, and strong credential separation can slow routine administration, so restore validation and lifecycle management are as important as backup creation.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Documentation:&lt;/STRONG&gt; &lt;A href="https://learn.microsoft.com/azure/architecture/security/ransomware-resilient-backup-architecture/" target="_blank" rel="noopener noreferrer"&gt;Design a ransomware-resilient backup architecture&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/azure/backup/azure-backup-data-protection-best-practices" target="_blank" rel="noopener noreferrer"&gt;Azure Backup security best practices&lt;/A&gt; | &lt;A href="https://learn.microsoft.com/azure/backup/backup-azure-system-state" target="_blank" rel="noopener noreferrer"&gt;Back up Windows Server system state&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;---&lt;/P&gt;
&lt;P&gt;This isn't everything you can do, but it's a start. What other techniques do you use to harden your Windows Server deployments?&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 22 Jul 2026 20:30:09 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/some-tools-and-techniques-for-hardening-windows-server/ba-p/4539840</guid>
      <dc:creator>OrinThomas</dc:creator>
      <dc:date>2026-07-22T20:30:09Z</dc:date>
    </item>
    <item>
      <title>Cut Your Azure Blob Storage Bill in Half: A Practical Walkthrough of Object Storage TCO</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/cut-your-azure-blob-storage-bill-in-half-a-practical-walkthrough/ba-p/4534575</link>
      <description>&lt;P&gt;Hello Folks!&lt;/P&gt;
&lt;P&gt;If you have ever opened your monthly Azure invoice, stared at the object storage line, and quietly wondered how it grew so much, this one is for you. At the Microsoft Azure Infra Summit 2026, Benedict Berger and George Trossell from the Azure Storage Engineering team walked through a real customer scenario and showed how to bring that bill down without touching a single application.&lt;/P&gt;
&lt;P&gt;📺 &lt;STRONG&gt;Watch the session:&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;DIV class="lia-embeded-content" contenteditable="false"&gt;&lt;IFRAME src="https://www.youtube.com/embed/a_YZLhnJrcg?si=hEweMhuIQAnRkRPQ" width="100%" title="YouTube video player" allowfullscreen="allowfullscreen" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" frameborder="0" style="aspect-ratio: 16/9; height: auto;" sandbox="allow-scripts allow-same-origin allow-forms"&gt;
&lt;/IFRAME&gt;&lt;/DIV&gt;
&lt;H2&gt;Why IT Pros Should Care&lt;/H2&gt;
&lt;P&gt;Storage is one of those services we configure once at account creation, then never revisit. Redundancy, default tier, lifecycle rules. All decided on day one, then forgotten. Meanwhile, applications get built on top, dashboards get wired up, and the bill keeps climbing in a department nobody really audits.&lt;/P&gt;
&lt;P&gt;Here is what you get when you make storage TCO a first-class part of your operating model:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;A defensible understanding of capacity, transactions, and data retrieval charges (the three real cost drivers).&lt;/LI&gt;
&lt;LI&gt;Fewer surprise spikes when a cool tier read pattern runs hotter than expected.&lt;/LI&gt;
&lt;LI&gt;Cost optimization that runs on its own, instead of a quarterly cleanup project nobody volunteers for.&lt;/LI&gt;
&lt;LI&gt;Storage standards baked into your Infrastructure as Code, so cost-efficient defaults travel with every new account.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;In short, this is one of the highest-leverage cost levers you have in Azure. And unlike compute right-sizing, you can act on most of it from the portal in an afternoon.&lt;/P&gt;
&lt;H2&gt;What Storage TCO Actually Means on Object Storage, a Technical Overview&lt;/H2&gt;
&lt;P&gt;When Benedict and George talk about Total Cost of Ownership on Azure Blob Storage, they mean four moving parts:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;Capacity&lt;/STRONG&gt;. The per-gigabyte cost of the data you store, which varies by access tier and by redundancy.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Transactions&lt;/STRONG&gt;. Every read, write, list, and metadata call against the storage account. Priced in packages of 10,000 operations.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Data retrieval&lt;/STRONG&gt;. A per-gigabyte fee that applies when you read from cool or cold tiers. It is free on hot.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Network egress&lt;/STRONG&gt;. The charge for moving data out of an Azure region.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;The trap most teams fall into is looking only at the per-gigabyte capacity column and picking the cheapest tier they see. That ignores the fact that as data gets cooler, transaction and retrieval costs climb sharply, and cold has a 90-day early deletion penalty that can erase your savings outright. Microsoft Learn documents this trade-off clearly in the &lt;A href="https://learn.microsoft.com/en-us/azure/storage/blobs/access-tiers-overview" target="_blank"&gt;access tiers overview&lt;/A&gt;, where you can see the minimum retention windows and the relationship between storage cost and access cost across hot, cool, cold, and archive.&lt;/P&gt;
&lt;P&gt;Redundancy is the other dial. LRS keeps three copies in a single zone. ZRS spreads three copies across three zones in the region. GRS adds an asynchronous secondary in a paired region. The honest tradeoff George highlighted: redundancy protects your data, not your application. If your app is not zone-aware, ZRS alone will not keep you running through a zone outage. And GRS failover is a manual operation in most cases, with the secondary in read-only mode until you stand up new accounts to write into.&lt;/P&gt;
&lt;H2&gt;How It Works, Under the Hood&lt;/H2&gt;
&lt;P&gt;The session walked through a worked transaction example that finally made the math click for me. Picture a Spark job uploading 1,000 parquet files of 5 GB each into the hot tier, using an 8 MB block size.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Each 5 GB file is roughly 5,120 MB, divided by 8 MB blocks, which gives 640 put block operations.&lt;/LI&gt;
&lt;LI&gt;One additional put block list call commits the upload, so each object costs 641 write operations.&lt;/LI&gt;
&lt;LI&gt;Times 1,000 files, that is 641,000 operations, which works out to about 3.52 US dollars in that hour just for writes.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Now flip it. Read those same 1,000 files from the cool tier. The transaction count is similar, but you also pay a data retrieval fee on every gigabyte you pull back. That retrieval fee is where most teams get blindsided, because it does not show up on the hot tier at all.&lt;/P&gt;
&lt;P&gt;Block size matters too. Larger blocks mean fewer transactions per upload. And for small objects (under 128 KB), there is a new wrinkle to plan for: starting July 2026 for existing accounts and already in effect for new accounts created from July 2025, cooler tiers bill a 128 KB minimum object size. That means a 4 KB log file moved to cool gets charged as if it were 128 KB. The fix is either to leave small objects in hot, or bin-pack them into larger objects (a TAR or ZIP, for example) before tiering them down. The Microsoft Learn page on &lt;A href="https://learn.microsoft.com/en-us/azure/storage/blobs/access-tiers-best-practices" target="_blank"&gt;access tier best practices&lt;/A&gt; covers packing strategies in detail.&lt;/P&gt;
&lt;H2&gt;Real-World Value, Use Cases, and ROI&lt;/H2&gt;
&lt;P&gt;The customer in the session went from roughly 65,000 US dollars a month to around 25,000. That is not a marketing number, it is what happens when you apply the levers in order:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Right-size redundancy. Move non-production and easily reproducible data off LRS in production. Reserve GRS for the workloads where a compliance regulation actually requires a second region.&lt;/LI&gt;
&lt;LI&gt;Match tiers to access patterns. Use premium for bursty, latency-sensitive workloads. Hot for active reads and writes. Cool and cold only when you genuinely access the data infrequently and have budgeted for the retrieval fees.&lt;/LI&gt;
&lt;LI&gt;Buy reserved capacity for the steady-state portion of your footprint. A one or three year commitment unlocks a discount on block blob capacity. See &lt;A href="https://learn.microsoft.com/en-us/azure/storage/blobs/storage-blob-reserved-capacity" target="_blank"&gt;reserved capacity for Blob storage&lt;/A&gt; for terms and tier coverage.&lt;/LI&gt;
&lt;LI&gt;Kill wasteful transactions. Replace polling-for-changes with change feed. Replace recurring list-blob loops with a daily or weekly blob inventory report. Use conditional request headers (If-Modified-Since and friends) so reads skip unchanged objects.&lt;/LI&gt;
&lt;LI&gt;Pack small objects, or leave them in hot. Either is fine; tiering them down without packing is not.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;In short, the same scenario, with the same applications, runs at less than half the cost once you actually look at it.&lt;/P&gt;
&lt;H2&gt;Getting Started&lt;/H2&gt;
&lt;P&gt;Here is the order of operations I would follow tomorrow morning:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Pull a &lt;A href="https://learn.microsoft.com/en-us/azure/storage/blobs/blob-inventory" target="_blank"&gt;blob inventory report&lt;/A&gt; on your largest storage accounts to see what is actually there: tier mix, object sizes, last modified dates, snapshots, versions.&lt;/LI&gt;
&lt;LI&gt;Open the &lt;A href="https://azure.microsoft.com/en-us/pricing/calculator/" target="_blank"&gt;Azure pricing calculator&lt;/A&gt; and model your scenario with realistic transaction counts and retrieval volumes. Do not just compare per-GB prices.&lt;/LI&gt;
&lt;LI&gt;Audit your redundancy choices against the workload. If an account is LRS in production with no easy way to rebuild the data, change it.&lt;/LI&gt;
&lt;LI&gt;Enable Smart Tier on your zone-redundant accounts. New objects start in hot, get demoted to cool after 30 days of inactivity, and to cold after 90, with no charges for tier transitions, early deletions, or data retrieval. Anything accessed gets instantly promoted back to hot.&lt;/LI&gt;
&lt;LI&gt;For accounts that cannot use Smart Tier, write a &lt;A href="https://learn.microsoft.com/en-us/azure/storage/blobs/lifecycle-management-overview" target="_blank"&gt;lifecycle management policy&lt;/A&gt;. Keep the rules simple at first: tier down after 30 days, archive after 180, expire snapshots and versions on a schedule.&lt;/LI&gt;
&lt;LI&gt;Convert one of your existing lifecycle policies to ARM or Bicep, then commit it to source control. Add an Azure Policy that flags any new storage account that does not match your standard.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;That last step is the one that sticks. As Benedict put it, cost optimization must become part of your system, not an afterthought.&lt;/P&gt;
&lt;H2&gt;Resources&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/storage/blobs/access-tiers-overview" target="_blank"&gt;Access tiers for blob data, Microsoft Learn&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/storage/blobs/access-tiers-best-practices" target="_blank"&gt;Best practices for using blob access tiers, Microsoft Learn&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/storage/blobs/lifecycle-management-overview" target="_blank"&gt;Azure Blob Storage lifecycle management overview, Microsoft Learn&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/storage/blobs/storage-blob-reserved-capacity" target="_blank"&gt;Optimize costs for Blob storage with reserved capacity, Microsoft Learn&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/storage/blobs/blob-inventory" target="_blank"&gt;Enable Azure Storage blob inventory reports, Microsoft Learn&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://azure.microsoft.com/en-us/pricing/calculator/" target="_blank"&gt;Azure pricing calculator&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Keep Learning at the Summit&lt;/H2&gt;
&lt;P&gt;Catch the full &lt;A class="lia-external-url" href="https://www.youtube.com/playlist?list=PLjt5SKzX1iI8con7FJDB56G6hHqxGm7ki" target="_blank"&gt;Microsoft Azure Infra Summit 2026 session playlist here&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;Cheers!&lt;/P&gt;
&lt;P&gt;Pierre Roman&lt;/P&gt;</description>
      <pubDate>Wed, 22 Jul 2026 07:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/cut-your-azure-blob-storage-bill-in-half-a-practical-walkthrough/ba-p/4534575</guid>
      <dc:creator>Pierre_Roman</dc:creator>
      <dc:date>2026-07-22T07:00:00Z</dc:date>
    </item>
    <item>
      <title>Feeding the GPUs: File Storage for AI and Cloud-Native Workloads on Azure</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/feeding-the-gpus-file-storage-for-ai-and-cloud-native-workloads/ba-p/4534572</link>
      <description>&lt;P&gt;Hello Folks!&lt;/P&gt;
&lt;P&gt;If you are running AI workloads on Azure, you have probably learned the hard way that the wrong storage choice can leave a rack of very expensive GPUs sitting idle, waiting for data. In this session during the Microsoft Azure Infra Summit 2026, Wolfgang de Salvador and Reena Shah from the Azure Storage team walked through how Azure Managed Lustre and Azure Files map to the distinct stages of the AI pipeline, and why picking the right file system per stage is one of the highest leverage decisions you will make.&lt;/P&gt;
&lt;P&gt;📺 &lt;STRONG&gt;Watch the session:&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;DIV class="lia-embeded-content" contenteditable="false"&gt;&lt;IFRAME src="https://www.youtube.com/embed/RFqd9yIlTBM?si=N1_FwcubEHSgarbE" width="100%" title="YouTube video player" allowfullscreen="allowfullscreen" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" frameborder="0" style="aspect-ratio: 16/9; height: auto;" sandbox="allow-scripts allow-same-origin allow-forms"&gt;
&lt;/IFRAME&gt;&lt;/DIV&gt;
&lt;H2&gt;Why IT Pros Should Care&lt;/H2&gt;
&lt;P&gt;You probably did not get into IT to babysit checkpoint writes or debug Hugging Face egress bills at 2 a.m. But that is exactly the kind of work that lands on your plate when storage is not matched to the workload. Here is why MAIS28 matters for the IT pros, platform engineers, and Azure architects in the room:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;GPU time is the most expensive compute you will ever buy. Slow data loading and slow checkpoints turn that into burned cash.&lt;/LI&gt;
&lt;LI&gt;AI workloads are not one workload. Data prep, training, fine-tuning, and inferencing each have a different storage profile.&lt;/LI&gt;
&lt;LI&gt;Cloud-native AI on AKS and Azure Container Apps lives or dies on the ReadWriteMany experience. If model loading is slow or shared model caches do not exist, every cold start re-downloads hundreds of gigabytes.&lt;/LI&gt;
&lt;LI&gt;Storage choices ripple into security and compliance. Encryption in transit, redundancy, and snapshots are not optional in 2026.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;In short, this session is for anyone who has to answer the question, “What persistent volume should we use for this AI workload?” and wants a defensible answer.&lt;/P&gt;
&lt;H2&gt;What Azure Brings to the Table, Technical Overview&lt;/H2&gt;
&lt;P&gt;Wolfgang opened with the storage profile of every stage of an AI workflow. Data preparation needs hundreds of petabytes at the best TCO (think Azure Blob Storage as the durable core). Training and fine-tuning need extreme throughput so GPUs stay fed during data loading and so checkpoint writes complete fast. Inferencing needs fast model loads, low-latency KV cache, and grounded data for RAG. One filesystem does not fit all of those at once, and trying to make it fit is where teams overspend.&lt;/P&gt;
&lt;P&gt;Azure’s answer is a tiered, file-based portfolio that lines up with those stages:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure Managed Lustre (AMLFS)&lt;/STRONG&gt; is a fully managed, accelerator-tier filesystem. It scales to 25 PB of capacity and up to 512 GB/s of throughput, integrates with Azure Blob Storage as the durable core, and exposes a standard Lustre client plus a CSI driver for AKS.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure Files&lt;/STRONG&gt; is the natural ReadWriteMany choice for cloud-native AI on AKS and Azure Container Apps. It tops out at 256 TB of capacity and 10.4 GB/s of throughput, offers LRS and ZRS redundancy with snapshots and soft delete, and ships with a 99.99% SLA.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Azure Blob Storage&lt;/STRONG&gt; sits underneath both of these as the cheap, durable core for data prep and long-term retention.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The mental model the speakers used is “accelerator and core”. Blob is the core, durable and economical. AMLFS is the accelerator for training. Azure Files is the accelerator for inferencing and shared state. Pick the right pair for the stage you are running.&lt;/P&gt;
&lt;H2&gt;How It Works, Under the Hood&lt;/H2&gt;
&lt;P&gt;For training, Wolfgang showed the demo most folks came to see: a 32 x H100 ND H100 v5 AKS cluster deployed from the Azure AI Infrastructure repository, running a 30B-parameter GPT-3 training job backed by AMLFS. Two things matter here.&lt;/P&gt;
&lt;P&gt;First, AMLFS absorbs checkpoint write bursts. When 32 H100s flush state at the same time, you need a filesystem that can take the punch without stalling. AMLFS does, which keeps GPU utilization drops short and contained.&lt;/P&gt;
&lt;P&gt;Second, the AMLFS Lustre CSI driver for AKS supports both static and dynamic provisioning with availability-zone placement, and there are five SKU tiers from MLFS20 (cheapest by capacity) to MLFS500 (cheapest by bandwidth). That means you can pick a cost-performance point that matches your training budget instead of buying the top SKU and hoping for the best.&lt;/P&gt;
&lt;P&gt;For inferencing, Reena’s half of the session was just as practical. Five reasons Azure Files fits AKS ReadWriteMany workloads:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Standard Kubernetes RWM volume over NFS or SMB.&lt;/LI&gt;
&lt;LI&gt;256 TB capacity ceiling and up to 10.4 GB/s of throughput per share.&lt;/LI&gt;
&lt;LI&gt;LRS or ZRS redundancy with snapshots and soft delete for protection.&lt;/LI&gt;
&lt;LI&gt;99.99% SLA so it shows up in your availability math.&lt;/LI&gt;
&lt;LI&gt;Native support across AKS and Azure Container Apps, including serverless GPU.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The headline feature is &lt;STRONG&gt;Azure Files Provisioned v2&lt;/STRONG&gt;. In the old model, IOPS and throughput were a function of how much capacity you provisioned, which is wrong for AI shapes that need small capacity but very high IOPS and bandwidth. Provisioned v2 splits capacity, IOPS, and throughput into three independent knobs you can dial without downtime or remount. That alone changes the economics for a lot of inferencing patterns.&lt;/P&gt;
&lt;P&gt;The other big inferencing feature is &lt;STRONG&gt;NFS v4.1 encryption in transit&lt;/STRONG&gt;, delivered with the az-nfs utility and stunnel. You get AES-GCM TLS protection on the wire, with no Kerberos and no Active Directory needed, and the application has no idea it is happening. Reena’s live demo showed an AKS pod with an encrypted NFS mount, transparent to the workload.&lt;/P&gt;
&lt;P&gt;And then the pattern that ties it together: the &lt;STRONG&gt;shared model cache&lt;/STRONG&gt;. Download the model once into Azure Files, mount it across every replica via ReadWriteMany. No per-pod cold start, no re-download from Hugging Face, no egress bill. The demo used GPT-OSS 120B with VLLM on 32 x H100, and the pattern scales down to small fine-tuned models running on serverless GPU in Azure Container Apps.&lt;/P&gt;
&lt;H2&gt;Real-World Value&lt;/H2&gt;
&lt;P&gt;The session closed with the Viton case study. Viton is a Paris-based fashion AI startup. Their image-generation platform runs on Azure Container Apps serverless GPU, with Azure Service Bus for job routing and Azure Files NFS as the shared model store. Workers pull jobs, mount the shared model cache, generate the image, and scale to zero when the queue drains. The economics only work because they are not paying to re-download the model on every cold start, and because they only pay for GPU when there is work to do.&lt;/P&gt;
&lt;P&gt;The same pattern shows up across customer scenarios:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Training a foundation model on AKS&lt;/STRONG&gt; with AMLFS as the scratch tier and Blob as the durable archive.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Fine-tuning&lt;/STRONG&gt; smaller models where AMLFS checkpoints absorb the write bursts and the final artifact lands back in Blob.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Inferencing with VLLM&lt;/STRONG&gt; on AKS where Azure Files holds the model weights once and every replica reads from the same RWM mount.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Serverless inferencing&lt;/STRONG&gt; on Azure Container Apps with the same shared model cache pattern, but with scale-to-zero economics.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Honest tradeoff: Lustre is not the right filesystem for a 10-pod web app, and Azure Files is not the right filesystem for a 32-GPU training run. The whole point of the tiering is that you pick the right one per stage. Do not try to make one filesystem do all four jobs.&lt;/P&gt;
&lt;H2&gt;Getting Started&lt;/H2&gt;
&lt;P&gt;If you want to put this into practice this week:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Go to the &lt;A href="https://github.com/Azure/azurehpc-ai" target="_blank"&gt;Azure AI Infrastructure repository&lt;/A&gt; on GitHub. Wolfgang’s demo cluster came straight out of it, and you can spin up an AI-ready AKS cluster with GPU and InfiniBand operators in your own dev/test subscription.&lt;/LI&gt;
&lt;LI&gt;Install the &lt;A href="https://learn.microsoft.com/en-us/azure/azure-managed-lustre/use-csi-driver-kubernetes" target="_blank"&gt;Azure Managed Lustre CSI driver&lt;/A&gt; on AKS if you are running training or fine-tuning. Start with a smaller MLFS SKU and size up.&lt;/LI&gt;
&lt;LI&gt;Turn on &lt;A href="https://learn.microsoft.com/en-us/azure/storage/files/understanding-billing" target="_blank"&gt;Azure Files Provisioned v2&lt;/A&gt; on a new share, then dial capacity, IOPS, and throughput independently to match your inferencing shape.&lt;/LI&gt;
&lt;LI&gt;Enable &lt;A href="https://learn.microsoft.com/en-us/azure/storage/files/encryption-in-transit-for-nfs-shares" target="_blank"&gt;NFS v4.1 encryption in transit&lt;/A&gt; with az-nfs and stunnel before you put any sensitive workload on the wire.&lt;/LI&gt;
&lt;LI&gt;Try the shared model cache pattern. Pick one VLLM deployment, point it at an Azure Files RWM mount, and measure cold start time before and after.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2&gt;Resources&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/azure-managed-lustre/" target="_blank"&gt;Azure Managed Lustre documentation&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/azure-managed-lustre/use-csi-driver-kubernetes" target="_blank"&gt;Use the Azure Managed Lustre CSI driver with Azure Kubernetes Service&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/storage/files/" target="_blank"&gt;Azure Files documentation&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/storage/files/understanding-billing" target="_blank"&gt;Understand Azure Files billing (Provisioned v2)&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/storage/files/encryption-in-transit-for-nfs-shares" target="_blank"&gt;Encryption in transit for NFS Azure file shares&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/container-apps/gpu-serverless-overview" target="_blank"&gt;Azure Container Apps serverless GPUs&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://github.com/Azure/azurehpc-ai" target="_blank"&gt;Azure AI Infrastructure repository on GitHub&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Keep Learning at the Summit&lt;/H2&gt;
&lt;P&gt;Catch the full &lt;A class="lia-external-url" href="https://www.youtube.com/playlist?list=PLjt5SKzX1iI8con7FJDB56G6hHqxGm7ki" target="_blank"&gt;Microsoft Azure Infra Summit 2026 session playlist here&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;Cheers!&lt;/P&gt;
&lt;P&gt;Pierre Roman&lt;/P&gt;</description>
      <pubDate>Tue, 21 Jul 2026 07:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/feeding-the-gpus-file-storage-for-ai-and-cloud-native-workloads/ba-p/4534572</guid>
      <dc:creator>Pierre_Roman</dc:creator>
      <dc:date>2026-07-21T07:00:00Z</dc:date>
    </item>
    <item>
      <title>Premium SSD v2 and Instant Access Snapshots: A Better, Faster, Cheaper Disk for Your Azure VMs</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/premium-ssd-v2-and-instant-access-snapshots-a-better-faster/ba-p/4534571</link>
      <description>&lt;P&gt;Hello Folks!&lt;/P&gt;
&lt;P&gt;If you have been running Premium SSD v1 because that is just what you have always done, this session from the Microsoft Azure Infra Summit 2026 is going to be a wake up call. Raymond Lui and Adam Li from the Azure Disk Storage team walked us through Premium SSD v2 (PV2 for short) and the new Instant Access Snapshots, and the punchline is simple. PV2 is faster, it is cheaper, and the operational story around it just keeps getting better.&lt;/P&gt;
&lt;P&gt;📺 &lt;STRONG&gt;Watch the session:&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;DIV class="lia-embeded-content" contenteditable="false"&gt;&lt;IFRAME src="https://www.youtube.com/embed/-7HMosnCR2o?si=xtWYFesuJIdFkPxQ" width="100%" title="YouTube video player" allowfullscreen="allowfullscreen" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" frameborder="0" style="aspect-ratio: 16/9; height: auto;" sandbox="allow-scripts allow-same-origin allow-forms"&gt;
&lt;/IFRAME&gt;&lt;/DIV&gt;
&lt;H2&gt;Why IT Pros Should Care&lt;/H2&gt;
&lt;P&gt;If you are an infrastructure person, a SQL DBA, an SAP Basis admin, or anyone who has ever had to right-size a VM around its storage tier, this matters to you. In short, Premium SSD v2 changes the rules around how you provision block storage in Azure. Here is what stood out from the session:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;4x more IOPS and 2x more throughput compared to Premium SSD v1, on a matched configuration that costs 42% less.&lt;/LI&gt;
&lt;LI&gt;Sub-millisecond average latency, with a top configuration of 800,000 IOPS and 20 GB/s of throughput on a single VM.&lt;/LI&gt;
&lt;LI&gt;Capacity, IOPS, and throughput are decoupled. You dial each one independently, in 1 GB increments, instead of buying a tiered SKU.&lt;/LI&gt;
&lt;LI&gt;3,000 baseline IOPS and 125 MB/s throughput included on every disk, with no extra cost.&lt;/LI&gt;
&lt;LI&gt;Live Resize. You can grow disk size, IOPS, or throughput on a running VM with no restart required.&lt;/LI&gt;
&lt;LI&gt;Instant Access Snapshots make restores feel actually instant, with up to 10x faster hydration and 90% lower read latency during hydration.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;That is a lot of wins on one slide. Let’s break it down.&lt;/P&gt;
&lt;H2&gt;What Premium SSD v2 Is, Technical Overview&lt;/H2&gt;
&lt;P&gt;Premium SSD v2 is Azure’s purpose-built block storage for I/O-intensive enterprise workloads. Microsoft Learn describes it as designed for workloads that need sub-millisecond disk latency, high IOPS, and high throughput at a low cost. The target list is broad: SQL Server, Oracle, MariaDB, SAP, Cassandra, MongoDB, big data and analytics, gaming, and stateful containers running on AKS.&lt;/P&gt;
&lt;P&gt;The architectural shift that Raymond highlighted is independent scaling. With Premium SSD v1, you bought a fixed SKU. If you wanted more IOPS, you had to buy more capacity, even if you did not need it. With PV2, capacity, IOPS, and throughput are three separate dials. You provision capacity in 1 GB increments, then you set IOPS and throughput to match what your workload actually needs. If you over-provisioned, you tune it down. If you under-provisioned, you tune it up, and the VM keeps running.&lt;/P&gt;
&lt;P&gt;Raymond highlighted three primary use cases in the session:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;SAP workloads&lt;/STRONG&gt;, including SAP application VMs, SAP HANA databases, and non-HANA databases like Oracle, DB2, and SQL Server in SAP environments.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;SQL Server&lt;/STRONG&gt;. According to a GigaOM benchmark cited in the session, SQL Server on PV2 delivered 51% more transactions per second and 39% lower cost per transaction compared to AWS EC2, with a 9% lower 3-year TCO.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Big data and analytics replacing local SSD&lt;/STRONG&gt;. This one is a bit of a surprise. On D-series VMs, PV2 delivered over 1,400 MB/s of throughput compared to 720 MB/s from local SSD. That means you can run Spark or Databricks workloads on cheaper VM SKUs (without local storage) and still get more performance than you had before.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Premium SSD v2 supports a 4k physical sector size by default, with 512E available for legacy applications. There are a few honest tradeoffs to know about. PV2 disks cannot be used as an OS disk, and they cannot be used with Azure Compute Gallery. PV2 also does not support host caching. For regions with availability zones, PV2 disks can only be attached to zonal VMs, so plan your VM placement accordingly.&lt;/P&gt;
&lt;H2&gt;How It Works, Under the Hood&lt;/H2&gt;
&lt;P&gt;Raymond covered the architecture briefly, and it is worth understanding. PV2 uses direct VM-to-storage-node communication, with 3-replica durability behind the scenes. That direct path is part of how it gets sub-millisecond latency consistently.&lt;/P&gt;
&lt;P&gt;For Instant Access Snapshots, Adam walked through the architectural difference between the classic incremental snapshot path and the new Instant Access path. With classic incremental snapshots for PV2 and Ultra Disk, the snapshot is created, then the data has to copy in the background to Standard HDD before the snapshot is usable for restore. That copy could take a while on a large disk, and restored disks would then hydrate slowly, which dragged down read latency until hydration finished.&lt;/P&gt;
&lt;P&gt;With Instant Access, the snapshot is usable the moment it exists. The data stays in the same high-performance storage as the source disk for a configurable duration (60 to 300 minutes, controlled by the InstantAccessDurationMins parameter). At the same time, Azure copies the snapshot data to Standard ZRS in the background for long-term retention. When the Instant Access window expires, the snapshot transitions to a regular incremental snapshot, sitting on cheap durable storage. You get the speed and the long-term durability without running two separate workflows.&lt;/P&gt;
&lt;P&gt;In short, your VM can boot and run at near-full performance while the data hydrates in the background. There are some limits to keep in mind. Instant Access counts toward the existing limit of three in-progress snapshots per disk, and you can create up to 15 disks concurrently from all instant access snapshots of a single disk.&lt;/P&gt;
&lt;H2&gt;Real-World Value (Use Cases, ROI, Scenarios)&lt;/H2&gt;
&lt;P&gt;Adam closed his portion of the session with a BCDR demo. He cloned 12 disks of an M-series production database into a recovery VM, attached them, and was immediately running roughly 500,000 IOPS at single-digit-millisecond latency. No waiting for hydration. No degraded performance window. That is a meaningful improvement to your Recovery Time Objective (RTO).&lt;/P&gt;
&lt;P&gt;A few scenarios where this combination really pays off:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Pre-deployment safety nets.&lt;/STRONG&gt; Take an instant access snapshot before a big upgrade. If something goes sideways, roll back in seconds instead of hours.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Rapid scale-out for stateful apps.&lt;/STRONG&gt; Spin up multiple disk copies of a primary instance in seconds. You can even place them across availability zones in the same region.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Dev/test environment refresh.&lt;/STRONG&gt; Clone production into dev or test on demand, with full performance from the first I/O. No more “we’ll refresh dev next quarter” because the restore takes too long.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;SAP HANA always-on operations.&lt;/STRONG&gt; Live Resize means you can scale IOPS or throughput up on a running database during a load spike, without a maintenance window.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Right-sizing to cut spend.&lt;/STRONG&gt; If you have been paying for VM SKUs purely to get local SSD throughput, PV2 may let you drop to a smaller, cheaper VM and still hit higher numbers.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;One nuance came up in the live Q&amp;amp;A. Jens asked a great question about profiling: how do you know when PV2 is the right choice versus Standard SSD? Raymond’s guidance was direct. If the workload needs high IOPS or high throughput, PV2 is generally the right call. The VM SKU also needs to support “Premium Disk” capability for PV2 to attach, so check that compatibility first.&lt;/P&gt;
&lt;H2&gt;Getting Started&lt;/H2&gt;
&lt;P&gt;Concrete first steps so you can start kicking the tires:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;Confirm region and zone support.&lt;/STRONG&gt; Use az vm list-skus --resource-type disks --query "[?name=='PremiumV2_LRS']" to see which regions and availability zones are supported in your subscription.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Pick a Premium-capable VM in a supported zone.&lt;/STRONG&gt; Remember, PV2 is zonal in AZ regions. Decide on the zone before you create the VM.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Provision a disk.&lt;/STRONG&gt; Start with default performance (3,000 IOPS, 125 MB/s) and a small capacity. You are paying for the dials you turn up; defaults are reasonable for most starting points.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Plan your v1 to v2 migration.&lt;/STRONG&gt; Raymond demoed two paths. Option A: detach the disk from a running VM and convert it (the VM keeps running on its other disks). Option B: stop and deallocate the VM, then convert in place. Both preserve data, and you can raise IOPS and throughput as part of the conversion.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Try Instant Access Snapshots.&lt;/STRONG&gt; Add --instant-access-duration-in-minutes (or the equivalent ARM/PowerShell parameter) to your existing snapshot command. That is all the change you need to enable it.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;For AKS users&lt;/STRONG&gt;, define a storage class with skuName: PremiumV2_LRS and let dynamic provisioning take it from there.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2&gt;Resources&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/virtual-machines/disks-types" target="_blank"&gt;Select a disk type for Azure IaaS VMs (managed disks)&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/virtual-machines/disks-deploy-premium-v2" target="_blank"&gt;Deploy a Premium SSD v2 managed disk&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/virtual-machines/disks-convert-types" target="_blank"&gt;Convert managed disks storage between different disk types&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/virtual-machines/disks-instant-access-snapshots" target="_blank"&gt;Instant access snapshots for Azure managed disks&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/virtual-machines/use-premium-ssd-v2-with-availability-set" target="_blank"&gt;Use Premium SSD v2 with VMs in an availability set&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/aks/use-premium-v2-disks" target="_blank"&gt;Use Azure Premium SSD v2 disks on Azure Kubernetes Service&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/virtual-machines/managed-disks-overview" target="_blank"&gt;Azure managed disks overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/virtual-machines/workloads/sap/hana-vm-premium-ssd-v2" target="_blank"&gt;SAP HANA Azure virtual machine Premium SSD v2 storage configurations&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Keep Learning at the Summit&lt;/H2&gt;
&lt;P&gt;Catch the full &lt;A class="lia-external-url" href="https://www.youtube.com/playlist?list=PLjt5SKzX1iI8con7FJDB56G6hHqxGm7ki" target="_blank"&gt;Microsoft Azure Infra Summit 2026 session playlist here&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;Cheers!&lt;/P&gt;
&lt;P&gt;Pierre Roman&lt;/P&gt;</description>
      <pubDate>Mon, 20 Jul 2026 07:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/premium-ssd-v2-and-instant-access-snapshots-a-better-faster/ba-p/4534571</guid>
      <dc:creator>Pierre_Roman</dc:creator>
      <dc:date>2026-07-20T07:00:00Z</dc:date>
    </item>
    <item>
      <title>Network Security Perimeter for Azure Event Hubs: Hardening Your Data Streams</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/network-security-perimeter-for-azure-event-hubs-hardening-your/ba-p/4538352</link>
      <description>&lt;H3&gt;What is Network Security Perimeter for Azure Event Hubs?&lt;/H3&gt;
&lt;P&gt;Azure Event Hubs now supports Network Security Perimeter (NSP), a logical network isolation boundary that lets you define a security perimeter around your PaaS resources and control public network access through perimeter-based access rules.&lt;/P&gt;
&lt;P&gt;In practical terms, this means you can now group Event Hubs resources within a perimeter, apply consistent network access policies across them, and prevent unauthorized inbound traffic at the PaaS boundary level. It's not a firewall replacement, it's a compliance and segmentation tool that works alongside your existing NSGs and private endpoints.&lt;/P&gt;
&lt;P&gt;Before NSP, managing network access to Event Hubs involved:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Private endpoints (which route traffic over private networks)&lt;/LI&gt;
&lt;LI&gt;IP firewall rules (which block public access from specific CIDR blocks)&lt;/LI&gt;
&lt;LI&gt;Virtual Network Service Endpoints (which restrict traffic to VNets)&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Network Security Perimeter adds a declarative, organization-wide layer: you define which resources belong inside the perimeter, and then manage access rules once, and those policies apply consistently across all perimeter members. Changes to the perimeter automatically cascade to all enrolled resources.&lt;/P&gt;
&lt;H3&gt;Why ITPros Should Care&lt;/H3&gt;
&lt;P&gt;If you're managing Event Hubs in a regulated industry like healthcare, finance, or government, you know the pressure. Compliance auditors want proof that data pipelines are segmented, isolated, and protected from lateral movement. Network Security Perimeter directly addresses that.&lt;/P&gt;
&lt;H3&gt;Operational Value&lt;/H3&gt;
&lt;P&gt;Network Security Perimeter delivers three immediate operational wins:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Single Source of Truth for Access Rules. Instead of managing firewall rules on each Event Hubs namespace independently, you manage rules once at the perimeter level. Reduce configuration drift, reduce the attack surface, reduce human error.&lt;/LI&gt;
&lt;LI&gt;Compliance and Audit Readiness. Demonstrate network isolation to auditors with a clear diagram: "All Event Hubs in the perimeter are protected by these rules." That narrative matters for SOC 2, FedRAMP, HIPAA, and PCI-DSS compliance. You can export perimeter configurations and attach them to compliance documentation.&lt;/LI&gt;
&lt;LI&gt;Simplified Onboarding. When a new Event Hubs namespace joins the organization, add it to the perimeter and it inherits all access rules automatically. No manual rule-by-rule configuration. No weeks of back-and-forth with security teams.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Secondary benefits include:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Reduced blast radius during incidents, if an application is compromised, perimeter rules limit what it can access.&lt;/LI&gt;
&lt;LI&gt;Simplified network topology diagrams for architecture reviews.&lt;/LI&gt;
&lt;LI&gt;Faster mean time to remediation (MTTR) when security issues arise.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;Real-World Example: Securing a Multi-Tenant Event Hub Deployment&lt;/H3&gt;
&lt;P&gt;Let's walk through a practical scenario. You're an ITPro at a financial services firm. You have three Event Hubs namespaces:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;hubs-prod-transactions (production trading data)&lt;/LI&gt;
&lt;LI&gt;hubs-prod-compliance (regulatory event streams)&lt;/LI&gt;
&lt;LI&gt;hubs-staging-dev (development and testing)&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Your security policy mandates:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Production namespaces should only accept traffic from specific applications (IP-restricted).&lt;/LI&gt;
&lt;LI&gt;Staging can accept traffic from developer VNets but not from the internet.&lt;/LI&gt;
&lt;LI&gt;All outbound access to external services must be logged and monitored.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;Step 1: Define Your Perimeter&lt;/H3&gt;
&lt;P&gt;First, create a Network Security Perimeter in the Azure Portal or via Azure CLI:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;az network perimeter create --resource-group rg-security --name nsp-financialservices --location eastus&lt;/LI-CODE&gt;
&lt;P&gt;This creates the perimeter container. Think of it as a logical security zone.&lt;/P&gt;
&lt;H3&gt;Step 2: Enroll Event Hubs Resources&lt;/H3&gt;
&lt;P&gt;Add your Event Hubs namespaces to the perimeter:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;az network perimeter access-rule create --resource-group rg-security --perimeter-name nsp-financialservices --name allow-prod-apps --direction Inbound --access Allow --protocols Tcp --source-address-prefix 10.0.0.0/8 --destination-port-range 5671-5672
&lt;/LI-CODE&gt;
&lt;P&gt;Enroll the Event Hubs namespace:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;az network perimeter resource create --resource-group rg-security --perimeter-name nsp-financialservices --resource-name hubs-prod-transactions --resource-type "Microsoft.EventHub/namespaces"&lt;/LI-CODE&gt;
&lt;P&gt;You've now enrolled your production Event Hubs namespace. It inherits the "allow-prod-apps" rule, only traffic from your internal VNET (10.0.0.0/8) is permitted.&lt;/P&gt;
&lt;H3&gt;Step 3: Define Access Rules&lt;/H3&gt;
&lt;LI-CODE lang=""&gt;$ns = "hubs-prod-transactions" $hub = "transactions-hub" $key = (az eventhubs namespace authorization-rule keys list --resource-group rg-prod --namespace-name $ns --name RootManageSharedAccessKey --query primaryConnectionString --output tsv)&lt;/LI-CODE&gt;
&lt;P&gt;Create rules that reflect your security policy. Allow internal compliance applications:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;az network perimeter access-rule create --resource-group rg-security --perimeter-name nsp-financialservices --name allow-compliance-writers --direction Inbound --access Allow --protocols Tcp --source-address-prefix 10.50.0.0/16 --destination-port-range 5671-5672&lt;/LI-CODE&gt;
&lt;P&gt;Deny all other public traffic:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;az network perimeter access-rule create --resource-group rg-security --perimeter-name nsp-financialservices --name deny-internet --direction Inbound --access Deny --protocols "*" --source-address-prefix "*" --destination-port-range "*"&lt;/LI-CODE&gt;
&lt;P&gt;Now your Event Hubs accept traffic only from specific internal subnets. Everything else is rejected at the PaaS boundary.&lt;/P&gt;
&lt;H3&gt;Step 4: Validate Connectivity&lt;/H3&gt;
&lt;P&gt;Test that legitimate applications can still reach Event Hubs:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;$ns = "hubs-prod-transactions" $hub = "transactions-hub" $key = (az eventhubs namespace authorization-rule keys list --resource-group rg-prod --namespace-name $ns --name RootManageSharedAccessKey --query primaryConnectionString --output tsv)&lt;/LI-CODE&gt;
&lt;P&gt;Check logs in Azure Monitor:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;az monitor log-analytics query --workspace $(az monitor log-analytics workspace list --query "[0].id" -o tsv) --analytics-query "AzureDiagnostics | where ResourceProvider=='MICROSOFT.EVENTHUB' | summarize by NetworkSecurityPerimeter_s"&lt;/LI-CODE&gt;
&lt;P&gt;If you see accepted connections logged with your perimeter name, you're good. If you see denied connections from unexpected IPs, you've caught a security issue before it impacts production.&lt;/P&gt;
&lt;H3&gt;Step 5: Monitor and Alert&lt;/H3&gt;
&lt;P&gt;Set up alerts for denied traffic:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;az monitor metrics alert create --name "NSP-Denied-Connections" --resource-group rg-security --scopes /subscriptions/{subId}/resourceGroups/rg-security/providers/Microsoft.Network/networkSecurityPerimeters/nsp-financialservices --condition "avg ConnectionRejectedCount &amp;gt; 5" --window-size 5m --evaluation-frequency 1m --action email-admin@company.com&lt;/LI-CODE&gt;
&lt;P&gt;Now you'll be notified if someone attempts to access Event Hubs from an unauthorized source. Your security posture just went from reactive to proactive.&lt;/P&gt;
&lt;H3&gt;Technical Details: How NSP Works Under the Hood&lt;/H3&gt;
&lt;H3&gt;Perimeter Architecture&lt;/H3&gt;
&lt;P&gt;Network Security Perimeter operates at the Azure platform level, not in your VNets. Here's the flow:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Connection arrives at Event Hubs public IP.&lt;/LI&gt;
&lt;LI&gt;Azure evaluates the source IP/protocol against NSP rules.&lt;/LI&gt;
&lt;LI&gt;If allowed, connection is routed to the namespace.&lt;/LI&gt;
&lt;LI&gt;If denied, connection is dropped and logged.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;This happens before TLS handshake, reducing CPU overhead and improving response times. Denied connections generate zero namespace load.&lt;/P&gt;
&lt;H3&gt;Rule Evaluation Order&lt;/H3&gt;
&lt;P&gt;NSP rules are evaluated in this order:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Explicit Allow rules (matched first wins)&lt;/LI&gt;
&lt;LI&gt;Explicit Deny rules&lt;/LI&gt;
&lt;LI&gt;Implicit Deny (default action)&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Best practice: Create your Allow rules first (be specific about what you permit), then add Deny rules for anything not explicitly allowed. This ensures you don't accidentally block legitimate traffic.&lt;/P&gt;
&lt;H3&gt;Integration with Existing Security Tools&lt;/H3&gt;
&lt;P&gt;NSP works alongside (not instead of):&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Private Endpoints: NSP adds a policy layer; private endpoints route traffic over Azure backbone. Use both.&lt;/LI&gt;
&lt;LI&gt;IP Firewall: NSP provides namespace-level access control; IP firewall is still available for per-namespace rules.&lt;/LI&gt;
&lt;LI&gt;VNet Service Endpoints: NSP complements VNet endpoints by adding perimeter-wide policies.&lt;/LI&gt;
&lt;LI&gt;Managed Identity + RBAC: NSP is transport-layer security; identity-based access control remains separate.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;Performance Considerations&lt;/H3&gt;
&lt;P&gt;NSP introduces minimal latency (&amp;lt;1ms typically). Azure evaluates rules in parallel and caches common decisions. For high-throughput Event Hubs:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Keep rules simple and specific (avoid wildcard ranges if possible).&lt;/LI&gt;
&lt;LI&gt;Use CIDR blocks instead of individual IPs where applicable.&lt;/LI&gt;
&lt;LI&gt;Monitor connection acceptance rates in Azure Monitor.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;Comprehensive Resources&lt;/H3&gt;
&lt;P&gt;Official Microsoft Documentation:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/virtual-network/network-security-perimeter/network-security-perimeter-overview" data-test-app-aware-link="" target="_blank"&gt;Network Security Perimeter Overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/event-hubs/network-security" data-test-app-aware-link="" target="_blank"&gt;Event Hubs Network Security&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/event-hubs/configure-network-security-perimeter" data-test-app-aware-link="" target="_blank"&gt;Configuring NSP for Event Hubs&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/cli/azure/network/perimeter" data-test-app-aware-link="" target="_blank"&gt;Azure CLI: az network perimeter&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/event-hubs/authorize-access-azure-active-directory" data-test-app-aware-link="" target="_blank"&gt;Azure RBAC for Event Hubs&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/event-hubs/event-hubs-protocol-guide" data-test-app-aware-link="" target="_blank"&gt;Azure Event Hubs Protocol Guide&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;Closing: Perimeter Security for Modern Data Streams&lt;/H3&gt;
&lt;P&gt;Network Security Perimeter for Event Hubs is a quiet but powerful addition to Azure's security toolkit. You get the ability to enforce organization-wide network policies without having to reconfigure every namespace individually. You can demonstrate perimeter-based isolation to auditors. You can catch lateral-movement attacks before they happen.&lt;/P&gt;
&lt;P&gt;For ITPros managing event-driven architectures, message processors, IoT data streams, financial transactions, this capability directly improves your security posture and reduces operational overhead.&lt;/P&gt;
&lt;P&gt;I encourage you to:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Audit your current Event Hubs deployments. How many namespaces? How many security policies are you managing today?&lt;/LI&gt;
&lt;LI&gt;Design your perimeter boundaries. Group namespaces by security zone (prod, staging, dev) or by business unit.&lt;/LI&gt;
&lt;LI&gt;Start with one perimeter in a dev environment. Define rules. Validate connectivity. Then expand to staging and production.&lt;/LI&gt;
&lt;LI&gt;Document your perimeter architecture and rules. Include it in your security runbook and architecture reviews.&lt;/LI&gt;
&lt;LI&gt;Set up monitoring and alerting. Denied connections are a leading indicator of either misconfiguration or attack attempts.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;The networking challenges in cloud are complex. Network Security Perimeter gives you a declarative, policy-driven way to solve them at scale. Take advantage of it, and let me know how it changes your security workflows.&lt;/P&gt;
&lt;P&gt;Keep your networks hardened, and your data flowing safe.&lt;/P&gt;
&lt;P&gt;Cheers!&lt;/P&gt;
&lt;P&gt;Pierre Roman&lt;/P&gt;</description>
      <pubDate>Fri, 17 Jul 2026 20:26:45 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/network-security-perimeter-for-azure-event-hubs-hardening-your/ba-p/4538352</guid>
      <dc:creator>Pierre_Roman</dc:creator>
      <dc:date>2026-07-17T20:26:45Z</dc:date>
    </item>
    <item>
      <title>Az Update - Week 2 of the return editions</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/az-update-week-2-of-the-return-editions/ba-p/4538337</link>
      <description>&lt;P&gt;Hello Folks!&lt;/P&gt;
&lt;P&gt;This week's updates all focus on something we hear from IT pros and platform engineers all the time:&lt;/P&gt;
&lt;P&gt;How do we make our environments more secure, more manageable, and easier to modernize without adding more complexity?&lt;/P&gt;
&lt;P&gt;Whether you're running PostgreSQL workloads in Azure, securing Kubernetes storage, or planning your next wave of SQL Server migrations, this week's announcements bring practical improvements that can help reduce operational overhead while strengthening your overall platform strategy.&lt;/P&gt;
&lt;P&gt;We'll look at three newly available capabilities:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Update #1 - Generally Available: Microsoft Defender security assessments for Azure Database for PostgreSQL Flexible Server&lt;/LI&gt;
&lt;LI&gt;Update #2 - Generally Available: Encryption in Transit for Azure Files NFS Shares in Azure Kubernetes Service (AKS)&lt;/LI&gt;
&lt;LI&gt;Update #3 - Generally Available: Expanding Azure Arc SQL Migration with SQL Server on Azure Virtual Machines&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;As always, I'm approaching these updates from an infrastructure and operations perspective. I'll cover why each capability matters, what to watch out for before production deployment, and some practical steps you can take to start evaluating them in your own environment.&lt;/P&gt;
&lt;P&gt;Let's dig in.&lt;/P&gt;
&lt;H2&gt;Update #1 - Generally Available: Microsoft Defender security assessments for Azure Database for PostgreSQL Flexible Server&lt;/H2&gt;
&lt;H3&gt;Why ITPros should care&lt;/H3&gt;
&lt;P&gt;This release brings automated security posture assessment directly into managed PostgreSQL environments. For ITPros, this matters because database security is often treated separately from infrastructure security tooling, creating blind spots and silos.&lt;/P&gt;
&lt;P&gt;What changed is that Defender now runs native vulnerability scanning and compliance checks against PostgreSQL configurations, patches, and the ways a database could be exposed to security risks or attack opportunities. Instead of relying on external scanners or manual audits, you get platform-native assessments integrated with your existing Defender workflows.&lt;/P&gt;
&lt;P&gt;The operational impact is significant: you can now enforce security baselines at the database layer with the same consistency you apply to VMs and network resources, reducing the gap between infrastructure and data security accountability.&lt;/P&gt;
&lt;H3&gt;Operational value&lt;/H3&gt;
&lt;P&gt;Operationally, this improves your security baseline enforcement and reduces the need for separate database security assessment tools. It also strengthens how well you can demonstrate and prove that security controls are in place and working for compliance reviews where regulators expect consistent, documented security controls.&lt;/P&gt;
&lt;P&gt;Before production rollout, validate that Defender cost models fit your budget, that assessment frequency aligns with your change windows, and that remediation guidance maps to your patch and maintenance processes.&lt;/P&gt;
&lt;P&gt;Prerequisites include enabling Microsoft Defender for Cloud, registering the PostgreSQL Flexible Server provider, and ensuring network connectivity so assessments can reach the database endpoint.&lt;/P&gt;
&lt;H3&gt;Real-world example with step-by-step guidance&lt;/H3&gt;
&lt;OL&gt;
&lt;LI&gt;Enable Microsoft Defender for Cloud if not already active, and ensure PostgreSQL Flexible Server subscription coverage.&lt;/LI&gt;
&lt;LI&gt;Register the target PostgreSQL Flexible Server instances and confirm Defender has network visibility to the database endpoints.&lt;/LI&gt;
&lt;LI&gt;Run a baseline assessment and review initial findings to understand current security posture and common remediation patterns.&lt;/LI&gt;
&lt;LI&gt;Prioritise findings by severity and business impact, then schedule patches and configuration changes in maintenance windows.&lt;/LI&gt;
&lt;LI&gt;Monitor ongoing assessments and track remediation progress through Defender dashboards, validating that fixes reduce exposure scores.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H3&gt;Technical details including code examples&lt;/H3&gt;
&lt;P&gt;This example validates that Defender is actively assessing your PostgreSQL estate. The sequence checks Defender status, confirms PostgreSQL registration, and retrieves current assessment scores.&lt;/P&gt;
&lt;P&gt;Run these queries in a pilot subscription first to understand data structure and expected output before scaling to production databases.&lt;/P&gt;
&lt;LI-CODE lang=""&gt;az account set --subscription &amp;lt;subscriptionId&amp;gt; az security sql-vulnerability-assessment baseline show --resource-group &amp;lt;rg&amp;gt; --server-name &amp;lt;postgresServer&amp;gt; --database-name &amp;lt;databaseName&amp;gt;

az security pricing show --subscription &amp;lt;subscriptionId&amp;gt; --query "[?name=='VirtualMachines' || name=='SqlServers' || name=='StorageAccounts'].[name,pricingTier]" -o table

az provider show --namespace Microsoft.DBforPostgreSQL --query "registrationState" -o tsv&lt;/LI-CODE&gt;
&lt;P&gt;Expected behaviour: Defender status shows active, PostgreSQL instances are registered with the provider, and pricing tier reflects your coverage level. If assessments do not run, check network rules, managed identity permissions, and Defender plan activation. If baseline data is missing, trigger a manual scan and wait for completion.&lt;/P&gt;
&lt;H3&gt;Comprehensive Resources&lt;/H3&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;A href="https://azure.microsoft.com/updates?id=567527" data-test-app-aware-link="" target="_blank"&gt;Azure update: Microsoft Defender security assessments for Azure Database for PostgreSQL Flexible Server&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/defender-for-cloud/defender-for-cloud-introduction" data-test-app-aware-link="" target="_blank"&gt;Microsoft Defender for Cloud overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/postgresql/flexible-server/concepts-security" data-test-app-aware-link="" target="_blank"&gt;Azure Database for PostgreSQL security&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/defender-for-cloud/defender-for-sql-on-machines" data-test-app-aware-link="" target="_blank"&gt;SQL vulnerability assessments in Defender for Cloud&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/defender-for-cloud/enable-defender-for-cloud" data-test-app-aware-link="" target="_blank"&gt;Enable Defender for Cloud&lt;/A&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2&gt;Update #2 - Generally Available: Encryption in Transit for Azure Files NFS Shares in Azure Kubernetes Service (AKS)&lt;/H2&gt;
&lt;H3&gt;Why ITPros should care&lt;/H3&gt;
&lt;P&gt;This release closes a significant gap in data protection for Kubernetes workloads consuming NFS shares from Azure Files. Previously, NFS traffic between AKS nodes and Azure Files was unencrypted, creating compliance and security risks for sensitive workloads.&lt;/P&gt;
&lt;P&gt;What changed is that you can now enforce encryption for NFS communication at the Azure Files layer, not just at the application layer. This is important because traditional NFS lacks built-in encryption, and relying on network isolation alone is increasingly insufficient.&lt;/P&gt;
&lt;P&gt;For ITPros managing regulated workloads (healthcare, finance, PII-sensitive data), this removes a control gap. Encryption in transit now becomes a platform-native feature instead of a workaround, reducing architecture complexity and improving auditability.&lt;/P&gt;
&lt;H3&gt;Operational value&lt;/H3&gt;
&lt;P&gt;The operational value is stronger compliance posture and reduced attack surface for data in motion between containers and storage. It also simplifies the security story when auditors ask about data protection controls.&lt;/P&gt;
&lt;P&gt;Before enabling in production, validate that NFS-over-TLS introduces acceptable latency overhead for your workload patterns, test failover and reconnection behaviour under encryption, and confirm that monitoring and logging still work correctly.&lt;/P&gt;
&lt;P&gt;Prerequisites include running AKS with Azure CNI or Kubenet networking, having Azure Files with NFS 4.1 enabled, and ensuring the NFS client libraries on container images support TLS.&lt;/P&gt;
&lt;H3&gt;Real-world example with step-by-step guidance&lt;/H3&gt;
&lt;OL&gt;
&lt;LI&gt;Create an Azure Files NFS share with encryption in transit enabled and confirm TLS version alignment with your security standards.&lt;/LI&gt;
&lt;LI&gt;Deploy a test AKS workload that mounts the NFS share and validate that pods mount successfully with encrypted traffic.&lt;/LI&gt;
&lt;LI&gt;Run performance baselines (throughput, latency, CPU overhead) before and after enabling encryption to document operational expectations.&lt;/LI&gt;
&lt;LI&gt;Monitor pod logs and Azure Files metrics during the test to confirm no silent failures or unexpected throttling occurs.&lt;/LI&gt;
&lt;LI&gt;Roll out to production workloads in stages, with clear rollback criteria tied to application latency and error rates.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H3&gt;Technical details including code examples&lt;/H3&gt;
&lt;P&gt;This example validates that your AKS cluster can successfully mount NFS shares with encryption enabled. The sequence checks cluster networking, confirms NFS connectivity, and tests mount success.&lt;/P&gt;
&lt;P&gt;Run these commands in a non-production cluster first to validate environment readiness before touching production storage.&lt;/P&gt;
&lt;LI-CODE lang=""&gt;az aks show --resource-group &amp;lt;rg&amp;gt; --name &amp;lt;clusterName&amp;gt; --query "networkProfile.{networkPlugin:networkPlugin,networkPolicy:networkPolicy,podCidr:podCidr}" -o jsonc

az storage account show --resource-group &amp;lt;rg&amp;gt; --name &amp;lt;storageAccount&amp;gt; --query "{name:name,kind:kind,accessTier:accessTier}" -o jsonc

kubectl get pvc -A --all-namespaces -o wide kubectl describe pv &amp;lt;pvName&amp;gt; | grep -i nfs&lt;/LI-CODE&gt;
&lt;P&gt;Expected behaviour: cluster networking is properly configured, storage account kind supports NFS, and PVC/PV resources show NFS mount points. If mounts fail, check network security group rules, storage account firewall allowances, and subnet delegation. If latency increases, monitor resource utilisation and adjust workload placement if needed.&lt;/P&gt;
&lt;H3&gt;Comprehensive Resources&lt;/H3&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;A href="https://azure.microsoft.com/updates?id=567787" data-test-app-aware-link="" target="_blank"&gt;Azure update: Encryption in Transit for Azure Files NFS Shares in Azure Kubernetes Service (AKS)&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/storage/files/files-nfs-protocol" data-test-app-aware-link="" target="_blank"&gt;Azure Files NFS support&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/aks/azure-files-volume" data-test-app-aware-link="" target="_blank"&gt;Mount Azure Files with NFS in AKS&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/storage/common/storage-security-guide" data-test-app-aware-link="" target="_blank"&gt;Azure storage security&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/aks/concepts-network" data-test-app-aware-link="" target="_blank"&gt;AKS networking concepts&lt;/A&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2&gt;Update #3 - Generally Available: Expanding Azure Arc SQL Migration with SQL Server on Azure Virtual Machines&lt;/H2&gt;
&lt;H3&gt;Why ITPros should care&lt;/H3&gt;
&lt;P&gt;This capability brings SQL Server migration into the Azure Arc operational footprint, creating a unified migration and inventory experience. For ITPros, this matters because SQL Server modernisation is often fragmented across multiple tools and teams.&lt;/P&gt;
&lt;P&gt;What changed is that you can now discover, assess, and execute SQL migrations through Arc-native workflows, using the same permissions and governance model you already have for infrastructure and hybrid resources.&lt;/P&gt;
&lt;P&gt;The operational gain is consistency: discovery data feeds migration planning, assessments surface blockers early, and rollout can be controlled through the same change and approvals processes you use for other infrastructure migrations.&lt;/P&gt;
&lt;H3&gt;Operational value&lt;/H3&gt;
&lt;P&gt;Operationally, this reduces tooling sprawl and improves coordination between infrastructure and database teams. Arc becomes your single control plane for tracking migration progress, managing runbooks, and collecting audit evidence.&lt;/P&gt;
&lt;P&gt;Before production use, validate that your SQL Server inventory is complete, that migration blockers are understood and addressed, and that your maintenance windows can accommodate expected cutover timings.&lt;/P&gt;
&lt;P&gt;Prerequisites include Azure Arc agent deployment on source VMs, Azure Database Migration Service readiness, and network connectivity to target Azure SQL resources.&lt;/P&gt;
&lt;H3&gt;Real-world example with step-by-step guidance&lt;/H3&gt;
&lt;OL&gt;
&lt;LI&gt;Deploy Azure Arc agents to SQL Server VMs and confirm all instances report healthy status with complete inventory data.&lt;/LI&gt;
&lt;LI&gt;Run Arc-integrated SQL Server assessments to identify compatibility issues, dependencies, and recommended migration targets.&lt;/LI&gt;
&lt;LI&gt;Pilot migration for a non-critical workload to establish runbook patterns, measure cutover time, and validate post-migration validation procedures.&lt;/LI&gt;
&lt;LI&gt;Execute validation tests: connectivity, login success, database consistency checks, job execution, and application integration tests.&lt;/LI&gt;
&lt;LI&gt;Scale migration in waves using documented runbooks, with gates for monitoring data health and application performance after each cutover.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H3&gt;Technical details including code examples&lt;/H3&gt;
&lt;P&gt;This example validates Arc agent health and SQL Server discovery completeness. The sequence ensures your Arc infrastructure is ready for migration workflows.&lt;/P&gt;
&lt;P&gt;Run these commands as part of your pre-migration checklist to catch configuration gaps before committing to migration timelines.&lt;/P&gt;
&lt;LI-CODE lang=""&gt;az account show --output table az connectedmachine list --resource-group &amp;lt;rg&amp;gt; --query "[].{name:name,status:status,osName:osName}" -o table

az resource list --resource-type Microsoft.AzureArcData/sqlServerInstances --query "[].{name:name,resourceGroup:resourceGroup,location:location}" -o table

az connectedmachine machine extension list --resource-group &amp;lt;rg&amp;gt; --machine-name &amp;lt;vmName&amp;gt; --query "[].{name:name,provisioningState:provisioningState}" -o table&lt;/LI-CODE&gt;
&lt;P&gt;Expected behaviour: Arc agents report healthy status, SQL Server instances are fully discovered with accurate inventory, and required extensions are provisioned successfully. If discovery is incomplete, check Arc agent connectivity, extension deployment, and SQL service running status on source VMs. If migration pre-checks fail, verify SQL Server version compatibility and review Defender logs for blocking issues.&lt;/P&gt;
&lt;H3&gt;Comprehensive Resources&lt;/H3&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;A href="https://azure.microsoft.com/updates?id=567362" data-test-app-aware-link="" target="_blank"&gt;Azure update: Expanding Azure Arc SQL Migration with SQL Server on Azure Virtual Machines&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/sql/sql-server/azure-arc/overview" data-test-app-aware-link="" target="_blank"&gt;Azure Arc SQL Server Overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/azure-arc/servers/overview" data-test-app-aware-link="" target="_blank"&gt;Azure Arc-enabled servers&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/azure-sql/virtual-machines/windows/sql-server-on-azure-vm-iaas-what-is-overview" data-test-app-aware-link="" target="_blank"&gt;SQL Server on Azure Virtual Machines&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/dms/dms-overview" data-test-app-aware-link="" target="_blank"&gt;Azure Database Migration Service&lt;/A&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;For any new capability this week, if they map to your operational roadmap, run a controlled pilot, measure the impact, and then scale with confidence. That is how you move the needle on modernisation while managing risk.&lt;/P&gt;
&lt;P&gt;Cheers!&lt;/P&gt;
&lt;P&gt;Pierre Roman&lt;/P&gt;</description>
      <pubDate>Fri, 17 Jul 2026 18:37:42 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/az-update-week-2-of-the-return-editions/ba-p/4538337</guid>
      <dc:creator>Pierre_Roman</dc:creator>
      <dc:date>2026-07-17T18:37:42Z</dc:date>
    </item>
    <item>
      <title>Azure Elastic SAN: Pooled, Cloud-Native Block Storage That Actually Acts Like a SAN</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/azure-elastic-san-pooled-cloud-native-block-storage-that/ba-p/4534554</link>
      <description>&lt;P&gt;Hello Folks!&lt;/P&gt;
&lt;P&gt;If you have ever lived through a Friday night SAN expansion, racking new shelves and praying the zoning held together, the idea of getting that same shared block storage model in Azure (without owning a single fibre channel cable) sounds almost too good to be true. In his session at the Microsoft Azure Infra Summit 2026, Kiran Cherukuwada, Principal PM in Azure Storage, walked us through exactly how Azure Elastic SAN does that, and where it fits next to the other block storage options on Azure.&lt;/P&gt;
&lt;P&gt;📺 &lt;STRONG&gt;Watch the session:&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;DIV class="lia-embeded-content" contenteditable="false"&gt;&lt;IFRAME src="https://www.youtube.com/embed/SF313FsAwMU?si=qvnNKcMdWL7mIXJ2" width="100%" title="YouTube video player" allowfullscreen="allowfullscreen" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" frameborder="0" style="aspect-ratio: 16/9; height: auto;" sandbox="allow-scripts allow-same-origin allow-forms"&gt;
&lt;/IFRAME&gt;&lt;/DIV&gt;
&lt;H2&gt;Why IT Pros Should Care&lt;/H2&gt;
&lt;P&gt;Most of us were taught a simple rule. One workload, one disk, size it for peak, move on. That rule has been kind to managed disks, but it gets expensive fast when you have dozens or hundreds of workloads that all peak at different times. Elastic SAN flips the model. You provision a pool of capacity and performance once, then carve volumes out of it for many workloads.&lt;/P&gt;
&lt;P&gt;Here is why that matters for IT pros:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;You stop over-provisioning each workload to its own peak; the SAN absorbs the bursts.&lt;/LI&gt;
&lt;LI&gt;You get a SAN-style resource hierarchy (SAN, volume groups, volumes) that looks and behaves like the on-prem model you already know.&lt;/LI&gt;
&lt;LI&gt;iSCSI connectivity means a wide compute footprint, including Azure Virtual Machines, Azure Kubernetes Service, Azure Container Instances, Azure VMware Solution, and Nutanix Cloud Clusters.&lt;/LI&gt;
&lt;LI&gt;You can drive storage throughput over VM network bandwidth, which often lets you keep a smaller (and cheaper) VM SKU.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;In short, if you have many IO-intensive workloads sharing one region, Elastic SAN is the lever that turns “buy peak for every workload” into “buy combined peak for the group.”&lt;/P&gt;
&lt;H2&gt;What Azure Elastic SAN Is, Technical Overview&lt;/H2&gt;
&lt;P&gt;Azure Elastic SAN is the industry’s first fully managed SAN storage service in the cloud. It brings the on-prem SAN consumption model to Azure as a single managed pool of block storage, shared across many workloads, accessed over the industry-standard iSCSI protocol.&lt;/P&gt;
&lt;P&gt;Inside the service you get three resources, matching the on-prem mental model:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;The Elastic SAN itself.&lt;/STRONG&gt; Top-level resource. This is where you provision overall capacity and performance, and where billing happens.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Volume groups.&lt;/STRONG&gt; Where you set network rules (service or private endpoints) and security policies. Any policy you apply here is inherited by every volume in the group, so a volume group is effectively your workload boundary.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Volumes.&lt;/STRONG&gt; The LUNs that you mount on compute. They show up as raw block devices on a VM, as iSCSI targets to a Kubernetes node, or as VMware data stores on AVS.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;A single SAN can scale to a petabyte of capacity, 2 million IOPS, and 80 GB/s of throughput. It is locally redundant by default, with a zone-redundant option, and shared volume support is there for clustered solutions like SQL Server Failover Cluster Instances and Azure VMware Solution. Network isolation is delivered via service endpoints and private endpoints, and data is encrypted at rest. Incremental snapshots are supported for fast point-in-time restore, and snapshots can be exported to managed disk snapshots when you need a hardened copy for backup or DR purposes.&lt;/P&gt;
&lt;P&gt;Where does it land in the block storage portfolio? Kiran framed it simply. Premium SSD v2 is the best price/performance for &lt;STRONG&gt;dedicated&lt;/STRONG&gt; per-workload performance. Ultra Disk is for the mission-critical, every-microsecond-matters workloads. Elastic SAN is the best price/performance option &lt;STRONG&gt;at scale&lt;/STRONG&gt;, when you have many workloads that can share a storage pool.&lt;/P&gt;
&lt;H2&gt;How It Works, Under the Hood&lt;/H2&gt;
&lt;P&gt;The economics live in the provisioning model. You buy two types of units:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Base unit.&lt;/STRONG&gt; Each base unit gives you 1 TiB of capacity plus 5,000 IOPS and 200 MB/s. Roughly 8 cents per GiB per month in East US.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Capacity-only unit.&lt;/STRONG&gt; Each capacity-only unit gives you 1 TiB of capacity but no extra performance. About 25 percent cheaper, around 6 cents per GiB per month in East US.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The pattern Kiran showed is “size for performance first, then top up capacity.” A 250 TiB SAN delivering 1 million IOPS and 40 GB/s came out to roughly 200 base units plus 50 capacity-only units, landing around 20 grand per month for the whole pool.&lt;/P&gt;
&lt;P&gt;The magic ingredient is &lt;STRONG&gt;dynamic performance sharing.&lt;/STRONG&gt; With traditional disks you provision each workload to its own peak. With Elastic SAN, you provision the &lt;STRONG&gt;combined&lt;/STRONG&gt; peak. So a SQL Server needing 60,000 IOPS, an AVS cluster needing 40,000, and an Oracle workload needing 100,000 IOPS look like 200,000 IOPS of dedicated disk. But if they never peak simultaneously, you can land a 150,000 IOPS SAN and let each workload hit its peak on demand. That is real money back.&lt;/P&gt;
&lt;P&gt;The second lever is &lt;STRONG&gt;throughput over network bandwidth.&lt;/STRONG&gt; Because Elastic SAN connects over iSCSI, storage I/O flows through the VM’s network pipe, not the VM’s disk throughput cap. Most VMs have far more network bandwidth than disk bandwidth, so you can drive higher storage throughput from a smaller VM SKU. That smaller SKU is cheaper to run, and (this is the quiet win) it can also cut per-core database licensing costs. As one attendee asked in the live Q&amp;amp;A, “Why is it possible to go beyond the VM disk throughput limit with SAN?” The answer: iSCSI traffic uses VM network bandwidth like any other VM-to-VM traffic, so the disk throttle does not apply.&lt;/P&gt;
&lt;P&gt;One honest tradeoff: that same network bandwidth is also used by your app-tier-to-database traffic. So if you are planning to push storage hard, size the VM with both flows in mind.&lt;/P&gt;
&lt;H2&gt;Real-World Value&lt;/H2&gt;
&lt;P&gt;Where does this actually pay off?&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Mixed enterprise workloads on Azure VMs.&lt;/STRONG&gt; SQL Server, Oracle, custom OLTP, sharing one SAN. Kiran’s demo ran SQL TPCC, an AVS cluster benchmark, and an Oracle OLTP load simultaneously off a single 30-base-unit SAN, and the metrics blade showed exactly how each volume group consumed performance.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Extending Azure VMware Solution storage.&lt;/STRONG&gt; Instead of buying expensive vSAN nodes just to grow storage, you connect AVS to an Elastic SAN datastore. Gen2 AVS private clouds skip the ExpressRoute gateway requirement and let you use a single private endpoint on the volume group.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Container Storage.&lt;/STRONG&gt; Azure Container Storage v2 with Elastic SAN backing is generally available. The fast attach and detach behavior means that even if a node or cluster goes down, the data sits on the SAN and persists.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Lift and shift from on-prem SAN.&lt;/STRONG&gt; Kiran shared one migration example: a workload with 100-plus vCPUs running off a mid-tier all-flash SAN array landed on Elastic SAN with roughly 64 percent TCO savings and performance that exceeded the original array.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;In short, this is a “many workloads, one pool” story. If you have one heavy workload, premium SSD v2 may be a better fit.&lt;/P&gt;
&lt;H2&gt;Getting Started&lt;/H2&gt;
&lt;P&gt;Here is a practical order of operations:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;Size the SAN.&lt;/STRONG&gt; Add up the combined peak IOPS and throughput for the workloads you plan to consolidate, then pick base units to cover performance and capacity-only units to top up storage.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Lock down the network.&lt;/STRONG&gt; Access is closed by default. Choose service endpoints or private endpoints per volume group, and open them only to the right subnets.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Place compute in the same zone.&lt;/STRONG&gt; For best latency, deploy your VMs (or AVS cluster) in the same region and availability zone as the SAN.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Tune the client.&lt;/STRONG&gt; Use Gen 5 (D, E, or M series) VMs with Accelerated Networking on, configure the iSCSI initiator, set up native MPIO on Windows or Linux, and use the Connect scripts from the portal which default to 32 sessions per volume.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Watch the metrics.&lt;/STRONG&gt; The SAN’s Metrics tab shows transactions, ingress, and egress at the SAN, volume group, and individual volume level. Drop the granularity to one minute when you are troubleshooting.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Plan snapshots.&lt;/STRONG&gt; Use Elastic SAN volume snapshots for fast dev/test restores. Export to managed disk snapshots when you need hardened backup or cross-region DR.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;If you are coming from on-prem, the partnership with Cirrus Data (free in the Azure Marketplace) is the recommended path to migrate storage at the block level.&lt;/P&gt;
&lt;H2&gt;Resources&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/storage/elastic-san/" target="_blank"&gt;Azure Elastic SAN documentation hub&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/storage/elastic-san/elastic-san-introduction" target="_blank"&gt;What is Azure Elastic SAN (introduction)&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/storage/elastic-san/elastic-san-planning" target="_blank"&gt;Plan for an Azure Elastic SAN deployment&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/storage/elastic-san/elastic-san-best-practices" target="_blank"&gt;Azure Elastic SAN configuration best practices&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/azure/storage/elastic-san/elastic-san-snapshots" target="_blank"&gt;Snapshot Azure Elastic SAN volumes&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Keep Learning at the Summit&lt;/H2&gt;
&lt;P&gt;Catch the full &lt;A class="lia-external-url" href="https://www.youtube.com/playlist?list=PLjt5SKzX1iI8con7FJDB56G6hHqxGm7ki" target="_blank"&gt;Microsoft Azure Infra Summit 2026 session playlist here&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;Cheers!&lt;/P&gt;
&lt;P&gt;Pierre Roman&lt;/P&gt;</description>
      <pubDate>Fri, 17 Jul 2026 07:15:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/azure-elastic-san-pooled-cloud-native-block-storage-that/ba-p/4534554</guid>
      <dc:creator>Pierre_Roman</dc:creator>
      <dc:date>2026-07-17T07:15:00Z</dc:date>
    </item>
    <item>
      <title>When Compute Sits and Waits: Fixing the Hidden Storage Bottleneck in EDA and HPC with Azure NetApp Files</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/when-compute-sits-and-waits-fixing-the-hidden-storage-bottleneck/ba-p/4534550</link>
      <description>&lt;P&gt;Hello Folks!&lt;/P&gt;
&lt;P&gt;If you have ever stared at an HPC pipeline and wondered why the queue depth keeps climbing while every CPU graph looks lazy, this session from the Microsoft Azure Infrastructure Summit 2026 is going to feel very familiar. Ron Hogue from the Azure Storage team and Ranga Sankar from NetApp spend twenty-seven minutes diagnosing the problem that nobody wants to admit out loud, that the bottleneck is not compute, not the scheduler, and not licenses. It is storage. And then they walk through exactly how Azure NetApp Files (ANF) solves it.&lt;/P&gt;
&lt;P&gt;📺 &lt;STRONG&gt;Watch the session:&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;DIV class="lia-embeded-content" contenteditable="false"&gt;&lt;IFRAME src="https://www.youtube.com/embed/MfyRyJqrI8c?si=ZGRBScVHnXh1SLEs" width="100%" title="YouTube video player" allowfullscreen="allowfullscreen" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" frameborder="0" style="aspect-ratio: 16/9; height: auto;" sandbox="allow-scripts allow-same-origin allow-forms"&gt;
&lt;/IFRAME&gt;&lt;/DIV&gt;
&lt;H2&gt;Why IT Pros Should Care&lt;/H2&gt;
&lt;P&gt;If you support a semiconductor, simulation, rendering, or scientific computing team, you have lived this story. Engineers ask for more cores. You buy them. Wall-clock times barely move. Managers then ask for more tool licenses, because surely that must be the constraint. Spoiler, it usually is not.&lt;/P&gt;
&lt;P&gt;Here is why this matters to anyone running shared infrastructure on Azure:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;EDA and HPC pipelines hammer storage with millions of tiny reads and writes, metadata operations like file creation, rename, and unlink, all happening concurrently across thousands of processes.&lt;/LI&gt;
&lt;LI&gt;That pattern breaks generic cloud file storage that was tuned for large sequential I/O.&lt;/LI&gt;
&lt;LI&gt;ANF brings the same NetApp ONTAP data services that EDA teams have used on-premises for decades, delivered as a native Azure service.&lt;/LI&gt;
&lt;LI&gt;You can move design environments to Azure without re-architecting tools, schedulers, or scratch path conventions.&lt;/LI&gt;
&lt;LI&gt;The economics changed in the last twelve months with the Flexible service level and Cool Access, so the old “ANF is too expensive for scratch” argument deserves a fresh look.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;In short, if your job involves keeping expensive engineers and expensive licenses busy, the storage layer deserves your attention.&lt;/P&gt;
&lt;H2&gt;What Azure NetApp Files Brings to EDA and HPC, the technical overview&lt;/H2&gt;
&lt;P&gt;ANF is a first-party Azure service running NetApp ONTAP on bare-metal infrastructure inside Azure datacenters. It speaks NFS v3, NFS v4.1, and SMB, with dual-protocol options, and it preserves the file semantics that EDA tools assume. Things like POSIX permissions, fast metadata operations, snapshots, and consistent low latency.&lt;/P&gt;
&lt;P&gt;Ron and Ranga framed the session in eight layers (Problem, Foundation, Accelerator, Scale, AI Ready, Guardrails, Optimize, Proof). Three pieces did most of the heavy lifting.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Foundation, Migration Assistant with SnapMirror.&lt;/STRONG&gt; This is how you get the data into Azure without rewriting your inventory. SnapMirror replicates from on-premises ONTAP or Cloud Volumes ONTAP into ANF while preserving metadata, permissions, snapshots, and directory structure, with continuous synchronization and minimal downtime. For EDA flows where a missing ACL can invalidate an entire run, that fidelity is not optional.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Accelerator, Cache Volumes.&lt;/STRONG&gt; Built on NetApp FlexCache, these volumes front your authoritative dataset (whether it lives in on-prem ONTAP or in Cloud Volumes ONTAP) and pull hot reads close to your Azure compute. Tool libraries, PDKs, shared reference data, all served at sub-millisecond latency without copying petabytes around. There is still one source of truth, which keeps your data governance story clean.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Scale, Large Volumes with Breakthrough mode.&lt;/STRONG&gt; A single ANF large volume in Breakthrough mode scales to 2 PiB and delivers throughput in the tens of GiB per second by fanning I/O across six storage endpoints. That lets you collapse sharded namespaces (the classic /proj1, /proj2, /proj3 split that nobody loves) into fewer high-throughput volumes that behave predictably under contention.&lt;/P&gt;
&lt;H2&gt;How it works, under the hood&lt;/H2&gt;
&lt;P&gt;Cache Volumes are the FlexCache pattern most ONTAP customers already know. You peer a cluster, point a cache at an origin, and the cache propagates data on demand. Ron demoed this live: cluster peering with on-prem ONTAP, a cache volume created in ANF, and reads served from cache while the authoritative copy stayed home.&lt;/P&gt;
&lt;P&gt;Large Volumes in Breakthrough mode are where the architecture gets interesting. Instead of a single mount point pinned to a single storage endpoint, Breakthrough mode exposes six storage endpoints for one logical volume. Clients can mount, balance I/O, and aggregate throughput across all six. Microsoft published Linux scale-out benchmarks showing a single 50 TiB large volume in Breakthrough mode sustaining roughly 50,000 MiB/s of sequential reads and approaching 1.8 million 8 KiB random read IOPS using twelve VMs (see the Microsoft Learn link in Resources).&lt;/P&gt;
&lt;P&gt;For shared environments, ANF added user and group quotas with real-time consumption reporting and hard limits. Ranga demoed quota rules that stopped a runaway simulation generating millions of scratch files from starving the rest of the team. If you have ever had to send the “who filled the scratch volume” email at 2am, this feature alone might justify the trip.&lt;/P&gt;
&lt;P&gt;On the cost side, the Flexible service level decouples capacity from throughput. You buy a capacity pool, you pick throughput independently, with 128 MiB/s of baseline throughput included and a ceiling of up to 640 MiB/s per provisioned TiB (which is roughly five times the Ultra service level). Cool Access transparently tiers cold blocks to Azure Blob behind the same file mount point.&lt;/P&gt;
&lt;P&gt;In short, you stop paying for throughput you do not need on archival volumes, and you stop overprovisioning capacity to chase throughput on small scratch volumes.&lt;/P&gt;
&lt;H2&gt;Real-world value, use cases, ROI, scenarios&lt;/H2&gt;
&lt;P&gt;The Proof section closed with SPEC Storage 2020 EDA Blended results that are worth reading carefully:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;A single large volume in Breakthrough mode sustained 2,880 EDA jobsets at about 0.51 ms overall response time.&lt;/LI&gt;
&lt;LI&gt;Six volumes scaled linearly to 17,280 jobsets at about 0.60 ms.&lt;/LI&gt;
&lt;LI&gt;FIO measurements approached 2 million IOPS.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The honest tradeoff: those numbers come from a benchmark, not your environment, and the SPEC tables show latency climbing at the highest load points. So treat the result as evidence that the platform behaves predictably under EDA-style concurrency, not as a guarantee for every workflow. That predictability is the part that matters. EDA leads do not lose sleep over peak throughput, they lose sleep over latency that drifts when concurrency rises.&lt;/P&gt;
&lt;P&gt;Practical scenarios where this lands:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Burst regression and verification runs to Azure during tape-out crunches, with Cache Volumes keeping your on-prem golden tree authoritative.&lt;/LI&gt;
&lt;LI&gt;Full migration of EDA environments to Azure for teams whose datacenters are out of capacity or out of lease.&lt;/LI&gt;
&lt;LI&gt;HPC simulation workloads (CFD, weather, seismic, life sciences) that share the same metadata-heavy I/O profile.&lt;/LI&gt;
&lt;LI&gt;AI-adjacent pipelines that need to read design data with file semantics and surface the same bytes as objects to Fabric, OneLake, or Databricks via the dual file and object access pattern Ron mentioned in the AI Ready section.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Getting Started, concrete first steps&lt;/H2&gt;
&lt;P&gt;You do not need a multi-quarter project to get a useful pilot moving.&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Register the Azure NetApp Files resource provider in your target subscription and request quota in the regions you care about.&lt;/LI&gt;
&lt;LI&gt;Stand up a small capacity pool, start with the Flexible service level so you can dial throughput independently.&lt;/LI&gt;
&lt;LI&gt;Create a test volume, mount it from a representative VM (HBv4 or Ev5 family are good starting points for HPC and EDA respectively).&lt;/LI&gt;
&lt;LI&gt;Run your real workload, not just FIO. Use a representative regression batch or simulation job and watch the metadata patterns in the ANF metrics blade.&lt;/LI&gt;
&lt;LI&gt;If you have on-premises ONTAP, peer a Cache Volume against it to test FlexCache behavior with your actual datasets before committing to a full migration.&lt;/LI&gt;
&lt;LI&gt;Layer in user and group quotas before you open the volume to a wider team. Trust me on this one.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2&gt;Resources&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/azure-netapp-files/" target="_blank" rel="noopener"&gt;Azure NetApp Files documentation&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/azure-netapp-files/azure-netapp-files-service-levels" target="_blank" rel="noopener"&gt;Service levels for Azure NetApp Files, including the Flexible tier&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/azure-netapp-files/performance-large-volume-breakthrough-mode-linux" target="_blank" rel="noopener"&gt;Azure NetApp Files large volume breakthrough mode performance benchmarks for Linux&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/azure-netapp-files/solutions-benefits-azure-netapp-files-electronic-design-automation" target="_blank" rel="noopener"&gt;Benefits of using Azure NetApp Files for Electronic Design Automation (EDA)&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/azure-netapp-files/cache-volumes" target="_blank" rel="noopener"&gt;Azure NetApp Files cache volumes&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/azure-netapp-files/manage-quotas-user-group" target="_blank" rel="noopener"&gt;Manage user and group quotas on Azure NetApp Files volumes&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/azure-netapp-files/manage-cool-access" target="_blank" rel="noopener"&gt;Manage Azure NetApp Files cool access&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Keep Learning at the Summit&lt;/H2&gt;
&lt;P&gt;Catch the full &lt;A class="lia-external-url" href="https://www.youtube.com/playlist?list=PLjt5SKzX1iI8con7FJDB56G6hHqxGm7ki" target="_blank" rel="noopener"&gt;Microsoft Azure Infra Summit 2026 session playlist&lt;/A&gt; here.&lt;/P&gt;
&lt;P&gt;Cheers!&lt;/P&gt;
&lt;P&gt;Pierre Roman&lt;/P&gt;</description>
      <pubDate>Thu, 16 Jul 2026 07:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/when-compute-sits-and-waits-fixing-the-hidden-storage-bottleneck/ba-p/4534550</guid>
      <dc:creator>Pierre_Roman</dc:creator>
      <dc:date>2026-07-16T07:00:00Z</dc:date>
    </item>
    <item>
      <title>Modernize VDI with Azure Files and Entra Cloud-Native Identities</title>
      <link>https://techcommunity.microsoft.com/t5/itops-talk-blog/modernize-vdi-with-azure-files-and-entra-cloud-native-identities/ba-p/4534545</link>
      <description>&lt;P&gt;Hello Folks!&lt;/P&gt;
&lt;P&gt;If you have ever run a Virtual Desktop Infrastructure (VDI) estate, you know the recurring riddle. The session hosts are designed to be stateless and pooled, yet every user expects a persistent Outlook profile, their OneDrive cache, their pinned apps, and a sub-ten-second logon. In this session at the Microsoft Azure Infrastructure Summit 2026, Adam Groves and Priyanka Gangal from the Azure Files team showed how Azure Files plus Microsoft Entra ID finally let you deliver that experience without dragging domain controllers along for the ride.&lt;/P&gt;
&lt;P&gt;📺 &lt;STRONG&gt;Watch the session:&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;DIV class="lia-embeded-content" contenteditable="false"&gt;&lt;IFRAME src="https://www.youtube.com/embed/cTKkvfH0KZw?si=y6r6ROtU04TznVej" width="100%" title="YouTube video player" allowfullscreen="allowfullscreen" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" frameborder="0" style="aspect-ratio: 16/9; height: auto;" sandbox="allow-scripts allow-same-origin allow-forms"&gt;
&lt;/IFRAME&gt;&lt;/DIV&gt;
&lt;H2&gt;Why IT Pros Should Care&lt;/H2&gt;
&lt;P&gt;VDI has always been a balancing act between elasticity and continuity. The compute layer wants to be ephemeral. The user wants to be at home. Bridging those two worlds used to mean a stack of identity plumbing that quietly grew until it became its own platform. This session changes that math.&lt;/P&gt;
&lt;P&gt;Here is what jumped out for me:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Cloud-only identity for SMB.&lt;/STRONG&gt; Azure Files now authenticates pure Microsoft Entra ID users and groups, including B2B guests, directly over SMB Kerberos. No on-premises Active Directory required, no Entra Connect required, no line of sight to a domain controller.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;NTFS ACLs and Kerberos preserved.&lt;/STRONG&gt; You keep the security model your apps already understand. Permissions still live on the file system, tickets still come over SMB, and FSLogix does not care that the identity stack underneath is different.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Performance built for the spike.&lt;/STRONG&gt; Metadata caching is generally available and rolling out by default. Concurrent file handles per share are moving from 2,000 today to 10,000, with a roadmap toward 30,000 to 50,000. That means fewer storage accounts to shard across when 9 AM hits.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Zonal placement and smarter alerts.&lt;/STRONG&gt; Premium LRS now lets you co-locate the share with its session hosts inside the same availability zone, and new percentage-based metrics finally make alert thresholds portable across shares of any size.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;In short, the boring identity and storage plumbing that propped up VDI for a decade is being collapsed into something you can actually run as a cloud-native service.&lt;/P&gt;
&lt;H2&gt;What This Is, A Technical Overview&lt;/H2&gt;
&lt;P&gt;Let’s set the table. VDI on Azure (whether you run Azure Virtual Desktop, Citrix on Azure, or Omnissa Horizon) uses pooled session hosts. Those hosts are intentionally stateless so they can be patched, scaled, and recycled without ceremony. The user’s identity is “Connie Cloud” today, and on a different host tomorrow.&lt;/P&gt;
&lt;P&gt;FSLogix solves the continuity half of the puzzle. It packages the user’s profile and Office data containers (the profile container and the ODFC, the Office Data Folder Container) as VHDX files that get dynamically attached when Connie logs on and detached when she signs out. Those VHDX files need to live somewhere durable, fast, and reachable over SMB from any host in the pool.&lt;/P&gt;
&lt;P&gt;That is precisely what Azure Files delivers. It is a fully managed SMB file share service that integrates cleanly with FSLogix profile containers and App Attach image stores for AVD. The reference architecture and sizing guidance are documented on Microsoft Learn for anyone who wants the official map.&lt;/P&gt;
&lt;P&gt;The historic friction was identity. Until recently, SMB authentication to Azure Files required either on-premises AD DS joined to the storage account or hybrid identities synced through Entra Connect. That meant keeping domain controllers (and the network paths to reach them) alive purely to satisfy storage authentication. As of this year, Azure Files supports pure Microsoft Entra ID identities for SMB Kerberos, which closes that loop.&lt;/P&gt;
&lt;H2&gt;How It Works, Under the Hood&lt;/H2&gt;
&lt;P&gt;Here is the simplified flow Adam and Priyanka walked through during the demo.&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;The user (an Entra-only account, no on-prem footprint) signs into an Entra-joined AVD session host with single sign-on.&lt;/LI&gt;
&lt;LI&gt;The session host needs to mount the user’s FSLogix profile container from an Azure Files share.&lt;/LI&gt;
&lt;LI&gt;The host requests a Kerberos service ticket. Because the share has Microsoft Entra Kerberos authentication enabled, Entra ID issues that ticket directly, no on-prem KDC involved.&lt;/LI&gt;
&lt;LI&gt;The SMB connection is established, the share-level RBAC role (for example Storage File Data SMB Share Contributor) is checked, and then the directory and file ACLs (standard NTFS) are evaluated.&lt;/LI&gt;
&lt;LI&gt;FSLogix attaches the VHDX, the profile loads, Outlook is happy, OneDrive is happy, and Connie’s pinned taskbar shows up exactly the way she left it.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;A few details worth filing away:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Two layers of authorization.&lt;/STRONG&gt; Share-level access uses Azure RBAC roles. Item-level access uses NTFS ACLs. Both still apply, which is why your existing permissions model carries over cleanly.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;B2B guest support.&lt;/STRONG&gt; Vendor and contractor accounts that come in as guests in your tenant can be granted access to file shares without needing a synced shadow account.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Metadata caching is the unlock.&lt;/STRONG&gt; VDI is metadata-heavy: directory enumerations, file opens, renames, and closes hammer the share at logon. Metadata caching reduces P50 latency on those operations by roughly 80 to 90% and roughly doubles metadata transaction throughput, which is what makes the higher concurrent handle limits realistic. The full SMB performance reference on Microsoft Learn lays out the knobs.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Zonal placement.&lt;/STRONG&gt; Premium LRS lets you pin the share to the same availability zone as your session host pool, so the SMB traffic does not bounce across zones.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Real-World Value&lt;/H2&gt;
&lt;P&gt;Where does this show up in your operations review?&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Retire orphan domain controllers.&lt;/STRONG&gt; Plenty of shops have a couple of DCs in Azure that exist only so Azure Files can authenticate. Cloud-native Entra ID lets you turn those off and shrink the identity attack surface.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Simpler M&amp;amp;A and vendor onboarding.&lt;/STRONG&gt; Adding a partner organization or a new acquisition no longer requires forest trusts or a sync project. Invite guests, assign them to a group, grant the group access to the share.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Fewer storage accounts and shares to manage.&lt;/STRONG&gt; Higher concurrent handle limits mean you can consolidate users that you previously had to spread across many accounts just to dodge the 2,000-handle ceiling. Less sprawl, less monitoring, fewer naming conventions to remember.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Predictable logon times at scale.&lt;/STRONG&gt; Metadata caching is the kind of feature you only notice when it is missing. With it on by default, large host pools see flatter logon latency curves during the morning rush.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Operational consistency.&lt;/STRONG&gt; Percentage-based metrics let you set a single rule like “alert at 10% remaining capacity” and apply it cleanly to a 5 TiB share and a 100 TiB share without bespoke thresholds.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;In short, the ROI conversation moves from “how do we keep VDI running” to “how much of the supporting cast can we delete.”&lt;/P&gt;
&lt;H2&gt;Getting Started&lt;/H2&gt;
&lt;P&gt;If you want to kick the tires this week, here is a practical starting path.&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;Inventory your VDI identity story.&lt;/STRONG&gt; Are you running hybrid because the apps need it, or because Azure Files used to need it? If it’s the second one, you have a candidate workload for cloud-only identity.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Spin up a pilot Premium SSD Azure Files share&lt;/STRONG&gt; in the same region (and ideally the same availability zone) as a small AVD host pool.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Enable Microsoft Entra Kerberos authentication&lt;/STRONG&gt; on the storage account. The configuration is now in the standard Azure portal, no more side trips to the fileperms portal.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Assign Azure RBAC roles&lt;/STRONG&gt; at the share level (Storage File Data SMB Share Reader, Contributor, or Elevated Contributor as appropriate) to your Entra groups.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Set NTFS ACLs&lt;/STRONG&gt; on the directories that will host FSLogix containers, and point FSLogix at the share’s UNC path.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Test with a cloud-only user&lt;/STRONG&gt; (no on-prem identity at all) to confirm the end-to-end flow.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Turn on metadata caching and the new metrics&lt;/STRONG&gt; and set percentage-based alerts so you find the limits before your users do.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2&gt;Resources&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/storage/files/" target="_blank"&gt;Azure Files documentation&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/storage/files/virtual-desktop-workloads" target="_blank"&gt;Use Azure Files for virtual desktop workloads&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/storage/files/storage-files-identity-auth-hybrid-identities-enable" target="_blank"&gt;Enable Microsoft Entra Kerberos authentication for hybrid and cloud-only identities on Azure Files&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/storage/files/smb-performance" target="_blank"&gt;Improve performance for SMB Azure file shares (metadata caching, multichannel, handle limits)&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/storage/common/redundancy-premium-file-shares" target="_blank"&gt;Data redundancy for Premium file shares&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Keep Learning at the Summit&lt;/H2&gt;
&lt;P&gt;Catch the full Microsoft Azure Infra Summit 2026 session playlist here: https://www.youtube.com/playlist?list=PLjt5SKzX1iI8con7FJDB56G6hHqxGm7ki&lt;/P&gt;
&lt;P&gt;Cheers!&lt;/P&gt;
&lt;P&gt;Pierre Roman&lt;/P&gt;</description>
      <pubDate>Wed, 15 Jul 2026 07:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/itops-talk-blog/modernize-vdi-with-azure-files-and-entra-cloud-native-identities/ba-p/4534545</guid>
      <dc:creator>Pierre_Roman</dc:creator>
      <dc:date>2026-07-15T07:00:00Z</dc:date>
    </item>
  </channel>
</rss>

