pierre roman
127 TopicsZonal Resiliency in Azure: Application-Centric Goals, Recovery Plans, and Drills
Hello Folks If you have ever stared at a multi-tier app in Azure and asked yourself, “Is this actually going to survive a zone outage?”, you are not alone. In session MAIS23 of the Microsoft Azure Infra Summit 2026, Bhavya, Aditya, and Chaya from the Azure Resiliency product team walked us through the new Resiliency in Azure experiences (formerly Azure Business Continuity Center) and showed how to stop treating resiliency as a per-resource checkbox and start treating it as an application-level outcome. Why IT Pros Should Care Most of us have lived this story. An app is “in the cloud”, spread across IaaS VMs, PaaS databases, an app service plan, and a shared Azure Firewall managed by some other team. Then a zonal blip hits, and suddenly nobody can answer the simple question: was this app supposed to be zone resilient or not? The session opened with a customer scenario called Zava, a fast-growing insurance company running a claims app at 99.9 percent availability that just lost more than $40,000 in revenue in one week because of zonal outages. That is the price tag the speakers put on the problem, and it lines up with the patterns I see every week. Here is why this matters to IT pros: You finally get a single pane to see zonal resiliency posture across IaaS, PaaS, and shared services. Resiliency goals are set at the application level, not buried inside each resource blade. You get tailored Azure Advisor recommendations plus an Azure Copilot guided flow that emits remediation scripts. You can run zone-down drills powered by Azure Chaos Studio without stitching together five different tools. Recovery plans orchestrate failover in a defined order, with on-demand readiness checks before the next real outage. In short, less guessing, less spreadsheet bookkeeping, and a lot more confidence that the app will behave the way you told the business it would. What Resiliency in Azure Does, a Technical Overview The team has rebranded Azure Business Continuity Center to Resiliency in Azure. It is a unified solution that covers infra, data, and cyber resiliency in one place. Today the focus is zonal resiliency, with regional disaster recovery (and proper RPO/RTO goals) on the roadmap. The central concept is the service group. A service group is a logical application unit that can span subscriptions and resource groups. You add the VMs, databases, app service plans, Redis caches, and other Azure resources that make up an application, and from that point on, resiliency operations work against the whole app, not one resource at a time. There are two views you will spend most of your time in: Resource resiliency, a zonal configuration summary across the (roughly 20) resource types supported today. Service group resiliency, the same summary but pivoted to the application level, so you can prioritize the apps that need attention first. The speakers were honest about scope. Goals today are a simple intent (“this service group should be evaluated for zonal resilience”). Once additional pillars like regional DR ship, goals will expand to include RPO and RTO targets. I appreciate that they did not oversell it. How It Works, Under the Hood Once a service group exists, the workflow has three big building blocks. Each one solves a problem I bet you have hit. Goals and recommendations. You assign a zonal resiliency goal to the service group, and Azure Advisor surfaces tailored recommendations for the resources inside it. Two details I liked: The view shows cost implications before you flip the switch. Some Azure services have no cost delta for zone redundancy. Others do. You see it inline, not in a separate calculator tab. There is an Azure Copilot guided remediation flow that walks you through the recommendation and, at the end, emits a script. That script accounts for resource-type corner cases (SKU changes, redeploys, and so on) and is meant to be run through your automation pipeline. You can also exclude a resource with a reason (“not critical, zonal redundancy not required”) or manually attest a resource when your own custom solution already provides resiliency that the platform cannot auto-detect. That escape hatch is important, because real environments always have a few weird cases. Application-centric recovery plans. Instead of failing over one resource at a time, a recovery plan orchestrates the entire app. It auto-detects existing solutions (Azure Site Recovery for VMs, for example), lets you group and order the resources for failover, and excludes resources that are already configured for high availability (no point failing them over if they did not go down). You can run an on-demand readiness check any time the app structure changes, so you find configuration drift before an outage finds it for you. Zone-down drills powered by Azure Chaos Studio. A zone-down drill template identifies the service group resources, pre-populates the right native faults per resource type (think a Redis cache fault, a VM scale set shutdown, and so on), bundles in identity and permission checks, monitoring, and the recovery plan you already built. When you execute, you pick the region and the target zone, the drill runs a pre-validation check, injects the fault, runs failover, then reprotection and failback, and tracks all of it as a single job in the execution report. Per-resource metrics let you visualize the actual downtime each component experienced. If a native fault is not what you want, you can override with a custom runbook. That last point is the part I think a lot of folks miss. A drill is not just fault injection. It is fault injection plus failover plus reprotection plus failback, all measured and attestable in one place. Real-World Value Back to Zava. They needed to answer three questions: what is our current zonal resiliency posture across these Azure services, what should we prioritize against our 99.9 percent target, and how do we validate that we will actually perform during an outage? Resiliency in Azure answers all three without forcing the platform team to write a 200-line PowerShell script. Use cases that should be on your shortlist: Regulated workloads (insurance, healthcare, financial services) that need to evidence drills for compliance. The notes and manual attestation features were clearly designed with auditors in mind. Apps with mixed estates, where a central platform team owns shared services (firewalls, identity) and app teams own everything else. Service groups can be parented to mirror that org structure. Apps with custom resiliency solutions that the platform cannot detect. Manual attestation keeps the dashboard honest without forcing you to refactor. Game-day rehearsals. The pre-built zone-down template means you can run a meaningful drill in an afternoon instead of standing up a custom Chaos Studio experiment from scratch. The honest tradeoff: zone redundancy is not free for every service, and not every resource type is in scope yet (around 20 today). Plan accordingly, exclude what is not critical, and attest what is covered by something else. Getting Started Here is the path I would take on a Monday morning: Open the Azure portal and search for Resiliency. You will land on the Resiliency in Azure page that replaces the old Business Continuity Center. Create a service group. Add resources directly, or add resource groups if each resource group is already an application boundary in your environment. Assign the zonal resiliency goal to the service group. Review the summary tiles. Exclude or manually attest the resources that need it. Walk the Advisor recommendations. Use the Copilot guided flow to generate a remediation script and run it through your automation. Build an application-centric recovery plan, group and order the resources, run an on-demand readiness check. Create a zone-down drill from the template, validate identity, monitoring, and faults, then execute the drill in a non-production zone first. Resources Resiliency in Azure documentation Zonal resources and zone resiliency Azure service groups overview Azure Advisor reliability recommendations Azure Chaos Studio documentation Azure Site Recovery overview Keep Learning at the Summit Catch the full Microsoft Azure Infra Summit 2026 session playlist here: https://www.youtube.com/playlist?list=PLjt5SKzX1iI8con7FJDB56G6hHqxGm7ki Cheers! Pierre Roman79Views0likes0CommentsAgentic Migrations and Modernization: How the Azure Migrate Agent Keeps Your Intent Alive End to End
Hello Folks! If you have ever tried to move a few hundred VMs, a pile of databases, and a couple of web apps from on-prem to Azure, you already know the hard part is not the tooling. The hard part is keeping context, intent, and momentum alive across weeks of planning, hand-offs, and decisions. In session MAIS15 at the Microsoft Azure Infra Summit 2026, Ankur Gupta (Senior Product Manager on the Azure Migrate team) walked us through the new Azure Migrate agent, an AI layer that sits on top of Azure Migrate and carries your intent from “I have an idea” all the way to “the landing zone is deployed.” Why IT Pros Should Care Ankur opened with a line that stuck with me. Infrastructure complexity has far outpaced human scale. We have MySQL here, PostgreSQL there, web apps, storage devices, networking gear, multiple dashboards, multiple alerts, and we are all expected to move faster than ever, with fewer mistakes. In short, migrations rarely fail because someone picked the wrong tool. They fail because the system between the stages breaks. Here is why the agentic approach matters for the folks in the trenches: It keeps context across the entire lifecycle, so the intent you set on day one is still the intent at execution. It guides you when you are stuck, instead of leaving you to figure out which of three discovery methods is the right one. It compresses tasks that used to take days of analysis (think side-by-side business cases) into a few hours. It connects IT ops, architects, and developers through a single thread of information, including a clean handoff to GitHub Copilot for code work. It builds on the Azure Migrate portal you already know, so nothing you have learned goes to waste. That last point is important. The portal does not go away. The agent is a layer on top. You can still do everything you do today. What the Azure Migrate Agent Is, technical overview Azure Migrate has always been Microsoft’s hub for discover, assess, and migrate. What Ankur showed at MAIS15 is the next evolution. Azure Migrate is becoming a migration control plane that spans the whole lifecycle (Decide, Plan, Execute), and the Azure Migrate agent is the conversational, guidance-oriented layer that ties it all together. In Ankur’s words, the agent is educational and guidance-oriented. You ask one natural-language question, like “how should I plan moving my VMware workloads to Azure,” and you get the next steps you actually need to take. Behind the scenes, the agent is doing three things very well. It maintains state across the entire lifecycle. Preferences you set early stay with you. It carries context across discovery, plan, and execute. You can jump around, repeat steps, change your mind, and the agent remembers what your goal was. It recommends the right next move based on what it has learned about your intent. This is the heart of the “agentic” part. The agent is not a chatbot grafted onto a portal. It is a stateful workflow runner that remembers you. How It Works, under the hood The session walked through a full VMware-to-Azure scenario, and the flow is worth seeing because it shows how the pieces snap together. The agent supports three discovery methods today: appliance-based discovery, RV Tools, and the new Azure Migrate collector. The collector is the new lightweight option. It ships as a set of PowerShell scripts you run on a machine that can reach your vCenter, it produces a zip file, and you upload that zip to your Azure Migrate project. No appliance to deploy, no inbound network plumbing. Once the inventory lands, the agent reads it. Ankur asked for a summary and got a card showing 207 VMs, 177 SQL databases, one PostgreSQL instance, and some web apps. He then asked for all servers with an out-of-support OS, got a list of around 50, and tagged them right inside the conversation so he could refer back to them later. Next came the business case. Ankur asked the agent to generate one based on a modernize preference. A few minutes later, he had Azure cost, on-prem cost, and projected savings. Then he asked for a second business case for lift and shift, and a side-by-side comparison. The agent ran it, showed that the on-prem cost in the lift-and-shift comparison was higher and that lift-and-shift TCO savings were actually higher in that specific scenario, and gave him the data points he needed to bring the decision to leadership. From there, Ankur moved to application assessment, this time back in the portal. He created an assessment for two apps (Airsonic and Parts Unlimited), let the high-confidence plan run, and got a modernize recommendation with 100% readiness and a target cost of about $580 per month, plus an emissions estimate of 32 kgs of CO2. App Service for the web tier, Azure Database for PostgreSQL for the data tier. Both were flagged “ready with conditions,” with clickable links into why. Then came one of my favorite parts. Ankur connected GitHub for a Copilot Assessment, which adds code-level insights on top of the infrastructure readiness assessment. The system recalculated, and the migration effort estimate sharpened up. Finally, the agent built a wave plan from the assessment, then generated a platform landing zone aligned with Azure best practices. He could ask the agent about chosen defaults, request changes to the deployment mechanism, swap in a third-party firewall, or apply naming conventions. The agent produced a downloadable Infrastructure-as-Code template and handed it to a cloud architect, who refined it in their IDE using GitHub Copilot. That last handoff is the bridge between Azure Migrate’s planning world and the developer world. Real-World Value (use cases, ROI, scenarios) So where does this actually pay off? A few scenarios stood out. Pitching the business case to leadership. Ankur framed the demo around “I need to pitch a migration proposal to the planning committee.” Generating modernize, lift-and-shift, and Azure VMware Solution business cases used to be days of spreadsheet work. With the agent, it is hours. Cleaning up legacy debt. Tagging out-of-support servers in one conversational step lets you plan upgrades without exporting CSVs and slicing them by hand. Mixed estates with web apps and databases. The agent surfaces App Service and Azure Database for PostgreSQL targets, gives SKU recommendations, and flags the warnings worth investigating. Closing the IT-to-developer gap. The GitHub Copilot Assessment and the IaC handoff to the IDE means developers and architects work from the same context. Reducing intent drift on long migrations. Multi-week journeys lose their original intent. The agent remembers. In short, the ROI here is measured in calendar time, not just dollars. And honestly, in fewer late-night calls when something goes sideways because nobody remembered the original decision. Tradeoffs worth flagging: the agentic capabilities are landing in preview, and outputs are advisory. You still need human review, testing, and governance on every recommendation. That is by design. Getting Started (concrete first steps) Here is a practical onramp. Stand up an Azure Migrate project in the Azure portal if you do not already have one. Pick a discovery method that fits your environment. If you cannot deploy an appliance, try the new collector. Download the PowerShell scripts, run them from a host that can reach vCenter, and upload the zip. Bring in the Azure Migrate agent from inside the portal and ask it to summarize your discovered inventory. Generate at least two business cases (modernize and lift-and-shift). Compare them. Run an application assessment on a small, representative set of apps. Connect GitHub and add a Copilot Assessment for code-level insight. Ask the agent to build a wave plan and a platform landing zone template, then push the IaC to your repo for the architects. Start small, build the muscle, and scale out. Resources Azure Migrate documentation (Microsoft Learn) About Azure Migrate, including the Azure Copilot migration agent (Microsoft Learn) GitHub Copilot modernization overview (Microsoft Learn) GitHub Copilot modernization agent overview (Microsoft Learn) GitHub Copilot modernization documentation (Microsoft Learn) Assess and migrate a .NET project with GitHub Copilot modernization (Microsoft Learn) GitHub Copilot modernization overview for .NET (Microsoft Learn) Keep Learning at the Summit Catch the full Microsoft Azure Infra Summit 2026 session playlist here Microsoft Azure Infra Summit 2026 Cheers! Pierre Roman73Views0likes0CommentsFrom Alert to Resolved: Building a Self-Healing Azure Platform with SRE Agent
Hello Folks! It’s 3 a.m. Your phone lights up. A critical workload that spans multiple clouds is on fire, ownership is fuzzy, the alert routed to the wrong team first, and now it’s your problem. You sit up in bed, cold and groggy, and start the ritual. Open the runbook. Pull logs from one place. Pull metrics from another. Stare at three dashboards. None of them tell the whole story. So you build a theory. The theory is wrong. The clock keeps ticking. The customer impact keeps climbing. Every wrong turn costs you time, context, and confidence. That is the scene Lee Oommen opened with at MAIS14, and it is the reason Azure SRE Agent exists. In this session, Lee walks through the four classic SRE pain points and shows how an agentic operations platform compresses MTTR from hours to minutes. I am going to unpack what he showed, why it matters for IT pros, and how to get your hands on it. Why IT Pros Should Care If you carry a pager, write runbooks, or get pulled into post-mortems, this one is for you. The clock is the enemy, not the incident. Lee said it plainly: there is almost always an expert who can fix the problem. The real damage comes from the minutes spent finding that expert and reconstructing context. Dashboards lie in both directions. False positives create alert fatigue. False negatives let the customer call you before your monitors do. Neither outcome is acceptable. RCAs take weeks because the answer never lives in one layer. Infrastructure, network, deployments, dependencies, databases, app code. You need someone, or something, that can correlate across all of them in one pass. You did not become an SRE to be a dashboard watcher. Toil is the work that holds back the people who should be designing reliability into the next generation of services. What is the SRE Agent Let’s start with what it is not. It is not a dashboard. It is not a monitoring tool. It is not a chatbot. Azure SRE Agent is an end-to-end agentic operations platform. Think of it as a senior SRE who sits inside your team, works 24 by 7, never gets tired, never misses a signal, and is fluent in your stack. Reasons over telemetry, not just text. It pulls metrics, logs, traces, deployment history, and activity logs, then correlates across them. Takes governed actions. Every action runs inside the permission boundary you define. You decide whether the agent proposes a fix, asks for approval, or acts autonomously. Authors RCAs in minutes, not weeks. It traces the root cause in a single flow across the entire stack and produces the report immediately after remediation. Remembers. It captures organizational memory from every incident, every chat, and every scheduled task, then applies it to the next investigation. Lee called it the operations half of DevOps, and that framing stuck with me. We have automated build and deploy. The operations side has been stuck in toil. SRE Agent closes that loop. From Alert to Resolved (the workflow) Lee demonstrated the full loop live. Here is what it looks like end to end. Detect. The agent integrates with Azure Monitor, PagerDuty, or ServiceNow. When an alert fires, the agent acknowledges it within seconds. The human does not have to wake up cold. Investigate. The agent runs diagnostics in parallel across the connected resources. It queries Log Analytics, App Insights, Azure Monitor metrics, and any third-party observability tools you wired in through MCP connectors. Correlate. It uses distributed tracing, cross-workspace KQL queries, and time-based signal alignment to connect dots across services that do not even share trace IDs. It also checks past incidents in memory to see if this looks familiar. Diagnose. It produces a root cause analysis with the relevant evidence linked inline. No more reconstruction exercise across multiple teams. Propose or Act. Based on your run mode and the permissions granted, the agent either proposes a fix and waits for approval, or executes the remediation autonomously. Lee demonstrated both. He set up a bad slot swap on Azure App Service, generated HTTP 500 errors, watched the agent acknowledge the alert, investigate, identify the bad slot, ask for permission, and then perform the slot swap to restore the service. Close the loop. The agent files a GitHub or Azure DevOps issue with full context, opens a pull request with proposed code changes when appropriate, and writes a session insights summary you can review. Three modes to interact with it: Interactive. Chat with the agent like a copilot. Most customers start here to build trust. Reactive. Event-driven. The agent reacts to incidents from Azure Monitor, PagerDuty, or ServiceNow. One agent per incident platform. Proactive. Scheduled tasks that run every five minutes, every hour, daily, or weekly. Certificate health audits, well-architected framework assessments, cost optimization sweeps, compliance checks. Lee showed a scheduled task that flagged a certificate expiring in 50 days before it could ever fire an alert. Real-World Value This is where the conversation gets practical. A few things from Lee’s demo and the live Q&A that I want to call out. Multi-cloud reality. SRE Agent lives in Azure but is not limited to Azure. Custom runbooks, Python execution, MCP servers, and connectors let it orchestrate across AWS, GCP, and on-premises. Treat it as the central SRE brain. Your data stays yours. Each agent gets a dedicated data store in your subscription and resource group. Memory, knowledge, threads, and session insights live in your chosen region. Nothing is used to train the model provider. Encryption at rest, TLS 1.2 in transit, Azure RBAC, managed identity, customer-managed keys all apply. Identity boundary you already know. The agent uses standard Azure managed identity. Grant the identity RBAC on any cross-subscription resource it needs to reach. Least privilege still applies. Region availability. At session time, agents can be deployed in EastUS2, Sweden Central, and Australia East. The list is updating roughly monthly. Canada is coming. An agent in one region can act on resources globally, but if you have data residency rules, deploy the agent inside the same jurisdiction. Private endpoints today. If your Log Analytics Workspace or databases are fully locked down behind private endpoints with public access disabled, the agent currently needs a VNet-integrated Azure Function as a proxy. Microsoft is actively working on injecting agents directly into private networks. Memory is the multiplier. A principal engineer is more valuable than a junior engineer because of pattern recognition. SRE Agent captures that pattern recognition for the whole team, every time it investigates. Getting Started The pattern is simple, and Lee summarized it cleanly: you teach the tool, you make the connections, and it works for you. Provision the agent. Go to sre.azure.com or the Azure portal, pick a subscription and resource group, pick a region, and stand it up. Takes a few minutes. Onboard it like a new engineer. Tell it about your team, your workloads, and your procedures. Upload runbooks, troubleshooting guides, wikis, and architecture docs to the knowledge base. If you do not have these documents, ask the agent to draft them for you. Connect your observability stack. Azure Monitor, Log Analytics, App Insights are wired in by default. Add third-party tools through MCP connectors. Wire in your incident platform. Azure Monitor, PagerDuty, or ServiceNow. One agent per platform. Grant code access. Connect your GitHub or Azure DevOps repositories so the agent can reason over application code, propose fixes, and open pull requests. Pick your run mode. Start in interactive mode while you build trust. Move to approval-gated reactive mode. Graduate to autonomous mode on safe operations once you have the audit trail you trust. Resources Azure SRE Agent documentation on Microsoft Learn Azure SRE Agent product docs Get Started guide Automate incident response Official Microsoft SRE Agent GitHub repository (issues, labs, resources) Watch the Rest of the Summit If you found this useful, the rest of the Azure Infra Summit 2026 is packed with sessions on identity, AKS, deployment, storage, networking, and resiliency. Grab the full playlist here and binge what is relevant to your stack:Microsoft Azure Infra Summit 2026 Big thanks to Lee Oommen for walking us through this. The 3 a.m. pager scenario is something every one of us has lived, and seeing an agent take the first hour of that incident off your plate is a tangible win. Cheers! Pierre Roman99Views1like0CommentsDesigning Azure Networks That Scale: From Small Deployments to Enterprise-Grade
Hello Folks! If you have ever spent a long afternoon untangling overlapping CIDR ranges, chasing down a broken VNet peering, or trying to remember which UDR points to which firewall, this MAIS 2026 session is going to feel uncomfortably familiar. Jon Ormond (Principal PM, Azure Networking) brought along Jay Li and Jeff Lovett from the Azure Networking team to walk through what actually happens when an Azure network grows from a handful of VNets into a real enterprise estate, and where most teams hit the wall. The headline they kept coming back to is simple. Azure networks do not usually fail because they were built wrong on day one. They fail because they did not evolve fast enough. Scale is not a smooth ramp. It is a step function, and every step adds an order of magnitude of complexity. Why IT Pros Should Care You may be running three VNets today. That is fine. But the day a second team shows up, or you cross into a second region, or somebody asks for hybrid connectivity to the datacenter, your operating model changes whether you planned for it or not. The session is built around two pivots every growing Azure environment hits: Management and control inside Azure (VNets, peerings, routes, security rules). Connectivity and hybrid (VPN, ExpressRoute, Virtual WAN, reliability). Both of those break quietly. By the time you notice, you are already firefighting drift, broken peerings, or unpredictable latency from on-prem. Bottom line, here is what you take away: Design for the next stage, not the one you are in. Put the management layer in before complexity outpaces manual effort. Treat reliability as a design choice, not an afterthought. Start Small, Plan to Grow One VNet, one subnet, one workload. Nothing wrong with that. You can manage it with the portal, a spreadsheet for CIDR tracking, and a calm heart. The problem is that the jump from “one VNet” to “a few VNets across teams” is not gradual. As soon as you have a second team that needs isolation, you are into hub and spoke territory. Ten spokes feels manageable. Fifty spokes across multiple subscriptions does not. And by the time you hit a hundred, the spreadsheet is a liability. Jay made the case that the smartest move at small scale is not to stay manual until it hurts. It is to put Azure Virtual Network Manager (AVNM) in early, even if you only have three VNets. AVNM lets you declare intent once and let the platform handle the rest: IP address management (IPAM) so new spokes get non-overlapping CIDRs automatically. Network groups with tag-based dynamic membership so VNets land in the right group the moment they exist. Connectivity (hub and spoke or mesh) without hand-built peerings. Security admin rules pushed centrally across the estate. Routing intent so traffic flows through the right firewall by default. The honest tradeoff: AVNM is one more thing to learn and operate, and it adds cost. The counter-question Jay kept asking is, “What is the cost of drift?” One overlapping CIDR or one missing UDR at 100 VNets can cascade into an outage that takes days to unwind. That is the real tradeoff. Mid-Stage Patterns: Hub and Spoke, Peering, and the First Cracks The hub and spoke topology is the workhorse of Azure networking and the pattern the Cloud Adoption Framework recommends for most enterprises. It centralises shared services (firewall, DNS, ExpressRoute and VPN gateways, Private DNS zones) in a hub VNet, and connects spoke VNets through peerings. Where teams get into trouble at this stage: Peering sprawl. Every new spoke needs a peering, sometimes two if you want transitive paths. Doing this by hand across subscriptions is where human error lives. Route table drift. UDRs copied from spoke to spoke get out of sync. One spoke routes through the firewall, another bypasses it. Now you have a compliance problem. Security rule drift. NSGs and security policies start as a copy paste exercise and end as a forensic exercise. CIDR collisions. “Just give me a /24” turns into a multi day investigation when the new spoke overlaps with on-prem. Jay’s point on this was sharp. The mistake is not the topology. Hub and spoke is the right pattern. The mistake is staying manual on top of it. AVNM network groups let you say, “any VNet tagged environment=production joins the production group, gets the production security baseline, peers to the production hub, and inherits the routing intent that sends east-west traffic through the firewall.” No tickets, no copy paste, no drift. If you are already deployed via Azure Landing Zones (ALZ) with Bicep or Terraform, AVNM is not a replacement, it is another construct in your template. As Jon put it in the chat, it is “just another object” in your ALZ, and the two layers work together rather than competing. Enterprise Scale: Virtual WAN, Segmentation, and Governance At some point hub and spoke stops scaling cleanly. You start adding regions. Branch offices show up. You need SD-WAN integration, more than 30 IPsec tunnels, or transitive routing between VPN and ExpressRoute. That is when Microsoft pushes you toward Azure Virtual WAN. Virtual WAN is a Microsoft managed global transit network. You deploy regional virtual hubs and connect everything (Azure VNets, branches, remote users, ExpressRoute circuits) into them with consistent routing and security. The trade up is real: Any to any connectivity by default. Hub to hub mesh is built in. Routing intent and policies for centralised internet egress and east-west inspection through Azure Firewall or a partner NVA in a secured hub. Branch scale. Tens or hundreds of sites stop being a custom integration project. Operational simplification. Microsoft owns the hub control plane so you stop babysitting peerings. For hybrid connectivity itself, Jeff walked the curve every customer travels: VPN Gateway is the on-ramp. Cheap, fast to stand up, good enough until public internet latency, throughput, or regulatory requirements force a change. ExpressRoute circuits give you dedicated bandwidth from 50 Mbps to 100+ Gbps, with predictable performance and over 200 service providers worldwide. Scalable ExpressRoute virtual network gateways grow and shrink with usage, so you deploy once and stop re-architecting every time traffic changes. ExpressRoute Metro is the headliner. Same price as a standard circuit, but the redundant device lives in a second, physically distinct co-location facility across town. Building fire, flood, or power outage in one site, and your traffic keeps flowing. Multiple circuits are still on the table when “this cannot fail” actually means it cannot fail. Honest tradeoff on Virtual WAN: it is opinionated, Microsoft managed, and you give up some of the granular control you have in a customer managed hub. For most enterprises that is a win. For the few with very specific routing requirements or heavy NVA investments, traditional hub and spoke with Azure Route Server can still be the right call. The CAF guidance lays this out in detail. Getting Started If you take one thing from this session, take this. Design for the next stage. Three concrete moves: Stand up AVNM now, even at small scale. Declare your intent for IPAM, connectivity, security, and routing once. Let new VNets inherit it. Pick your topology with eyes open. Hub and spoke for customer managed control, Virtual WAN for Microsoft managed global transit at scale. The CAF decision tree is the right starting point. Plan hybrid for failure, not for the sunny day. ExpressRoute with Metro by default. Multiple circuits for the workloads that genuinely cannot go down. Test the failover. Resources Azure Virtual Network Manager overview Azure ExpressRoute introduction About ExpressRoute virtual network gateways About Azure VPN Gateway About Azure Virtual WAN Hub-spoke network topology in Azure Define an Azure network topology (Cloud Adoption Framework) Virtual WAN network topology in an Azure landing zone Watch the Rest of the Summit This was one of many great sessions at the Microsoft Azure Infra Summit 2026. If you want to catch the keynotes, the deep dives on storage and AKS, and everything in between, the full playlist is here: Microsoft Azure Infra Summit 2026 Playlist Big thanks to Jon Ormond for moderating, and to Jay Li and Jeff Lovett for the practical, no-fluff walk through what actually breaks at scale and how to design ahead of it. Cheers! Pierre Roman179Views0likes0CommentsAzure Files, Reimagined: Top-Level Shares with Per-Share Networking, Billing, and Scale
Hello Folks! If you have ever wrestled with Azure Files inside a storage account, juggling shared RBAC, shared networking, and shared IOPS across a pile of shares that really should not live together, this session is going to address all that. During Microsoft Azure Infra Summit 2026, Vincent Du and Will Gries (both Product Managers on the Azure Files team) walked us through the new Microsoft.FileShares resource provider, a management model that promotes the file share itself to a top-level Azure resource. Why IT Pros Should Care For years, file shares lived inside a storage account, and that storage account dictated a lot of decisions for you. If one team needed a private endpoint and another needed a service endpoint, you either compromised or you created another storage account. If one share got hot and consumed all the IOPS, the other shares felt it too. Vincent and Will are on the team that built the new model to remove that compromise. Here is what changes for you as an IT pro: Each file share is its own Azure resource with its own RBAC, networking, billing, IOPS, and throughput. Per-share cost shows up directly in Azure Cost Management’s per-resource view, no more Excel guesswork. Encryption in transit is on by default for NFS shares, at no extra cost. Provisioning is dramatically faster. In their head-to-head demo, 200 shares finished in about 50 seconds on the new model versus about 720 seconds with the classic flow. A new MCP server lets you create and manage shares from GitHub Copilot in VS Code with natural language. In short, the new model trades the storage-account-as-gatekeeper pattern for something that feels a lot more like the rest of Azure (think VMs and disks, where the resource you care about is the resource you actually manage). What Microsoft.FileShares Does, a Technical Overview The new Microsoft.FileShares resource provider lets you deploy a file share without first standing up a storage account. When you go into the Azure portal, search for “File share,” and click create, you fill out a single create blade with the things that actually matter for that share: name, region, redundancy (LRS or ZRS), provisioned capacity, IOPS and throughput, networking, and tags. Microsoft Learn confirms the provisioned capacity range is 32 GiB to 262,144 GiB, and only LRS and ZRS redundancy are available at launch (see the Create a file share doc linked below). At GA, the new experience supports NFS 4.1 on the SSD media tier. SMB support, HDD support, customer-managed key encryption at rest, soft delete, and the AKS CSI driver integration are all on the roadmap and called out as the most-requested follow-ups. If you need those features today, the classic file share inside a storage account is still there for you. In the portal, Vincent showed off a small but meaningful detail: the icon color changed from blue (classic) to purple (new). It is a small thing, but when you are scanning a resource group, that visual cue saves you a click. How It Works Under the Hood The new model is built on the provisioned v2 billing structure. Microsoft Learn describes provisioned v2 as a billing model where you independently provision storage, IOPS, and throughput, and you pay for what you provision regardless of how much you actually use. This is a real shift from the older provisioned v1 model, where IOPS and throughput were a function of how much storage you provisioned. Will walked through the math. In his example, provisioning 14 TiB of storage on v1 gave 17,000 IOPS, about 1.5 GB/s throughput, and a bill of roughly $2,297. Moving to v2 with the exact same numbers was already noticeably cheaper. Then, because v2 lets you tune storage, IOPS, and throughput separately, he provisioned the exact storage he needed with slightly less IOPS and throughput, dropping the bill to roughly a third. For database-hot workloads you can dial IOPS up; for hot archive scenarios you can dial them down to the minimum. That kind of flexibility is genuinely useful. Encryption in transit deserves its own callout. The new shares default to encrypted NFS mounts using the AZNFS mount helper. Microsoft Learn explains that AZNFS wraps the NFS connection in a Stunnel-based TLS tunnel using AES-GCM, so you get TLS protection without needing Kerberos or external authentication. The helper installs cleanly on Ubuntu, RHEL, SUSE, Rocky, Oracle Linux, Alma Linux, and Azure Linux. If a workload genuinely cannot use the encrypted mount, you can uncheck the box and fall back to a traditional NFS mount. Networking is per share. You can attach a service endpoint or a private endpoint to each individual share, which means you can put a strict private-endpoint-only share next to a service-endpoint share for dev/test, all in the same resource group, without compromise. On the request side, classic shares throttle with a fixed window (you can burst, then you are locked out for the rest of the window). The new model uses a token-bucket algorithm (the same one Azure Resource Manager itself uses), which means you get a sustained refill rate. The team also gave you a separate delete bucket, so a big cleanup operation does not starve writes. That detail matters more than it sounds: batch cleanups against the classic model regularly crowd out new share creation. Real-World Value Where does this actually pay off? A few honest scenarios: Mission-critical and regulated workloads. A healthcare org with workloads at different sensitivity levels can put strict private-endpoint-only shares next to less sensitive service-endpoint shares without the storage-account ceiling. Chargeback and showback. With per-share resources, finance can pull a cost report that lines up to the team or project that owns each share. No more saying “we cannot itemize, the storage account is shared.” High-density tenants. The classic model effectively caps you at 34 file shares on an SSD provisioned v2 storage account (because of IOPS minimums) and 50 absolute. The new model goes up to 10,000 shares per subscription per region. That is a different game. Tuned database and analytics shares. Provisioned v2 lets you right-size IOPS to the workload. As Will showed, that can drop the bill to roughly a third for the right shape of workload. Faster deployment automation. A 14x improvement on a 200-share deployment is not a micro-optimization. If you spin up environments for CI, training, or per-customer tenants, that adds up quickly. The honest tradeoff: today, the new model is NFS-only on SSD. If you need SMB, HDD, customer-managed keys for NFS, or AKS CSI driver support, stay on the classic model for now. The team was upfront about that, and the GA-and-then-iterate roadmap is clear. Getting Started Here is the concrete path: Register the Microsoft.FileShares and Microsoft.Storage resource providers on your subscription (Subscriptions, Resource providers, Register). From the Azure portal, search for “File share” in the marketplace and click Create. Pick LRS or ZRS, set the capacity between 32 GiB and 262 TiB, and either accept the recommended IOPS/throughput or set them manually. On the Advanced tab, leave “Require encryption in transit” enabled (it is on by default) and pick a custom mount name if you want one distinct from the resource name. On the Networking tab, attach a service endpoint or a private endpoint, per share. Mount it on your Linux VM with the AZNFS mount helper. The portal generates the exact command for your distribution. If you live in IaC land, the Microsoft.FileShares ARM and Bicep types are available, and Terraform support is coming. If you live in AI-assisted dev land, install the Azure MCP server and ask Copilot in VS Code to create a share for you, pointing at an existing VNet. Resources Create an Azure file share with Microsoft.FileShares Understand Azure Files billing (provisioned v1 and v2) Encryption in Transit for NFS Azure file shares NFS file shares in Azure Files (protocol overview) Azure Files documentation home Keep Learning at the Summit Catch the full Microsoft Azure Infra Summit 2026 session playlist here. Cheers! Pierre Roman270Views0likes0CommentsCut Your Azure Blob Storage Bill in Half: A Practical Walkthrough of Object Storage TCO
Hello Folks! If you have ever opened your monthly Azure invoice, stared at the object storage line, and quietly wondered how it grew so much, this one is for you. At the Microsoft Azure Infra Summit 2026, Benedict Berger and George Trossell from the Azure Storage Engineering team walked through a real customer scenario and showed how to bring that bill down without touching a single application. 📺 Watch the session: Why IT Pros Should Care Storage is one of those services we configure once at account creation, then never revisit. Redundancy, default tier, lifecycle rules. All decided on day one, then forgotten. Meanwhile, applications get built on top, dashboards get wired up, and the bill keeps climbing in a department nobody really audits. Here is what you get when you make storage TCO a first-class part of your operating model: A defensible understanding of capacity, transactions, and data retrieval charges (the three real cost drivers). Fewer surprise spikes when a cool tier read pattern runs hotter than expected. Cost optimization that runs on its own, instead of a quarterly cleanup project nobody volunteers for. Storage standards baked into your Infrastructure as Code, so cost-efficient defaults travel with every new account. In short, this is one of the highest-leverage cost levers you have in Azure. And unlike compute right-sizing, you can act on most of it from the portal in an afternoon. What Storage TCO Actually Means on Object Storage, a Technical Overview When Benedict and George talk about Total Cost of Ownership on Azure Blob Storage, they mean four moving parts: Capacity. The per-gigabyte cost of the data you store, which varies by access tier and by redundancy. Transactions. Every read, write, list, and metadata call against the storage account. Priced in packages of 10,000 operations. Data retrieval. A per-gigabyte fee that applies when you read from cool or cold tiers. It is free on hot. Network egress. The charge for moving data out of an Azure region. The trap most teams fall into is looking only at the per-gigabyte capacity column and picking the cheapest tier they see. That ignores the fact that as data gets cooler, transaction and retrieval costs climb sharply, and cold has a 90-day early deletion penalty that can erase your savings outright. Microsoft Learn documents this trade-off clearly in the access tiers overview, where you can see the minimum retention windows and the relationship between storage cost and access cost across hot, cool, cold, and archive. Redundancy is the other dial. LRS keeps three copies in a single zone. ZRS spreads three copies across three zones in the region. GRS adds an asynchronous secondary in a paired region. The honest tradeoff George highlighted: redundancy protects your data, not your application. If your app is not zone-aware, ZRS alone will not keep you running through a zone outage. And GRS failover is a manual operation in most cases, with the secondary in read-only mode until you stand up new accounts to write into. How It Works, Under the Hood The session walked through a worked transaction example that finally made the math click for me. Picture a Spark job uploading 1,000 parquet files of 5 GB each into the hot tier, using an 8 MB block size. Each 5 GB file is roughly 5,120 MB, divided by 8 MB blocks, which gives 640 put block operations. One additional put block list call commits the upload, so each object costs 641 write operations. Times 1,000 files, that is 641,000 operations, which works out to about 3.52 US dollars in that hour just for writes. Now flip it. Read those same 1,000 files from the cool tier. The transaction count is similar, but you also pay a data retrieval fee on every gigabyte you pull back. That retrieval fee is where most teams get blindsided, because it does not show up on the hot tier at all. Block size matters too. Larger blocks mean fewer transactions per upload. And for small objects (under 128 KB), there is a new wrinkle to plan for: starting July 2026 for existing accounts and already in effect for new accounts created from July 2025, cooler tiers bill a 128 KB minimum object size. That means a 4 KB log file moved to cool gets charged as if it were 128 KB. The fix is either to leave small objects in hot, or bin-pack them into larger objects (a TAR or ZIP, for example) before tiering them down. The Microsoft Learn page on access tier best practices covers packing strategies in detail. Real-World Value, Use Cases, and ROI The customer in the session went from roughly 65,000 US dollars a month to around 25,000. That is not a marketing number, it is what happens when you apply the levers in order: Right-size redundancy. Move non-production and easily reproducible data off LRS in production. Reserve GRS for the workloads where a compliance regulation actually requires a second region. Match tiers to access patterns. Use premium for bursty, latency-sensitive workloads. Hot for active reads and writes. Cool and cold only when you genuinely access the data infrequently and have budgeted for the retrieval fees. Buy reserved capacity for the steady-state portion of your footprint. A one or three year commitment unlocks a discount on block blob capacity. See reserved capacity for Blob storage for terms and tier coverage. Kill wasteful transactions. Replace polling-for-changes with change feed. Replace recurring list-blob loops with a daily or weekly blob inventory report. Use conditional request headers (If-Modified-Since and friends) so reads skip unchanged objects. Pack small objects, or leave them in hot. Either is fine; tiering them down without packing is not. In short, the same scenario, with the same applications, runs at less than half the cost once you actually look at it. Getting Started Here is the order of operations I would follow tomorrow morning: Pull a blob inventory report on your largest storage accounts to see what is actually there: tier mix, object sizes, last modified dates, snapshots, versions. Open the Azure pricing calculator and model your scenario with realistic transaction counts and retrieval volumes. Do not just compare per-GB prices. Audit your redundancy choices against the workload. If an account is LRS in production with no easy way to rebuild the data, change it. Enable Smart Tier on your zone-redundant accounts. New objects start in hot, get demoted to cool after 30 days of inactivity, and to cold after 90, with no charges for tier transitions, early deletions, or data retrieval. Anything accessed gets instantly promoted back to hot. For accounts that cannot use Smart Tier, write a lifecycle management policy. Keep the rules simple at first: tier down after 30 days, archive after 180, expire snapshots and versions on a schedule. Convert one of your existing lifecycle policies to ARM or Bicep, then commit it to source control. Add an Azure Policy that flags any new storage account that does not match your standard. That last step is the one that sticks. As Benedict put it, cost optimization must become part of your system, not an afterthought. Resources Access tiers for blob data, Microsoft Learn Best practices for using blob access tiers, Microsoft Learn Azure Blob Storage lifecycle management overview, Microsoft Learn Optimize costs for Blob storage with reserved capacity, Microsoft Learn Enable Azure Storage blob inventory reports, Microsoft Learn Azure pricing calculator Keep Learning at the Summit Catch the full Microsoft Azure Infra Summit 2026 session playlist here Cheers! Pierre Roman172Views0likes0CommentsFeeding the GPUs: File Storage for AI and Cloud-Native Workloads on Azure
Hello Folks! If you are running AI workloads on Azure, you have probably learned the hard way that the wrong storage choice can leave a rack of very expensive GPUs sitting idle, waiting for data. In this session during the Microsoft Azure Infra Summit 2026, Wolfgang de Salvador and Reena Shah from the Azure Storage team walked through how Azure Managed Lustre and Azure Files map to the distinct stages of the AI pipeline, and why picking the right file system per stage is one of the highest leverage decisions you will make. 📺 Watch the session: Why IT Pros Should Care You probably did not get into IT to babysit checkpoint writes or debug Hugging Face egress bills at 2 a.m. But that is exactly the kind of work that lands on your plate when storage is not matched to the workload. Here is why MAIS28 matters for the IT pros, platform engineers, and Azure architects in the room: GPU time is the most expensive compute you will ever buy. Slow data loading and slow checkpoints turn that into burned cash. AI workloads are not one workload. Data prep, training, fine-tuning, and inferencing each have a different storage profile. Cloud-native AI on AKS and Azure Container Apps lives or dies on the ReadWriteMany experience. If model loading is slow or shared model caches do not exist, every cold start re-downloads hundreds of gigabytes. Storage choices ripple into security and compliance. Encryption in transit, redundancy, and snapshots are not optional in 2026. In short, this session is for anyone who has to answer the question, “What persistent volume should we use for this AI workload?” and wants a defensible answer. What Azure Brings to the Table, Technical Overview Wolfgang opened with the storage profile of every stage of an AI workflow. Data preparation needs hundreds of petabytes at the best TCO (think Azure Blob Storage as the durable core). Training and fine-tuning need extreme throughput so GPUs stay fed during data loading and so checkpoint writes complete fast. Inferencing needs fast model loads, low-latency KV cache, and grounded data for RAG. One filesystem does not fit all of those at once, and trying to make it fit is where teams overspend. Azure’s answer is a tiered, file-based portfolio that lines up with those stages: Azure Managed Lustre (AMLFS) is a fully managed, accelerator-tier filesystem. It scales to 25 PB of capacity and up to 512 GB/s of throughput, integrates with Azure Blob Storage as the durable core, and exposes a standard Lustre client plus a CSI driver for AKS. Azure Files is the natural ReadWriteMany choice for cloud-native AI on AKS and Azure Container Apps. It tops out at 256 TB of capacity and 10.4 GB/s of throughput, offers LRS and ZRS redundancy with snapshots and soft delete, and ships with a 99.99% SLA. Azure Blob Storage sits underneath both of these as the cheap, durable core for data prep and long-term retention. The mental model the speakers used is “accelerator and core”. Blob is the core, durable and economical. AMLFS is the accelerator for training. Azure Files is the accelerator for inferencing and shared state. Pick the right pair for the stage you are running. How It Works, Under the Hood For training, Wolfgang showed the demo most folks came to see: a 32 x H100 ND H100 v5 AKS cluster deployed from the Azure AI Infrastructure repository, running a 30B-parameter GPT-3 training job backed by AMLFS. Two things matter here. First, AMLFS absorbs checkpoint write bursts. When 32 H100s flush state at the same time, you need a filesystem that can take the punch without stalling. AMLFS does, which keeps GPU utilization drops short and contained. Second, the AMLFS Lustre CSI driver for AKS supports both static and dynamic provisioning with availability-zone placement, and there are five SKU tiers from MLFS20 (cheapest by capacity) to MLFS500 (cheapest by bandwidth). That means you can pick a cost-performance point that matches your training budget instead of buying the top SKU and hoping for the best. For inferencing, Reena’s half of the session was just as practical. Five reasons Azure Files fits AKS ReadWriteMany workloads: Standard Kubernetes RWM volume over NFS or SMB. 256 TB capacity ceiling and up to 10.4 GB/s of throughput per share. LRS or ZRS redundancy with snapshots and soft delete for protection. 99.99% SLA so it shows up in your availability math. Native support across AKS and Azure Container Apps, including serverless GPU. The headline feature is Azure Files Provisioned v2. In the old model, IOPS and throughput were a function of how much capacity you provisioned, which is wrong for AI shapes that need small capacity but very high IOPS and bandwidth. Provisioned v2 splits capacity, IOPS, and throughput into three independent knobs you can dial without downtime or remount. That alone changes the economics for a lot of inferencing patterns. The other big inferencing feature is NFS v4.1 encryption in transit, delivered with the az-nfs utility and stunnel. You get AES-GCM TLS protection on the wire, with no Kerberos and no Active Directory needed, and the application has no idea it is happening. Reena’s live demo showed an AKS pod with an encrypted NFS mount, transparent to the workload. And then the pattern that ties it together: the shared model cache. Download the model once into Azure Files, mount it across every replica via ReadWriteMany. No per-pod cold start, no re-download from Hugging Face, no egress bill. The demo used GPT-OSS 120B with VLLM on 32 x H100, and the pattern scales down to small fine-tuned models running on serverless GPU in Azure Container Apps. Real-World Value The session closed with the Viton case study. Viton is a Paris-based fashion AI startup. Their image-generation platform runs on Azure Container Apps serverless GPU, with Azure Service Bus for job routing and Azure Files NFS as the shared model store. Workers pull jobs, mount the shared model cache, generate the image, and scale to zero when the queue drains. The economics only work because they are not paying to re-download the model on every cold start, and because they only pay for GPU when there is work to do. The same pattern shows up across customer scenarios: Training a foundation model on AKS with AMLFS as the scratch tier and Blob as the durable archive. Fine-tuning smaller models where AMLFS checkpoints absorb the write bursts and the final artifact lands back in Blob. Inferencing with VLLM on AKS where Azure Files holds the model weights once and every replica reads from the same RWM mount. Serverless inferencing on Azure Container Apps with the same shared model cache pattern, but with scale-to-zero economics. Honest tradeoff: Lustre is not the right filesystem for a 10-pod web app, and Azure Files is not the right filesystem for a 32-GPU training run. The whole point of the tiering is that you pick the right one per stage. Do not try to make one filesystem do all four jobs. Getting Started If you want to put this into practice this week: Go to the Azure AI Infrastructure repository on GitHub. Wolfgang’s demo cluster came straight out of it, and you can spin up an AI-ready AKS cluster with GPU and InfiniBand operators in your own dev/test subscription. Install the Azure Managed Lustre CSI driver on AKS if you are running training or fine-tuning. Start with a smaller MLFS SKU and size up. Turn on Azure Files Provisioned v2 on a new share, then dial capacity, IOPS, and throughput independently to match your inferencing shape. Enable NFS v4.1 encryption in transit with az-nfs and stunnel before you put any sensitive workload on the wire. Try the shared model cache pattern. Pick one VLLM deployment, point it at an Azure Files RWM mount, and measure cold start time before and after. Resources Azure Managed Lustre documentation Use the Azure Managed Lustre CSI driver with Azure Kubernetes Service Azure Files documentation Understand Azure Files billing (Provisioned v2) Encryption in transit for NFS Azure file shares Azure Container Apps serverless GPUs Azure AI Infrastructure repository on GitHub Keep Learning at the Summit Catch the full Microsoft Azure Infra Summit 2026 session playlist here Cheers! Pierre Roman103Views0likes0CommentsPremium SSD v2 and Instant Access Snapshots: A Better, Faster, Cheaper Disk for Your Azure VMs
Hello Folks! If you have been running Premium SSD v1 because that is just what you have always done, this session from the Microsoft Azure Infra Summit 2026 is going to be a wake up call. Raymond Lui and Adam Li from the Azure Disk Storage team walked us through Premium SSD v2 (PV2 for short) and the new Instant Access Snapshots, and the punchline is simple. PV2 is faster, it is cheaper, and the operational story around it just keeps getting better. 📺 Watch the session: Why IT Pros Should Care If you are an infrastructure person, a SQL DBA, an SAP Basis admin, or anyone who has ever had to right-size a VM around its storage tier, this matters to you. In short, Premium SSD v2 changes the rules around how you provision block storage in Azure. Here is what stood out from the session: 4x more IOPS and 2x more throughput compared to Premium SSD v1, on a matched configuration that costs 42% less. Sub-millisecond average latency, with a top configuration of 800,000 IOPS and 20 GB/s of throughput on a single VM. Capacity, IOPS, and throughput are decoupled. You dial each one independently, in 1 GB increments, instead of buying a tiered SKU. 3,000 baseline IOPS and 125 MB/s throughput included on every disk, with no extra cost. Live Resize. You can grow disk size, IOPS, or throughput on a running VM with no restart required. Instant Access Snapshots make restores feel actually instant, with up to 10x faster hydration and 90% lower read latency during hydration. That is a lot of wins on one slide. Let’s break it down. What Premium SSD v2 Is, Technical Overview Premium SSD v2 is Azure’s purpose-built block storage for I/O-intensive enterprise workloads. Microsoft Learn describes it as designed for workloads that need sub-millisecond disk latency, high IOPS, and high throughput at a low cost. The target list is broad: SQL Server, Oracle, MariaDB, SAP, Cassandra, MongoDB, big data and analytics, gaming, and stateful containers running on AKS. The architectural shift that Raymond highlighted is independent scaling. With Premium SSD v1, you bought a fixed SKU. If you wanted more IOPS, you had to buy more capacity, even if you did not need it. With PV2, capacity, IOPS, and throughput are three separate dials. You provision capacity in 1 GB increments, then you set IOPS and throughput to match what your workload actually needs. If you over-provisioned, you tune it down. If you under-provisioned, you tune it up, and the VM keeps running. Raymond highlighted three primary use cases in the session: SAP workloads, including SAP application VMs, SAP HANA databases, and non-HANA databases like Oracle, DB2, and SQL Server in SAP environments. SQL Server. According to a GigaOM benchmark cited in the session, SQL Server on PV2 delivered 51% more transactions per second and 39% lower cost per transaction compared to AWS EC2, with a 9% lower 3-year TCO. Big data and analytics replacing local SSD. This one is a bit of a surprise. On D-series VMs, PV2 delivered over 1,400 MB/s of throughput compared to 720 MB/s from local SSD. That means you can run Spark or Databricks workloads on cheaper VM SKUs (without local storage) and still get more performance than you had before. Premium SSD v2 supports a 4k physical sector size by default, with 512E available for legacy applications. There are a few honest tradeoffs to know about. PV2 disks cannot be used as an OS disk, and they cannot be used with Azure Compute Gallery. PV2 also does not support host caching. For regions with availability zones, PV2 disks can only be attached to zonal VMs, so plan your VM placement accordingly. How It Works, Under the Hood Raymond covered the architecture briefly, and it is worth understanding. PV2 uses direct VM-to-storage-node communication, with 3-replica durability behind the scenes. That direct path is part of how it gets sub-millisecond latency consistently. For Instant Access Snapshots, Adam walked through the architectural difference between the classic incremental snapshot path and the new Instant Access path. With classic incremental snapshots for PV2 and Ultra Disk, the snapshot is created, then the data has to copy in the background to Standard HDD before the snapshot is usable for restore. That copy could take a while on a large disk, and restored disks would then hydrate slowly, which dragged down read latency until hydration finished. With Instant Access, the snapshot is usable the moment it exists. The data stays in the same high-performance storage as the source disk for a configurable duration (60 to 300 minutes, controlled by the InstantAccessDurationMins parameter). At the same time, Azure copies the snapshot data to Standard ZRS in the background for long-term retention. When the Instant Access window expires, the snapshot transitions to a regular incremental snapshot, sitting on cheap durable storage. You get the speed and the long-term durability without running two separate workflows. In short, your VM can boot and run at near-full performance while the data hydrates in the background. There are some limits to keep in mind. Instant Access counts toward the existing limit of three in-progress snapshots per disk, and you can create up to 15 disks concurrently from all instant access snapshots of a single disk. Real-World Value (Use Cases, ROI, Scenarios) Adam closed his portion of the session with a BCDR demo. He cloned 12 disks of an M-series production database into a recovery VM, attached them, and was immediately running roughly 500,000 IOPS at single-digit-millisecond latency. No waiting for hydration. No degraded performance window. That is a meaningful improvement to your Recovery Time Objective (RTO). A few scenarios where this combination really pays off: Pre-deployment safety nets. Take an instant access snapshot before a big upgrade. If something goes sideways, roll back in seconds instead of hours. Rapid scale-out for stateful apps. Spin up multiple disk copies of a primary instance in seconds. You can even place them across availability zones in the same region. Dev/test environment refresh. Clone production into dev or test on demand, with full performance from the first I/O. No more “we’ll refresh dev next quarter” because the restore takes too long. SAP HANA always-on operations. Live Resize means you can scale IOPS or throughput up on a running database during a load spike, without a maintenance window. Right-sizing to cut spend. If you have been paying for VM SKUs purely to get local SSD throughput, PV2 may let you drop to a smaller, cheaper VM and still hit higher numbers. One nuance came up in the live Q&A. Jens asked a great question about profiling: how do you know when PV2 is the right choice versus Standard SSD? Raymond’s guidance was direct. If the workload needs high IOPS or high throughput, PV2 is generally the right call. The VM SKU also needs to support “Premium Disk” capability for PV2 to attach, so check that compatibility first. Getting Started Concrete first steps so you can start kicking the tires: Confirm region and zone support. Use az vm list-skus --resource-type disks --query "[?name=='PremiumV2_LRS']" to see which regions and availability zones are supported in your subscription. Pick a Premium-capable VM in a supported zone. Remember, PV2 is zonal in AZ regions. Decide on the zone before you create the VM. Provision a disk. Start with default performance (3,000 IOPS, 125 MB/s) and a small capacity. You are paying for the dials you turn up; defaults are reasonable for most starting points. Plan your v1 to v2 migration. Raymond demoed two paths. Option A: detach the disk from a running VM and convert it (the VM keeps running on its other disks). Option B: stop and deallocate the VM, then convert in place. Both preserve data, and you can raise IOPS and throughput as part of the conversion. Try Instant Access Snapshots. Add --instant-access-duration-in-minutes (or the equivalent ARM/PowerShell parameter) to your existing snapshot command. That is all the change you need to enable it. For AKS users, define a storage class with skuName: PremiumV2_LRS and let dynamic provisioning take it from there. Resources Select a disk type for Azure IaaS VMs (managed disks) Deploy a Premium SSD v2 managed disk Convert managed disks storage between different disk types Instant access snapshots for Azure managed disks Use Premium SSD v2 with VMs in an availability set Use Azure Premium SSD v2 disks on Azure Kubernetes Service Azure managed disks overview SAP HANA Azure virtual machine Premium SSD v2 storage configurations Keep Learning at the Summit Catch the full Microsoft Azure Infra Summit 2026 session playlist here Cheers! Pierre Roman227Views0likes0CommentsNetwork Security Perimeter for Azure Event Hubs: Hardening Your Data Streams
What is Network Security Perimeter for Azure Event Hubs? Azure Event Hubs now supports Network Security Perimeter (NSP), a logical network isolation boundary that lets you define a security perimeter around your PaaS resources and control public network access through perimeter-based access rules. In practical terms, this means you can now group Event Hubs resources within a perimeter, apply consistent network access policies across them, and prevent unauthorized inbound traffic at the PaaS boundary level. It's not a firewall replacement, it's a compliance and segmentation tool that works alongside your existing NSGs and private endpoints. Before NSP, managing network access to Event Hubs involved: Private endpoints (which route traffic over private networks) IP firewall rules (which block public access from specific CIDR blocks) Virtual Network Service Endpoints (which restrict traffic to VNets) Network Security Perimeter adds a declarative, organization-wide layer: you define which resources belong inside the perimeter, and then manage access rules once, and those policies apply consistently across all perimeter members. Changes to the perimeter automatically cascade to all enrolled resources. Why ITPros Should Care If you're managing Event Hubs in a regulated industry like healthcare, finance, or government, you know the pressure. Compliance auditors want proof that data pipelines are segmented, isolated, and protected from lateral movement. Network Security Perimeter directly addresses that. Operational Value Network Security Perimeter delivers three immediate operational wins: Single Source of Truth for Access Rules. Instead of managing firewall rules on each Event Hubs namespace independently, you manage rules once at the perimeter level. Reduce configuration drift, reduce the attack surface, reduce human error. Compliance and Audit Readiness. Demonstrate network isolation to auditors with a clear diagram: "All Event Hubs in the perimeter are protected by these rules." That narrative matters for SOC 2, FedRAMP, HIPAA, and PCI-DSS compliance. You can export perimeter configurations and attach them to compliance documentation. Simplified Onboarding. When a new Event Hubs namespace joins the organization, add it to the perimeter and it inherits all access rules automatically. No manual rule-by-rule configuration. No weeks of back-and-forth with security teams. Secondary benefits include: Reduced blast radius during incidents, if an application is compromised, perimeter rules limit what it can access. Simplified network topology diagrams for architecture reviews. Faster mean time to remediation (MTTR) when security issues arise. Real-World Example: Securing a Multi-Tenant Event Hub Deployment Let's walk through a practical scenario. You're an ITPro at a financial services firm. You have three Event Hubs namespaces: hubs-prod-transactions (production trading data) hubs-prod-compliance (regulatory event streams) hubs-staging-dev (development and testing) Your security policy mandates: Production namespaces should only accept traffic from specific applications (IP-restricted). Staging can accept traffic from developer VNets but not from the internet. All outbound access to external services must be logged and monitored. Step 1: Define Your Perimeter First, create a Network Security Perimeter in the Azure Portal or via Azure CLI: az network perimeter create --resource-group rg-security --name nsp-financialservices --location eastus This creates the perimeter container. Think of it as a logical security zone. Step 2: Enroll Event Hubs Resources Add your Event Hubs namespaces to the perimeter: az network perimeter access-rule create --resource-group rg-security --perimeter-name nsp-financialservices --name allow-prod-apps --direction Inbound --access Allow --protocols Tcp --source-address-prefix 10.0.0.0/8 --destination-port-range 5671-5672 Enroll the Event Hubs namespace: az network perimeter resource create --resource-group rg-security --perimeter-name nsp-financialservices --resource-name hubs-prod-transactions --resource-type "Microsoft.EventHub/namespaces" You've now enrolled your production Event Hubs namespace. It inherits the "allow-prod-apps" rule, only traffic from your internal VNET (10.0.0.0/8) is permitted. Step 3: Define Access Rules $ns = "hubs-prod-transactions" $hub = "transactions-hub" $key = (az eventhubs namespace authorization-rule keys list --resource-group rg-prod --namespace-name $ns --name RootManageSharedAccessKey --query primaryConnectionString --output tsv) Create rules that reflect your security policy. Allow internal compliance applications: az network perimeter access-rule create --resource-group rg-security --perimeter-name nsp-financialservices --name allow-compliance-writers --direction Inbound --access Allow --protocols Tcp --source-address-prefix 10.50.0.0/16 --destination-port-range 5671-5672 Deny all other public traffic: az network perimeter access-rule create --resource-group rg-security --perimeter-name nsp-financialservices --name deny-internet --direction Inbound --access Deny --protocols "*" --source-address-prefix "*" --destination-port-range "*" Now your Event Hubs accept traffic only from specific internal subnets. Everything else is rejected at the PaaS boundary. Step 4: Validate Connectivity Test that legitimate applications can still reach Event Hubs: $ns = "hubs-prod-transactions" $hub = "transactions-hub" $key = (az eventhubs namespace authorization-rule keys list --resource-group rg-prod --namespace-name $ns --name RootManageSharedAccessKey --query primaryConnectionString --output tsv) Check logs in Azure Monitor: az monitor log-analytics query --workspace $(az monitor log-analytics workspace list --query "[0].id" -o tsv) --analytics-query "AzureDiagnostics | where ResourceProvider=='MICROSOFT.EVENTHUB' | summarize by NetworkSecurityPerimeter_s" If you see accepted connections logged with your perimeter name, you're good. If you see denied connections from unexpected IPs, you've caught a security issue before it impacts production. Step 5: Monitor and Alert Set up alerts for denied traffic: az monitor metrics alert create --name "NSP-Denied-Connections" --resource-group rg-security --scopes /subscriptions/{subId}/resourceGroups/rg-security/providers/Microsoft.Network/networkSecurityPerimeters/nsp-financialservices --condition "avg ConnectionRejectedCount > 5" --window-size 5m --evaluation-frequency 1m --action email-admin@company.com Now you'll be notified if someone attempts to access Event Hubs from an unauthorized source. Your security posture just went from reactive to proactive. Technical Details: How NSP Works Under the Hood Perimeter Architecture Network Security Perimeter operates at the Azure platform level, not in your VNets. Here's the flow: Connection arrives at Event Hubs public IP. Azure evaluates the source IP/protocol against NSP rules. If allowed, connection is routed to the namespace. If denied, connection is dropped and logged. This happens before TLS handshake, reducing CPU overhead and improving response times. Denied connections generate zero namespace load. Rule Evaluation Order NSP rules are evaluated in this order: Explicit Allow rules (matched first wins) Explicit Deny rules Implicit Deny (default action) Best practice: Create your Allow rules first (be specific about what you permit), then add Deny rules for anything not explicitly allowed. This ensures you don't accidentally block legitimate traffic. Integration with Existing Security Tools NSP works alongside (not instead of): Private Endpoints: NSP adds a policy layer; private endpoints route traffic over Azure backbone. Use both. IP Firewall: NSP provides namespace-level access control; IP firewall is still available for per-namespace rules. VNet Service Endpoints: NSP complements VNet endpoints by adding perimeter-wide policies. Managed Identity + RBAC: NSP is transport-layer security; identity-based access control remains separate. Performance Considerations NSP introduces minimal latency (<1ms typically). Azure evaluates rules in parallel and caches common decisions. For high-throughput Event Hubs: Keep rules simple and specific (avoid wildcard ranges if possible). Use CIDR blocks instead of individual IPs where applicable. Monitor connection acceptance rates in Azure Monitor. Comprehensive Resources Official Microsoft Documentation: Network Security Perimeter Overview Event Hubs Network Security Configuring NSP for Event Hubs Azure CLI: az network perimeter Azure RBAC for Event Hubs Azure Event Hubs Protocol Guide Closing: Perimeter Security for Modern Data Streams Network Security Perimeter for Event Hubs is a quiet but powerful addition to Azure's security toolkit. You get the ability to enforce organization-wide network policies without having to reconfigure every namespace individually. You can demonstrate perimeter-based isolation to auditors. You can catch lateral-movement attacks before they happen. For ITPros managing event-driven architectures, message processors, IoT data streams, financial transactions, this capability directly improves your security posture and reduces operational overhead. I encourage you to: Audit your current Event Hubs deployments. How many namespaces? How many security policies are you managing today? Design your perimeter boundaries. Group namespaces by security zone (prod, staging, dev) or by business unit. Start with one perimeter in a dev environment. Define rules. Validate connectivity. Then expand to staging and production. Document your perimeter architecture and rules. Include it in your security runbook and architecture reviews. Set up monitoring and alerting. Denied connections are a leading indicator of either misconfiguration or attack attempts. The networking challenges in cloud are complex. Network Security Perimeter gives you a declarative, policy-driven way to solve them at scale. Take advantage of it, and let me know how it changes your security workflows. Keep your networks hardened, and your data flowing safe. Cheers! Pierre Roman161Views1like0CommentsAz Update - Week 2 of the return editions
Hello Folks! This week's updates all focus on something we hear from IT pros and platform engineers all the time: How do we make our environments more secure, more manageable, and easier to modernize without adding more complexity? Whether you're running PostgreSQL workloads in Azure, securing Kubernetes storage, or planning your next wave of SQL Server migrations, this week's announcements bring practical improvements that can help reduce operational overhead while strengthening your overall platform strategy. We'll look at three newly available capabilities: Update #1 - Generally Available: Microsoft Defender security assessments for Azure Database for PostgreSQL Flexible Server Update #2 - Generally Available: Encryption in Transit for Azure Files NFS Shares in Azure Kubernetes Service (AKS) Update #3 - Generally Available: Expanding Azure Arc SQL Migration with SQL Server on Azure Virtual Machines As always, I'm approaching these updates from an infrastructure and operations perspective. I'll cover why each capability matters, what to watch out for before production deployment, and some practical steps you can take to start evaluating them in your own environment. Let's dig in. Update #1 - Generally Available: Microsoft Defender security assessments for Azure Database for PostgreSQL Flexible Server Why ITPros should care This release brings automated security posture assessment directly into managed PostgreSQL environments. For ITPros, this matters because database security is often treated separately from infrastructure security tooling, creating blind spots and silos. What changed is that Defender now runs native vulnerability scanning and compliance checks against PostgreSQL configurations, patches, and the ways a database could be exposed to security risks or attack opportunities. Instead of relying on external scanners or manual audits, you get platform-native assessments integrated with your existing Defender workflows. The operational impact is significant: you can now enforce security baselines at the database layer with the same consistency you apply to VMs and network resources, reducing the gap between infrastructure and data security accountability. Operational value Operationally, this improves your security baseline enforcement and reduces the need for separate database security assessment tools. It also strengthens how well you can demonstrate and prove that security controls are in place and working for compliance reviews where regulators expect consistent, documented security controls. Before production rollout, validate that Defender cost models fit your budget, that assessment frequency aligns with your change windows, and that remediation guidance maps to your patch and maintenance processes. Prerequisites include enabling Microsoft Defender for Cloud, registering the PostgreSQL Flexible Server provider, and ensuring network connectivity so assessments can reach the database endpoint. Real-world example with step-by-step guidance Enable Microsoft Defender for Cloud if not already active, and ensure PostgreSQL Flexible Server subscription coverage. Register the target PostgreSQL Flexible Server instances and confirm Defender has network visibility to the database endpoints. Run a baseline assessment and review initial findings to understand current security posture and common remediation patterns. Prioritise findings by severity and business impact, then schedule patches and configuration changes in maintenance windows. Monitor ongoing assessments and track remediation progress through Defender dashboards, validating that fixes reduce exposure scores. Technical details including code examples This example validates that Defender is actively assessing your PostgreSQL estate. The sequence checks Defender status, confirms PostgreSQL registration, and retrieves current assessment scores. Run these queries in a pilot subscription first to understand data structure and expected output before scaling to production databases. az account set --subscription <subscriptionId> az security sql-vulnerability-assessment baseline show --resource-group <rg> --server-name <postgresServer> --database-name <databaseName> az security pricing show --subscription <subscriptionId> --query "[?name=='VirtualMachines' || name=='SqlServers' || name=='StorageAccounts'].[name,pricingTier]" -o table az provider show --namespace Microsoft.DBforPostgreSQL --query "registrationState" -o tsv Expected behaviour: Defender status shows active, PostgreSQL instances are registered with the provider, and pricing tier reflects your coverage level. If assessments do not run, check network rules, managed identity permissions, and Defender plan activation. If baseline data is missing, trigger a manual scan and wait for completion. Comprehensive Resources Azure update: Microsoft Defender security assessments for Azure Database for PostgreSQL Flexible Server Microsoft Defender for Cloud overview Azure Database for PostgreSQL security SQL vulnerability assessments in Defender for Cloud Enable Defender for Cloud Update #2 - Generally Available: Encryption in Transit for Azure Files NFS Shares in Azure Kubernetes Service (AKS) Why ITPros should care This release closes a significant gap in data protection for Kubernetes workloads consuming NFS shares from Azure Files. Previously, NFS traffic between AKS nodes and Azure Files was unencrypted, creating compliance and security risks for sensitive workloads. What changed is that you can now enforce encryption for NFS communication at the Azure Files layer, not just at the application layer. This is important because traditional NFS lacks built-in encryption, and relying on network isolation alone is increasingly insufficient. For ITPros managing regulated workloads (healthcare, finance, PII-sensitive data), this removes a control gap. Encryption in transit now becomes a platform-native feature instead of a workaround, reducing architecture complexity and improving auditability. Operational value The operational value is stronger compliance posture and reduced attack surface for data in motion between containers and storage. It also simplifies the security story when auditors ask about data protection controls. Before enabling in production, validate that NFS-over-TLS introduces acceptable latency overhead for your workload patterns, test failover and reconnection behaviour under encryption, and confirm that monitoring and logging still work correctly. Prerequisites include running AKS with Azure CNI or Kubenet networking, having Azure Files with NFS 4.1 enabled, and ensuring the NFS client libraries on container images support TLS. Real-world example with step-by-step guidance Create an Azure Files NFS share with encryption in transit enabled and confirm TLS version alignment with your security standards. Deploy a test AKS workload that mounts the NFS share and validate that pods mount successfully with encrypted traffic. Run performance baselines (throughput, latency, CPU overhead) before and after enabling encryption to document operational expectations. Monitor pod logs and Azure Files metrics during the test to confirm no silent failures or unexpected throttling occurs. Roll out to production workloads in stages, with clear rollback criteria tied to application latency and error rates. Technical details including code examples This example validates that your AKS cluster can successfully mount NFS shares with encryption enabled. The sequence checks cluster networking, confirms NFS connectivity, and tests mount success. Run these commands in a non-production cluster first to validate environment readiness before touching production storage. az aks show --resource-group <rg> --name <clusterName> --query "networkProfile.{networkPlugin:networkPlugin,networkPolicy:networkPolicy,podCidr:podCidr}" -o jsonc az storage account show --resource-group <rg> --name <storageAccount> --query "{name:name,kind:kind,accessTier:accessTier}" -o jsonc kubectl get pvc -A --all-namespaces -o wide kubectl describe pv <pvName> | grep -i nfs Expected behaviour: cluster networking is properly configured, storage account kind supports NFS, and PVC/PV resources show NFS mount points. If mounts fail, check network security group rules, storage account firewall allowances, and subnet delegation. If latency increases, monitor resource utilisation and adjust workload placement if needed. Comprehensive Resources Azure update: Encryption in Transit for Azure Files NFS Shares in Azure Kubernetes Service (AKS) Azure Files NFS support Mount Azure Files with NFS in AKS Azure storage security AKS networking concepts Update #3 - Generally Available: Expanding Azure Arc SQL Migration with SQL Server on Azure Virtual Machines Why ITPros should care This capability brings SQL Server migration into the Azure Arc operational footprint, creating a unified migration and inventory experience. For ITPros, this matters because SQL Server modernisation is often fragmented across multiple tools and teams. What changed is that you can now discover, assess, and execute SQL migrations through Arc-native workflows, using the same permissions and governance model you already have for infrastructure and hybrid resources. The operational gain is consistency: discovery data feeds migration planning, assessments surface blockers early, and rollout can be controlled through the same change and approvals processes you use for other infrastructure migrations. Operational value Operationally, this reduces tooling sprawl and improves coordination between infrastructure and database teams. Arc becomes your single control plane for tracking migration progress, managing runbooks, and collecting audit evidence. Before production use, validate that your SQL Server inventory is complete, that migration blockers are understood and addressed, and that your maintenance windows can accommodate expected cutover timings. Prerequisites include Azure Arc agent deployment on source VMs, Azure Database Migration Service readiness, and network connectivity to target Azure SQL resources. Real-world example with step-by-step guidance Deploy Azure Arc agents to SQL Server VMs and confirm all instances report healthy status with complete inventory data. Run Arc-integrated SQL Server assessments to identify compatibility issues, dependencies, and recommended migration targets. Pilot migration for a non-critical workload to establish runbook patterns, measure cutover time, and validate post-migration validation procedures. Execute validation tests: connectivity, login success, database consistency checks, job execution, and application integration tests. Scale migration in waves using documented runbooks, with gates for monitoring data health and application performance after each cutover. Technical details including code examples This example validates Arc agent health and SQL Server discovery completeness. The sequence ensures your Arc infrastructure is ready for migration workflows. Run these commands as part of your pre-migration checklist to catch configuration gaps before committing to migration timelines. az account show --output table az connectedmachine list --resource-group <rg> --query "[].{name:name,status:status,osName:osName}" -o table az resource list --resource-type Microsoft.AzureArcData/sqlServerInstances --query "[].{name:name,resourceGroup:resourceGroup,location:location}" -o table az connectedmachine machine extension list --resource-group <rg> --machine-name <vmName> --query "[].{name:name,provisioningState:provisioningState}" -o table Expected behaviour: Arc agents report healthy status, SQL Server instances are fully discovered with accurate inventory, and required extensions are provisioned successfully. If discovery is incomplete, check Arc agent connectivity, extension deployment, and SQL service running status on source VMs. If migration pre-checks fail, verify SQL Server version compatibility and review Defender logs for blocking issues. Comprehensive Resources Azure update: Expanding Azure Arc SQL Migration with SQL Server on Azure Virtual Machines Azure Arc SQL Server Overview Azure Arc-enabled servers SQL Server on Azure Virtual Machines Azure Database Migration Service For any new capability this week, if they map to your operational roadmap, run a controlled pilot, measure the impact, and then scale with confidence. That is how you move the needle on modernisation while managing risk. Cheers! Pierre Roman101Views1like0Comments