azure kubernetes service
260 TopicsVirtual nodes on Azure Container Instances: a new compute layer for AKS
Meet virtual nodes on ACI Azure Kubernetes Service (AKS) gives you managed Kubernetes: the full Kubernetes API without operating the control plane yourself. Virtual nodes on Azure Container Instances go a step further, letting your pods run directly on Azure's serverless container platform, with the elasticity and with no capacity planning and no waiting for machines. Whether you already run AKS or want a managed Kubernetes that bursts without node management, this is for you. In short: virtual nodes on ACI attach Azure's serverless container platform to your cluster as Kubernetes nodes. Pods run as Hyper-V isolated containers, sized per pod rather than packed onto a fixed VM, up to 200 pods per virtual node. Run multiple virtual nodes, scaled as replicas, for more. They behave like any other pod: same kubectl, Helm, and GitOps. Kubernetes has always assumed a fixed set of machines underneath it. That assumption shapes everything above it: you size a node pool for a specific VM type in a specific region, you plan for peak rather than for average, and every workload on a node shares the same kernel and the same security boundary. Virtual nodes on ACI relaxs that assumption, which is what makes both elastic capacity and per container isolation possible without a different Kubernetes. If you've used the original AKS virtual nodes add-on (Virtual Kubelet based), this is not a rebrand. It is a new implementation that integrates far more deeply with Kubernetes, lifts most prior limitations (init containers, persistent volumes, managed identity, richer networking), and adds confidential containers as a first-class capability. The migration guide can be found here. Two capabilities carry the rest of this post: effortless burst capacity, and confidential containers. How virtual nodes on ACI work ACI runs every container as a Hyper-V isolated container, which means each one gets its own lightweight virtual machine boundary rather than sharing a kernel with its neighbors. Azure operates that platform. A virtual node connects it to your cluster. The cluster's control plane, the component that decides where each container runs, sees two kinds of destination: a small pool of virtual machines carrying cluster services, and one or more virtual nodes. From the application manifest's perspective, nothing changes. The pod lands on a virtual node; the virtual node hands it off to ACI. See Microsoft Learn: virtual nodes on ACI for the official capability and current limits. Virtual nodes on ACI in practice The rest of this post is hands on. You do not need to be a Kubernetes expert to follow it. kubectl is the command line tool for talking to a cluster, Helm installs packaged software into one, and a manifest is a text file describing what you want to run. If you have a cluster, everything below runs against it as written. The manifests behind the examples live in a companion demo repo. Setup is documented officially, and you can reproduce this end to end from the ACI virtual nodes documentation and the microsoft/virtualnodesOnAzureContainerInstances Helm repo. One requirement before you start: deploy into a delegated ACI subnet, meaning a subnet in your virtual network set aside for the ACI platform to place containers in. Size it for peak pod count plus headroom, since every pod consumes an address from it for its lifetime. Demo manifest files can be found in this repo, a personal sample repo provided as is and not a supported Microsoft artifact. Enable virtual nodes on ACI The virtual node is deployed via Helm. The Microsoft GitHub repo is itself a Helm repository, so a single helm install is all you strictly need. Cloning first, shown here, just makes it easier to customize values. Running kubectl get nodes afterward confirms the node registered. git clone https://github.com/microsoft/virtualnodesOnAzureContainerInstances.git helm install <yourReleaseName> ./virtualnodesOnAzureContainerInstances/Helm/virtualnode kubectl get nodes The virtual node appears alongside any existing capacity, ready to accept work. A virtual node is a Kubernetes node You target it the same way you would target any node. These few lines in a manifest say "run this on the virtual node": nodeSelector: virtualization: virtualnode2 kubernetes.io/os: linux tolerations: - key: virtual-kubelet.io/provider operator: Exists effect: NoSchedule That is the entire integration surface. No new API to learn, no separate deployment pipeline, no application changes. kubectl describe, kubectl logs, and kubectl exec, the standard commands for inspecting and troubleshooting, all work as they would anywhere else, including opening a shell inside a container running in a Hyper-V isolated boundary. Scaling stays trivial. kubectl scale deployment demo-deployment --replicas=10 lands every replica on the same virtual node, with no VMSS scale event, no provisioning latency, no climbing node-count chart. The same flow scales just as cleanly to hundreds. Cost follows the same shape. Each pod is billed per second against the cores and memory it requests, at ACI rates, and billing stops when the pod stops. Logs and metrics flow through the same path you already use, so existing dashboards and alerts keep working. One annotation makes a pod confidential Turning a regular container into a confidential one takes a single addition to its manifest: a policy that pins exactly which images, commands, environment variables, mounts, and capabilities are permitted inside the Trusted Execution Environment. The format is a base64 encoded Rego document, called a CCE (Container Confidential Enforcement) policy. You do not write that policy by hand. A tool generates it from the manifest you already have: az extension add -n confcom az confcom acipolicygen --virtual-node-yaml ./hello-world-deployment.yaml The tool pulls each image, hashes its layers, builds the allow-list, and injects the annotation back into the manifest. kubectl apply, and you're done. (acipolicygen has prerequisites of its own, including a working Docker installation; see the confcom documentation.) Here is why this is a genuinely new isolation primitive rather than a stronger version of an existing one. Most container security policy is enforced by software in the cluster, which means an attacker who compromises the host can potentially bypass it. This policy is enforced by the guest operating system inside the TEE instead. The underlying hardware, AMD SEV-SNP, also produces an attestation report, retrievable from inside the container, which is a cryptographic proof that the workload running is the workload you specified and nothing tampered with it. That is the guarantee regulated industries have been asking for, and increasingly the one AI workloads running untrusted code need too. The same per pod boundary is also what makes multi-tenancy on a single cluster realistic, though multi-tenancy in production still depends on your network and identity boundaries, which sit outside what the isolation layer itself provides. Background: Microsoft Learn: confidential containers on ACI. Wrapping up Virtual nodes on ACI give containers on Azure two things that were previously hard to deliver cleanly on Kubernetes: Effortless burst capacity on Azure's serverless container platform, billed per second for the cores and memory used, with no capacity planning and no waiting for machines. Confidential containers with hardware attested, per container isolation inside a Trusted Execution Environment. Virtual nodes are additive, not a replacement. Traditional node pools remain the right home for steady state, DaemonSet, and persistent volume workloads, and AKS features such as Node Auto Provisioning and Virtual Machine Node Pools already make that baseline more flexible. Virtual nodes on ACI absorb the spikes, the short-lived jobs, and the specialized isolation work on top. Where to start New to containers on Azure? Start with a small AKS cluster and add a virtual node from day one. You get a managed Kubernetes environment without having to guess your peak capacity in advance, and the elastic layer is there the first time you need it. Already running AKS? Add a virtual node to an existing cluster and move one bursty or short lived workload to it. Nothing else changes, and the comparison is immediate. Evaluating platforms? The capability that is hard to find elsewhere is the confidential containers path: hardware attested isolation per container, reachable through a standard Kubernetes manifest. The result: virtual nodes on ACI expand what AKS can run, with more capacity and stronger isolation, without changing the Kubernetes operating model you already use. Same kubectl, same manifests, same GitOps. New ceiling. For the high-level overview, official documentation, and Helm details, the Microsoft Learn is the source of truth. The companion repo holds the demo manifests used in this post. Acknowledgements I'd like to thank Gurpreet Virdi, Partner Group Engineering Manager, whose guidance shaped this post from the first outline through to publication. Her product leadership ensured this post reflects both the technical depth and the customer value of virtual nodes on ACI. Thanks to Gabriel Fuhrman, Senior Software Engineer, for his detailed technical review. His feedback refined the technical content and significantly improved the accuracy and depth of this post. Christopher Little, Principal CSA, shaped the enterprise adoption perspective, and Adam Sharif, CSA, reviewed the post from the earliest draft. Thanks also to Kirthi Maguluri, Senior Product Manager, and Varun Shandilya, Principal Product Manager, for their review of the blog.592Views1like0CommentsAction required: PromQL regex matching in Azure Monitor Workspace is becoming spec-compliant
TL;DR: Bare regex matchers like pod=~"foo" will soon return only exact matches, not substring matches. If you rely on =~ for "starts with" or "contains" semantics, update your queries to add .* before the change lands. What is changing AMW's Prometheus Query Service is moving to fully-anchored regex matching, aligning with the upstream Prometheus specification. After the change, every =~ and !~ matcher is wrapped with ^…$ before evaluation. This matches the behavior of self-hosted Prometheus, Grafana Cloud Metrics, and Amazon Managed Prometheus. "Regex matches are fully anchored. A match of env=~"foo" is treated as env=~"^foo$"." Reference: Prometheus querying basics — Instant vector selectors Behavior delta Given these timeseries: kube_pod_container_status_ready{pod="grafana-otel-collector-97b6c55f4-abc12"} kube_pod_container_status_ready{pod="grafana-otel-collector-97b6c55f4-def34"} kube_pod_container_status_ready{pod="grafana-otel-collector-97b6c55f4"} Query: kube_pod_container_status_ready{pod=~"grafana-otel-collector-97b6c55f4"} Environment Effective regex Series returned AMW PQS — current grafana-otel-collector-97b6c55f4 (unanchored) 3 (all matching pods) AMW PQS — after change ^grafana-otel-collector-97b6c55f4$ 1 (exact name only) Self-hosted Prometheus, Grafana Cloud, AMP ^grafana-otel-collector-97b6c55f4$ 1 (exact name only) If your query depends on the first row, it will return fewer — often zero — series after the change. What to do Audit any query, alert rule, recording rule, or dashboard variable that uses =~ or !~. Map each to the table below. Intent Update to Exact match No change necessary; both label="value" and label=~"value" should return exact matches after updating query engine version for fully anchored regex support. Starts with label=~"value.*" Ends with label=~".*value" Contains label=~".*value.*" Set of prefixes label=~"(a\|b\|c).*" Recommendation: Where the intent is exact match, switch from =~ to = It is clearer to readers and skips per-series regex evaluation. When ready, update AMW configuration to use the latest Query engine version. This can be done via the Azure Portal UI, Azure CLI, or PowerShell. Configuring AMW to use the latest Prometheus query engine version Azure Portal In the Azure portal, open your Azure Monitor workspace. Under Settings, select Properties. Select the Latest segment of the Query engine version for the AMW. Azure CLI Run the following command to update to the latest (2.0) metrics query engine version for the AMW: Az monitor account metrics-container update --subscription <subscriptionId> -g <resourcegroupName> -w <azmonWorkspaceName> -n default --version 2.0 Azure PowerShell Run the following command to use the latest Query engine version for the AMW: $resourceId = “/subscriptions/<subscriptionId>/resourceGroups/<resourcegroupName>/providers/Microsoft.Monitor/accounts/<azureMonitorWorkspaceName>/metricsContainers/default” Az resource update --ids $resourceId --api-version 2025-10-03 --set properties.version=2.0 --force-string To confirm: Az resource show --ids $resourceId --api-version 2025-10-03 --query properties.version --output tsv Expected output: 2.0 Common pitfalls Alternation anchors the whole expression. pod=~"foo|bar" becomes ^(?:foo|bar)$, not ^foo|bar$. Add .* and parentheses for prefix sets. . matches any character, including -. pod=~"my-pod.*" will match my-pod-xyz, my-podxyz, and my-pod9xyz. Use character classes like [-] for tighter matches. Empty-label matching is unchanged. label=~"" still matches series without that label. Timeline Public Preview rollout: September 2026 In Public Preview, newly created AMWs will default to the latest metrics query version (“2.0”) which features fully anchored regex matching. AMWs created prior to September 2026 will continue using version (“1.0”) unless users take action to update. This provides users with the opportunity to verify their query intent and update accordingly prior to cutover to the latest AMW metrics query engine version. General Availability: March 2027 In General Availability, all AMWs regardless of creation date will be updated to version “2.0” for cross-workspace compatibility. Action window: Audit and update queries before the General Availability date.243Views0likes0CommentsAzure Copilot announces general availability of the Troubleshooting Agent
Cloud operations are most effective when teams can turn insights into understanding and understanding into action. Whether you’re deploying new services, troubleshooting an issue, assessing resource health, or managing change, you need operational intelligence that helps you make faster, more informed decisions and continuously improve your environment. This is the promise of agentic cloud operations, a new operating model that uses AI powered agents equipped with deep contextual understanding and built in governance to help you run your cloud environments with speed and confidence. Azure Copilot brings AI-powered operations directly into Azure, helping you understand your environment, identify issues, determine likely causes and take informed action using the context already available in your Azure resources and services. Today, Azure Copilot is announcing the general availability of Troubleshooting Agent, a unified, built-in Azure Copilot capability that helps customers investigate and resolve operational issues faster. Available through both Azure Copilot and Support + Troubleshooting in the Azure portal, Troubleshooting Agent enables you to move from issue detection to resolution faster by bringing together troubleshooting insights, operational context, and recommended actions in a single experience. Deep expertise across Azure services Azure services support a wide variety of architectures, operational patterns, and workload requirements. While Azure Copilot provides consistent agentic troubleshooting across Azure, deeper service integrations enable specialized expertise and investigative capabilities for specific services. At general availability, Azure Compute and Azure Kubernetes Service (AKS) are the first services to offer these enhanced troubleshooting experiences. Azure Compute: You can investigate issues such as unexpected virtual machine restarts, connectivity and boot failures, high resource utilization, deployment and allocation failures, and unhealthy scale set instances. AKS: You can investigate issues such as pod restarts, stalled deployments, scheduling constraints, cluster networking, scaling behavior, upgrade regressions, and degraded performance. “This [Troubleshooting Agent] is really helpful for the person who is debugging, and they can come to a conclusion. Instead of going through the entire documentation, it is giving us the exact information on what it needs and what to look on.” – Senior Software Engineer, Lavelle Networks As the agentic troubleshooting capabilities of Azure Copilot continue to evolve, additional service specific skills and capabilities will be added to support a broader range of operational scenarios across Azure services through the same conversational experience. Grounded in Azure context, with customers in control Azure Copilot Troubleshooting Agent uses the resource information and supported diagnostics available for the affected resource while respecting your identity and Azure role based access control (RBAC). The agent explains the information behind its findings and recommends actions for review, allowing you to remain in control of remediation actions. Diagnostic depth varies by service, resource type, and scenario, but the goal remains consistent: helping you investigate, diagnose, and resolve operational issues more effectively using the context already available within your Azure environment. When additional assistance is needed, Azure Support remains available, but you can begin those engagements with a richer understanding of the issue, its likely causes, and potential remediation options. Available at no additional cost Troubleshooting Agent in Azure Copilot is available at no additional cost. There is no separate license or per-query charge, and existing Azure Support plans remain unchanged. Standard charges for the Azure resources continue to apply. Get started Open Azure Copilot and select Troubleshooting Agent in the Azure portal or go to Support + Troubleshooting, then describe the issue you are experiencing. Include the affected resource or subscription if it is not already clear from your current context. To get started, here are some prompt suggestions: "Why did my virtual machine restart unexpectedly?", "Investigate why my AKS deployment is failing.", "Why can't I connect to this resource?", "What recent platform issues have affected this resource?", "Check whether this resource has active errors or configuration issues." Troubleshooting Agent will use the supported resource diagnostics available for the investigation, explain what it finds, and recommend next steps. Looking ahead Connected cloud environments need a more intelligent approach to operations. As part of our broader vision for agentic cloud operations, Azure Copilot is your starting point when you need to understand, troubleshoot, and take action across your Azure environment. Troubleshooting Agent is an important step in that journey. Based on available context and diagnostics, we help you move from symptoms to informed next steps faster.4.5KViews2likes1CommentHow Azure uses AI to turn feedback into improved customer experience
Authors: eshaanbhattad, jenniferjhan, lakshminarasimha, bharadwajr The Challenge: Synthesizing fragmented feedback signals, to improve Azure's experience quality at scale Customers experience products and services end to end, but product experiences are often structured around individual service. One team may only know its top issues, while another may see only its own slice of experience. That structure makes it difficult to identify cross-cutting friction across the broader product experience. The feedback signals themselves are also fragmented. Customers share feedback through in-product surveys, support cases, field conversations, and social channels. Most product teams can see only part of that picture, making it hard to distinguish isolated comments from meaningful trends, understand which issues were having the greatest impact, and avoid missing critical feedback. For the Azure team, the challenges of fragmentation were amplified by scale. The team processed roughly 10,000 to 15,000 customer feedback reports each month, and synthesizing that feedback required 80 to 100 hours of expert analysis. Additional effort was needed to translate findings into consistent engineering work items. As feedback volume grew, manual analysis became increasingly unsustainable creating delays in identifying and addressing customer priorities. Compounding the challenge was the absence of an effective feedback loop to measure the impact of quality improvements. Teams struggled to justify investments in quality over new features because the return on those investments was difficult to quantify. The absence of a closed-loop measurement system made it difficult to consistently assess the customer impact of quality improvements. The team needed a system that could operate across organizational boundaries and across the product development lifecycle: Identify the most critical customer issues across fragmented feedback channels and product areas. Convert those insights into actionable engineering work, help teams address issues effectively, and measure outcomes to close the loop. The solution needed to preserve team-specific context, maintain auditability, continuously improve through feedback, and keep human experts in control of decisions that require judgment. To address these challenges, Microsoft launched the Great Experiences Matter (GEM) initiative. GEM is designed to analyze feedback signals in aggregate, with access controls and privacy safeguards designed to limit exposure of customer-identifiable information while helping teams identify patterns across channels. The Solution: an agentic feedback-to-fix loop with humans in control GEM created an AI-enabled feedback-to-fix workflow that connects customer listening, engineering action, and impact measurement across the product development lifecycle. The workflow uses Microsoft Foundry, Azure Data Explorer, Microsoft Fabric, Azure DevOps, and a set of custom agents to transform large volumes of qualitative feedback into prioritized insights and actionable engineering work. Fig 1. GEM AI-enabled automation workflow The system closes the loop through two connected motions. Find. Agentic workflows remove noise and duplicate reports, assess relevance and actionability, classify feedback against known issues from UX research, cluster related issues, and surface likely root causes. GEM builds on years of deep end-to-end UX research that has identified systemic friction across customer journeys and product boundaries. By continuously triangulating GEM signals with ongoing research, we combine broad, scalable listening with deep human insight to inform a more cohesive Azure experience. The results are surfaced through global scorecards for cross-cutting Azure issues and vertical scorecards tailored to individual product teams. Fig 2. GEM Global scorecard of top issues with Azure, data has been fictionalized to protect intellectual property Fix. The workflow creates Azure DevOps work items with the customer context, likely reproduction steps, recommended next actions, and an auditable trace of the supporting analysis. To date, 42% of the identified issues have been addressed through engineering action. The team is also extending an AI-assisted engineering workflow, using GitHub Copilot cloud agent, that can generate proposed fixes for straightforward issues. Engineers remain responsible for reviewing, refining, approving, and shipping any changes. The architecture is designed for inspection rather than blind automation. The recommendations include an auditable evidence trail allowing reviewers to inspect the source feedback, classifications, supporting references, confidence indicators, and recommendation actions. Human expertise enters the system at several points. Researchers shape the issue taxonomies and qualitative grounding. Product teams define ownership boundaries, business priorities, domain-specific vocabulary, and trusted sources that guide agent analysis. Engineers review and act on resulting work items, while leaders use scorecards to inform investment decisions. This context is captured in configuration files that evolve alongside the products they support. Teams can add new issue categories, refine keywords, clarify ownership boundaries, or identify trusted research sources. The next analysis cycle automatically incorporates the updated context without requiring changes to the underlying agents. That design creates a "feedback loop for the feedback loop". Teams review the root-cause analyses and work-item quality, identify gaps, and refine their configurations. This enables teams to continuously embed domain expertise into the workflow, improving how feedback is interpreted and prioritized without requiring changes to the underlying infrastructure. Their input improves subsequent runs, helping the system become more precise while preserving local product knowledge. The agent also maintains access to reports from previous runs and uses tools such as Web IQ and MCP servers to assess whether previously identified issues are improving, still require attention, or can be confidently closed. Fig 3. Example of the vertical feedback agent reasoning through customer feedback to find new work items One demonstrated example surfaced customer reports that a networking tool lacked IPv6 validation and support. The workflow generated an engineering work item describing the issue, customer impact, likely reproduction path, and recommended actions. The networking team reproduced the issue, validated the finding, and added it to its backlog. The goal is not to remove people from the process. Agents assume much of the cognitive load associated with sorting, clustering, tracing, and drafting, allowing experts to focus on judgment, prioritization, and implementation. Teams remain accountable for what is fixed, what is funded, and what is allowed to ship. The Impact: measurable experience gains at Microsoft production scale GEM began with manual interventions and is now scaling through AI-enabled workflows. The combined approach has produced measurable results across the Azure Portal and individual product experiences: The workflow aggregates and analyzes approximately 10,000 to 15,000 feedback reports each month across in-product, support, and social channels. Automated analysis reduced manual synthesis time by over 95 percent, turning a process that required 110 to 160 hours each month into a workflow that runs in under 60 minutes. Between October 2025 and April 2026, Azure Portal feedback rates declined by 30%. During that same period, GEM helped teams identify and prioritize experience improvements, creating a clearer link between customer feedback, engineering action, and outcome measurement. Service-level outcomes also demonstrate how better signals can drive business impact. For example, improvements to VM Connect experiences reduced overall Core Compute support volume by 1-2% per month, resulting in proportionate cost savings. The Azure Growth team increased subscription conversion by 9.1 percent after prioritizing issues highlighted through GEM. The value goes beyond speed. Leaders gain a more consistent basis for prioritization, and engineering teams receive work that is already connected to customer evidence and impact signals. Most importantly, every completed cycle creates new learning. Teams can measure changes in customer feedback and support volumes following improvements, incorporate partner input into future analyses, and continuously refine both the system and the products it helps improve. Key learnings and transferable practices The GEM experience offers several lessons for teams building agentic systems around complex, qualitative business processes: Start with real problems, not AI - Value comes from understanding the genuine business needs and applying AI where it is demonstrably better than existing approaches. Applying AI without a clearly defined problem often adds complexity without delivering meaningful value. Design for the end-to-end workflow - Value comes from connecting insights to the broader business process, including grounding in prior knowledge, prioritization, engineering action, post-fix measurement and reporting. Standalone AI output creates limited value, while an integrated workflow drives outcomes. Design for human judgment and accountability - Agents can reduce toil and cognitive load, but researchers, product managers, engineers, and leaders remain responsible for validating insights and determining appropriate actions. Ground agents in the knowledge of the teams they serve - Shared models require local context. Editable configuration files allow teams to define ownership, business priorities, releases, examples, and trusted sources without modifying underlying agents. Build observability and feedback mechanisms into the agentic system itself - Making analysis inspectable through reasoning traces, source context, and recommendations enables experts to identify gaps, improve outputs, and build trust over time. Build feedback mechanisms directly into the flow of work, making it effortless for users to provide input on the system. Tailor outputs to the people making decisions - Executives need trends and investment signals. Researchers need evidence and themes. Engineers need reproducible, actionable work. Effective systems deliver the right information to the right audience. Start small and iterate quickly - The AI landscape continues to evolve rapidly. Begin with a well-defined problem, measure outcomes, learn from feedback, and iterate as capabilities mature. Looking forward GEM continues to scale across the Azure Portal ecosystem. In addition to the global scorecard, vertical scorecards are now live with seven teams, expanding to the top 20 portal extensions representing more than 80% of portal traffic and feedback, with longer-term plans to extend coverage across the entire ecosystem. The roadmap includes expanded feedback ingestion, streamlined work-item tracking, AI-assisted remediation workflows, stronger evaluation, and a self-improving architecture. Proposed fixes would remain subject to engineer review, approval, and standard release controls before deployment. GEM is also developing AI-assisted pre-release governance workflows for production code that help identify potential quality issues during development. We will share more about these pre-release workflows in a future post. New tools and models will continue to evolve, but the enduring principle remains the same: combine enterprise-scale automation with clear ownership, trusted grounding, and human control. For Microsoft, Customer Zero means deploying these systems in real production environments, learning from the complexities, and sharing those lessons broadly. GEM shows what becomes possible when AI does more than summarize feedback. It helps an organization listen, act, measure outcomes, and continuously learn at customer scale. Microsoft's Customer Zero blog series gives an insider view of how Microsoft builds and operates Microsoft using our trusted, enterprise-grade agentic platform. Learn best practices from our engineering teams through real-world lessons, architectural patterns, and operational strategies for building, operating, and scaling AI-powered systems across the organization.916Views2likes0CommentsWiring Azure DevOps Pipeline Templates Without the Parameter Sprawl: The Manifest Facade Pattern
Where reusable pipelines start to hurt Who this is for: Platform and DevOps engineers who build shared Azure DevOps Pipeline Templates and want other teams to adopt them consistently. If you have ever built a set of shared Azure DevOps Pipeline Templates, you probably know how this goes. You write a whole library of clean, well-factored templates: provisioning infrastructure, deploying apps, scanning, promotion, and more. Each one is tidy on its own. Then the first team tries to actually use them together, and things get messy fast. That is exactly what happened to us on a customer engagement, while building a DevOps framework for their engineering teams. The templates themselves were fine. The trouble was wiring them together. Every consumer pipeline had to know how to call each template it used: the right order, which values one template fed another, and the full parameter list for every one of them. As the library grew, those consumer pipelines grew with it, long and repetitive. That friction turned into a real adoption problem. Getting started meant wading through pages of parameters, so teams put it off, copied whatever the team next door had, or quietly built their own thing instead. This post follows that story, from the tangle of parameters to the solution we landed on. Here is how it comes together. Where we started: Good templates, hard to adopt The DevOps framework we built had many templates. To keep this walkthrough concrete, we will follow just two of them: one that provisions infrastructure with Terragrunt (a wrapper around Terraform), and one that deploys an app to Kubernetes with Helm. Both are trimmed down to the essentials here so they are easy to follow. The infrastructure template, templates/infra/terragrunt-deploy.yml : parameters: - name: stack type: string - name: workingDir type: string steps: - script: | cd ${{ parameters.workingDir }} terragrunt plan -out=tfplan terragrunt apply -auto-approve tfplan displayName: "Deploy ${{ parameters.stack }}" The application template, templates/app/helm-deploy.yml : parameters: - name: releaseName type: string - name: chart type: string - name: namespace type: string - name: imageTag type: string steps: - script: | helm upgrade --install ${{ parameters.releaseName }} ${{ parameters.chart }} \ --namespace ${{ parameters.namespace }} \ --set image.tag=${{ parameters.imageTag }} displayName: "Deploy ${{ parameters.releaseName }}" There is nothing wrong with either file. The pain showed up in the pipeline that had to use them, where a team had to stitch the templates together by hand, remember the order, and repeat that boilerplate for every environment. Do that across a dozen services and you get long, copy-pasted pipelines that teams struggle to adopt. And because the wiring is done by hand, it leaves loopholes: a team can skip a step or override a parameter and slip past the guardrails you thought were in place. Challenge 1: Wiring the templates together The obvious move was to hide the wiring behind a single template that owned the order and the plumbing. We called it the orchestrator template: a consumer pipeline called that one template, and it wired up the rest. # templates/deployment-orchestrator.yml (the single orchestrator) parameters: - name: infraStack type: string - name: infraWorkingDir type: string - name: releaseName type: string - name: chart type: string - name: namespace type: string - name: imageTag type: string # ...and this list just kept growing stages: - stage: Infra jobs: - job: infra steps: - template: infra/terragrunt-deploy.yml parameters: stack: ${{ parameters.infraStack }} workingDir: ${{ parameters.infraWorkingDir }} - stage: App dependsOn: Infra jobs: - job: app steps: - template: app/helm-deploy.yml parameters: releaseName: ${{ parameters.releaseName }} chart: ${{ parameters.chart }} namespace: ${{ parameters.namespace }} imageTag: ${{ parameters.imageTag }} A team used it by calling that one template and passing a value for every parameter it exposed: # azure-pipelines.yml (a team's pipeline) parameters: - name: imageTag type: string extends: template: templates/deployment-orchestrator.yml parameters: infraStack: network infraWorkingDir: infra/network releaseName: orders-api chart: charts/orders-api namespace: orders imageTag: ${{ parameters.imageTag }} # ...and a value for every other parameter, too This solved the ordering and the copy-paste, but it handed us a new headache. This one template now had to expose every parameter of every template underneath it, and the list grew each time we added a capability. Worse, a real service usually needed more than one infrastructure stack and more than one Helm release. A flat list of parameters cannot express “two infrastructure stacks and three Helm releases” without silly names like infraStack1 , infraStack2 , and so on. Nobody could learn the thing. We had solved the wiring, only to trade it for a parameter problem. Challenge 2: The orchestrator’s parameter list explodes So how do you shrink that list? Step back and ask what you are really describing: a set of things to deploy. So the input should describe that set, not a long flat list of loose values. That is where the deployment manifest came in: one small config that lists the infrastructure to create and the apps to deploy. It reads about how you would expect: infrastructure: - stack: network workingDir: infra/network applications: - releaseName: orders-api chart: charts/orders-api namespace: orders Now the orchestrator can take that single manifest and loop over it, instead of exposing dozens of separate parameters. Its parameter list collapses to one, and a ${{ each }} loop turns each entry into a stage or job: # templates/deployment-orchestrator.yml (the orchestrator, now manifest-driven) parameters: - name: manifest type: object - name: imageTag type: string stages: - stage: Infra jobs: - ${{ each stack in parameters.manifest.infrastructure }}: - job: infra_${{ stack.stack }} steps: - template: infra/terragrunt-deploy.yml parameters: stack: ${{ stack.stack }} workingDir: ${{ stack.workingDir }} # ...an App stage loops over parameters.manifest.applications the same way And a consumer passes that manifest straight to the orchestrator: # azure-pipelines.yml (a team's pipeline) parameters: - name: imageTag type: string extends: template: templates/deployment-orchestrator.yml parameters: imageTag: ${{ parameters.imageTag }} manifest: infrastructure: - stack: network workingDir: infra/network applications: - releaseName: orders-api chart: charts/orders-api namespace: orders The parameter problem is solved. But hand-writing that whole manifest inside every pipeline is a lot to repeat, so the natural instinct is to pull it out into its own file. That is where things get tricky. Challenge 3: The manifest is read too late to shape the pipeline So we tried exactly that: we moved the manifest into a deployment-manifest.yml file and had the orchestrator read it back and expand it into stages and jobs. It sounded reasonable, but it did not work. To see why, you need to know how Azure DevOps builds a pipeline before it runs anything. An Azure DevOps pipeline actually happens in two phases, and they are further apart than people expect. First comes the build phase, before anything runs. Azure DevOps reads your YAML, pulls in every template, evaluates every ${{ }} expression, and unrolls every ${{ each }} loop. Out of this it produces one final, fully assembled pipeline. At this point no agent has started and no repo has been checked out. Only parameters and template expressions exist yet. Then comes the run phase. The assembled pipeline runs. An agent starts, checks out your code, and only now can a script open a file on disk. Here is the problem. Your deployment-manifest.yml does not exist as far as the pipeline is concerned until an agent checks out the repo, and that checkout happens during the run phase. So if the manifest is supposed to decide how many stages there are, or how many apps each get their own job, that decision has to be made earlier, while the pipeline is still being built. Important: The shape of the pipeline, its stages, jobs, and loops, is locked in while the pipeline is being built. A file you read during the run comes too late to change any of it. That was the wall we hit. The manifest was the right idea, but reading it from a file happened too late for the orchestrator to turn it into stages and jobs. So the manifest could not come from a file read during the run phase. It had to already exist, as an object, before the pipeline was assembled. The Manifest Facade pattern: Build the config while the pipeline is assembled The solution is to stop thinking of the manifest as a file to read, and start thinking of it as an object you build while the pipeline is being assembled. Parameters are available then. Template expressions can build a whole object then. So we added a small template, kept in the team’s own repo, with one job: take a couple of simple inputs, build the full manifest from them, and pass it to the shared orchestrator. We called it the builder, since assembling the manifest is its only job. It stays thin and lives next to the team’s pipeline, while the orchestrator stays in the shared platform repo. This is not a brand-new invention so much as a few familiar ideas working together, applied at pipeline build time. The manifest is a Parameter Object (Martin Fowler’s refactoring for collapsing a long parameter list into one structured value). The orchestrator is a Facade (the Gang of Four pattern for putting one simple, unified interface over a subsystem, here the underlying leaf templates). And the small template in the team’s repo is the piece that assembles that Parameter Object from a couple of inputs. What makes it an Azure DevOps pattern is the timing: the manifest is built as an object while the pipeline is assembled, so it can drive template expansion instead of sitting in a file that only gets read during the run. That is the twist we did not see written down anywhere, so we gave it a name: the Manifest Facade pattern. In practice it comes together in three small pieces: the consumer pipeline references the builder, the builder assembles the manifest and hands it to the shared orchestrator, and the orchestrator expands that manifest into stages and jobs. Here is how the files extend into one another: Here is each piece in turn. Step 1: The consumer references the builder The consumer pipeline extends the builder that lives in its own repo. It also declares the shared platform repo as a resource, so the orchestrator the builder calls is available while the pipeline is built. # azure-pipelines.yml (the team's pipeline) parameters: - name: imageTag displayName: Image tag type: string resources: repositories: - repository: platform type: git name: platform/pipeline-templates ref: refs/tags/v1.0.0 trigger: - main extends: template: config/deployment.yml@self parameters: environment: dev service: orders imageTag: ${{ parameters.imageTag }} Because the image tag is a runtime parameter, it is declared here in the consumer pipeline, which is what makes it appear in the Run pipeline panel for someone to fill in. Beyond that, the team passes only an environment and a service name, and the builder works out the rest. Step 2: The builder assembles the manifest The builder takes those inputs and assembles the whole manifest from them, then hands it to the orchestrator. Every structural value here is a parameter or a template expression, so all of it is ready while the pipeline is being assembled. # config/deployment.yml (builder, in the team's repo) parameters: - name: environment type: string - name: service type: string - name: imageTag type: string extends: template: deployment-orchestrator.yml@platform parameters: imageTag: ${{ parameters.imageTag }} manifest: schemaVersion: v1 infrastructure: - stack: network workingDir: infra/${{ parameters.environment }}/network applications: - releaseName: ${{ parameters.service }} chart: charts/${{ parameters.service }} namespace: ${{ parameters.service }} The builder assembles the manifest inline from the environment and service, then passes it straight to the orchestrator. The image tag is different: it is a runtime parameter the user supplies at queue time, so it is passed through to the orchestrator template and never becomes part of the manifest. The team gets a tiny interface, and the orchestrator still gets the full structure it needs. Step 3: The orchestrator turns the config into stages and jobs The orchestrator takes a single object and loops over it with ${{ each }} . Each stage and job gets generated while the pipeline is assembled. # deployment-orchestrator.yml (shared platform repo) parameters: - name: manifest type: object - name: imageTag type: string stages: # a ValidateManifest stage runs first (covered in the next section) - stage: Infra jobs: - ${{ each stack in parameters.manifest.infrastructure }}: - job: infra_${{ stack.stack }} steps: - template: infra/terragrunt-deploy.yml parameters: stack: ${{ stack.stack }} workingDir: ${{ stack.workingDir }} - stage: App dependsOn: Infra jobs: - ${{ each app in parameters.manifest.applications }}: - job: app_${{ app.releaseName }} steps: - template: app/helm-deploy.yml parameters: releaseName: ${{ app.releaseName }} chart: ${{ app.chart }} namespace: ${{ app.namespace }} imageTag: ${{ parameters.imageTag }} Because the manifest is a real object by the time the loops run, ${{ each }} unrolls it into actual jobs. List two stacks and you get two jobs. List two apps and you get two deploy jobs. And because the team only describes what to deploy, the sensitive wiring like service connections stays inside the platform templates, out of reach. The loophole is gone because the parameter is gone. Why the order of things matters The whole pattern comes down to who does what, and when. The builder turns a simple intent (“deploy orders to dev”) into a full manifest, while the pipeline is being assembled. The orchestrator turns that manifest into real stages and jobs, still while the pipeline is being assembled. Only the leaf templates do the real deployment work during the run: checkout, terragrunt apply , helm upgrade . Nothing about the shape of the pipeline waits for the run, so nothing about it depends on a file that only shows up after checkout. That is the whole trick. If you do have values you genuinely cannot know until the run, like a secret fetched from a vault or an artifact version an earlier stage writes to a variable, those still belong in runtime variables and variable groups. The manifest is for the structure and config you already know when you queue the build, which is almost always the part that was causing the pain. Validating the manifest against a schema Once the manifest became the one interface every team fills in, it needed a contract. We wrote that contract as a JSON Schema and kept it in the shared platform repo under schema/v1/ . The folder name is the version: backward-compatible additions go straight into v1 , and the day we need a breaking change we add a schema/v2/ alongside it, so existing consumers keep working while the schema evolves. Each manifest carries a schemaVersion so the orchestrator knows which contract to hold it to. On top of that, we validate in three layers, each catching a different class of mistake. Layer 1 is build time, for free. Because the orchestrator’s parameters are typed, with manifest declared as an object , Azure DevOps catches a range of structural problems while it expands the templates, before anything runs. For example, if a manifest left out infrastructure or misspelled it, the orchestrator’s ${{ each stack in parameters.manifest.infrastructure }} loop would have nothing valid to iterate over, and the failure would surface during template expansion rather than halfway through a deployment. Layer 2 is a version check that runs at build time. Both the ${{ if }} and parameters.manifest.schemaVersion are resolved while the pipeline is being assembled, so the orchestrator decides right then whether it understands the manifest’s version. If it does not, the only thing it generates is a single failing stage, and the real deployment stages are never built, so nothing runs against a contract the orchestrator does not know: # deployment-orchestrator.yml (shared platform repo) parameters: - name: manifest type: object stages: # Layer 2: reject schema versions this orchestrator does not understand - ${{ if not(containsValue(split('v1', ','), parameters.manifest.schemaVersion)) }}: - stage: UnsupportedSchemaVersion jobs: - job: fail steps: - script: | echo "##vso[task.logissue type=error]Unsupported manifest schemaVersion '${{ parameters.manifest.schemaVersion }}'" exit 1 # the ValidateManifest, Infra, and App stages below are generated only when the version is supported Layer 3 is a run-time check against the full schema. The first stage the orchestrator generates is ValidateManifest . Because the schema lives in the platform repo, the job checks that repo out first. Then it serializes the manifest object to JSON with the convertToJson expression, writes that JSON out to a file, and runs a JSON Schema checker to compare the file against the schema. If the manifest breaks the contract, the pipeline stops here, before any infrastructure or app stage runs: # Layer 3: validate the real manifest object against the JSON Schema - stage: ValidateManifest jobs: - job: validate steps: # the schema lives in the platform repo, so check it out first - checkout: platform - script: | echo '${{ convertToJson(parameters.manifest) }}' > manifest.json pip install check-jsonschema check-jsonschema --schemafile schema/${{ parameters.manifest.schemaVersion }}/deployment.schema.json manifest.json displayName: "Validate manifest against schema" Layer 1 is automatic, Layer 2 rejects an unknown contract at build time so the deployment stages are never generated, and Layer 3 confirms the actual values match the schema before any real work begins. Together they turn “the manifest looked right” into “the manifest is provably valid.” A few trade-offs to keep in mind No pattern comes without trade-offs. A few things worth weighing: The manifest has to be something template expressions can build. You can compose objects, loop, and branch with ${{ if }} , but there is no running arbitrary code while the pipeline is assembled. Anything fancier may need a prep step or a file generated upstream. Every team carries a small builder template. We think that is a fair trade, since it keeps their intent local and readable, but it is one more file per repo. Keep it thin and let the orchestrator hold the real logic. Treat the schema as living documentation. Because the manifest’s JSON Schema spells out every field and what it means, teams can read it to build their own manifest with confidence, instead of reverse-engineering the orchestrator. Pin the shared repo to a tag, like the v1.0.0 above, so a change to the orchestrator does not silently change everyone’s pipeline on the next run. Debugging takes a small shift in habit. When something looks off, use the pipeline’s preview to see the fully assembled YAML before it runs. It shows you exactly what the loops produced. Tip: Use the Azure DevOps pipeline preview to see the fully assembled YAML without running anything. It is the fastest way to confirm your manifest unrolled into the stages and jobs you expected. Wrapping up It is a journey a lot of platform teams will recognize. Clean templates, messy wiring, one orchestrator that fixes the order but drowns in parameters, a config file that brings back sanity, and then the surprise that reading it at the wrong moment means it can never shape the pipeline. The solution is small and it sticks. Put a thin builder template next to each team, let it assemble the manifest from a couple of simple inputs, and let a shared orchestrator turn that manifest into stages and jobs. Teams get an interface they can easily understand and adopt. The platform team keeps the wiring and the guardrails in one place. And the timing gap that trips up so many “just read the config file” attempts stops being a problem, because you are working with it instead of against it. Key takeaways Shared Pipeline Templates stall on adoption when every team has to wire them together by hand. A single orchestrator template fixes the ordering, but a flat parameter list does not scale to real adoption. A config file read during the pipeline run cannot shape the pipeline, because stages and jobs are decided earlier, while the pipeline is assembled. The Manifest Facade pattern builds the config as an object at assembly time, so one small input drives many templates, consistently and with the guardrails baked in. Give the manifest a versioned JSON Schema and validate it in layers, so a broken contract fails fast instead of halfway through a deployment. How are you handling template sprawl in your own pipelines? We would love to hear what has worked for your teams in the comments.632Views1like0CommentsHow Microsoft 365 built a platform engineering layer on AKS to ship faster at global scale
This Customer Zero story explains how Microsoft 365 built COSMIC, a platform engineering layer on top of Azure Kubernetes Service (AKS), to standardize how cloud services are deployed and operated at global scale. The goal was to eliminate repetitive infrastructure work for service teams, embed security and compliance by default, and enable developers to focus on delivering customer value instead of managing platform complexity.4.8KViews4likes0CommentsFind anomalies in Prometheus and OpenTelemetry metrics with Dynamic Thresholds (Preview)
Dynamic thresholds are extended to query-based metric alerts in Azure Monitor, allowing to detect and alert on anomalies in Azure Monitor managed Prometheus metrics and OpenTelemetry metrics stored in an Azure Monitor Workspace. This follows the introduction of Dynamic Thresholds for Log search alerts — Azure Monitor now offers consistent Dynamic Thresholds support across logs and metrics — platform metrics, log search queries, and now query-based metric alerts. A consistent anomaly-detection approach, wherever your signals live. Dynamic thresholds are not a single static formula. They apply a range of machine-learning models and algorithms to historical query results, learn each series’ normal rhythm — including hourly, daily, and weekly seasonality — and automatically fit the most appropriate baseline separately to every time series. This way, a single alert rule can monitor many resources or dimensions while each one gets its own independent, self-refining baseline. Why Dynamic Thresholds Matter Simpler configuration: Reduce the need to define, maintain, and continuously tune static thresholds inside PromQL alert logic. Adaptive monitoring: Let alert thresholds adjust to changing workload behavior, recurring traffic peaks, and seasonal usage patterns. At-scale intelligence: Monitor multiple time series with a single alert rule, while Azure Monitor learns an independent baseline for each resource or dimension combination. Example 1 — Spot CPU anomalies in AKS workloads Scenario: Monitor container CPU utilization across pods or deployments in AKS with a query-based metric alert built on Prometheus metrics. Example query: sum by (microsoft_resource_id, namespace, deployment, container) (rate(container_cpu_usage_seconds_total[5m])) / sum by (microsoft_resource_id, namespace, deployment, container) (container_spec_cpu_quota / container_spec_cpu_period) Why dynamic thresholds help: CPU usage of a Kubernetes workload changes with workload mix, deployment timing, scaling activity, and traffic patterns. Static thresholds can be difficult to tune across namespaces, deployments, and containers. Dynamic thresholds learn a separate baseline for each monitored time series — in this example, for every pod, deployment, and container combination — so genuine CPU spikes stand out while expected variation from autoscaling and traffic mix stays quiet. Example 2 — Catch application latency regressions sooner Scenario: Detect abnormal latency patterns in an application by alerting on custom OpenTelemetry metrics stored in an Azure Monitor Workspace. Example query: histogram_quantile(0.95, sum by (le, service_name, http_route, http_method) (rate(http_server_duration_seconds_bucket[5m]))) Why dynamic thresholds help: Application latency naturally changes with traffic, user behavior, and release cadence. Fixed thresholds can be noisy during peak periods and too loose during quiet ones. Dynamic thresholds learn a separate baseline for each time series — here, for every service, route, and method — so real p95 latency regressions surface even as traffic and release cadence shift throughout the day. Best practices for better results To get the best results from dynamic thresholds for PromQL-based alerts, design your query so Azure Monitor can learn a clear, stable signal over time: Keep the expression numeric. Dynamic thresholds work best when the query returns a continuous numeric signal rather than a Boolean true/false result. For example, use an expression that calculates CPU usage, not a Boolean comparison like CPU > 0.8. Use meaningful dimensions. Split by dimensions such as namespace, deployment, service, or route when you want separate baselines for different workloads or endpoints. Prefer stable entities. Use longer-lived dimensions or aggregate across short-lived entities so the model has enough consistent history to learn from. In Kubernetes, for example, deployment is usually a better baseline dimension than individual pod ID. Choose the right threshold behavior. Decide whether the alert should trigger on values above the learned upper bound, below the lower bound, or both. Start with medium sensitivity. Use Medium as a balanced default, then tune up or down based on noise and missed anomalies. Allow enough historical data. Dynamic thresholds improve as more history is collected. Initial seasonal patterns use recent history, and weekly seasonality becomes more effective after several weeks of data. Get started Ready to try it? Create a query-based metric alert with dynamic thresholds on your metrics in Azure Monitor Workspace. You can create such rules in the Azure portal, where the built-in preview chart shows when your dynamic threshold alert would have fired based on historical baseline analysis. Use the preview chart to tune both the PromQL query and the dynamic threshold sensitivity before enabling the rule. You can also create query-based metric alert rules using programmatic interfaces or resource templates. Figure 1. Dynamic thresholds preview chart showing the learned baseline and the points where an alert would have fired. Dynamic thresholds cut alert noise where it starts — at detection. The alerts that do fire connect into Azure Monitor’s broader AIOps experience, where the Azure Copilot Observability Agent can help correlate signals into investigated issues with explainable reasoning — with humans in control. Next steps Related blog: Anomaly detection made easy with Dynamic thresholds for Log search alerts Dynamic thresholds in Azure Monitor Query-based metric alerts overview Create query-based metric alerts Prometheus metrics in Azure Monitor OpenTelemetry on Azure Monitor Stay connected Follow the Azure Observability Blog for more updates on Azure Monitor, Prometheus-based monitoring, alerting, and troubleshooting experiences. We’ll continue sharing product updates, practical guidance, and examples to help you improve observability across your Azure environments. Feedback We’d love to hear how dynamic thresholds for query-based metric alerts work for your scenarios. Share your feedback through your Microsoft account team, Azure support channels, or the feedback options in the Azure portal so we can continue improving the experience.261Views0likes0CommentsIPv6 Dual-Stack Endpoints for Azure Container Registry (Public Preview)
By Johnson Shi, Aviral Takkar, Bin Du Introduction Two of the most common networking questions we hear from teams running Azure Container Registry (ACR) are: "Can my registry serve clients on IPv6 networks?" — Teams operating IPv6-only or dual-stack networks need their container registry reachable over IPv6. "How do we start moving registry traffic toward IPv6 without breaking anything?" — Organizations guarding against IPv4 address exhaustion, or operating under IPv6 transition mandates, want a migration path that doesn't disrupt existing IPv4 clients. Today, we're announcing the public preview of IPv6 dual-stack endpoints for Azure Container Registry for public endpoints and firewall rules, with IPv6 over private endpoints planned for GA. Set your registry's endpoint protocol to IPv4AndIPv6 , and its endpoints become reachable over both IPv4 and IPv6 — so IPv4-only, dual-stack, and IPv6-capable clients all connect to the same registry, each over whichever protocol their network stack selects. Key Takeaways ACR registries now support an endpointProtocol setting with two values: IPv4 (default) and IPv4AndIPv6 (dual stack, preview). Dual stack is additive — your registry continues serving IPv4 clients exactly as before. There is no IPv6-only mode. Dual stack requires dedicated data endpoints to be enabled ( --data-endpoint-enabled true ), and dedicated data endpoints require the Premium SKU. The service enforces this requirement. You can enable it today with Azure CLI 2.87.0 via az acr update --endpoint-protocol IPv4AndIPv6 . FQDN-based client firewall rules keep working unchanged; IP-based allowlists need to account for IPv6 traffic. Limitation: This public preview covers IPv6 for the registry's public endpoints and firewall rules only. IPv6 over private endpoints is planned for a future release. Limitation: ACR Tasks isn't supported on a registry that has IPv6 dual-stack enabled. Tasks does not work when the endpoint protocol isIPv6 dual-stack, including quick builds (with az acr build) and quick task runs (with az acr run). Support is planned for a future release. How to enable it On an existing registry (Azure CLI 2.87.0 or later) Dual stack requires dedicated data endpoints, so enable both in a single update: az acr update --name <your-registry> --data-endpoint-enabled true --endpoint-protocol IPv4AndIPv6 If dedicated data endpoints are already enabled, set the endpoint protocol on its own: az acr update --name <your-registry> --endpoint-protocol IPv4AndIPv6 Verify the configuration: az acr show --name <your-registry> --query "{endpointProtocol:endpointProtocol, dataEndpointEnabled:dataEndpointEnabled}" { "dataEndpointEnabled": true, "endpointProtocol": "IPv4AndIPv6" } Note: If your clients sit behind a firewall and you're enabling dedicated data endpoints for the first time, add firewall rules for <your-registry>.<region>.data.azurecr.io before enabling — switching from *.blob.core.windows.net to dedicated data endpoints changes where layer blobs are downloaded from. See Dedicated data endpoints for details. Reverting to IPv4 Dual stack is reversible at any time: az acr update --name <your-registry> --endpoint-protocol IPv4 Reverting the endpoint protocol leaves dedicated data endpoints enabled; disable them separately if desired. Scope of this preview This public preview enables IPv6 for the registry's public endpoints — the login server, dedicated data endpoints, and regional endpoints (if enabled). IPv6 over private endpoints isn't part of this preview. Support is planned for a future release. Until then, registries reached through a private endpoint continue to use IPv4. Additionally, IPv6 dual-stack support for ACR Tasks, including support for `az acr build` and `az acr run`, are not supported in the public preview. Support is planned for a future release. Requirements and how features compose Requirement Why Premium SKU Dedicated data endpoints are a Premium feature. Dedicated data endpoints enabled IPv4AndIPv6 requires dataEndpointEnabled: true ; the service rejects the setting otherwise. Azure CLI 2.87.0+ Adds --endpoint-protocol to az acr update . For geo-replicated registries, the endpoint protocol is a registry-level setting, and dedicated data endpoints exist in every replica region. Firewall guidance: rules based on registry FQDNs — the login server, dedicated data endpoints, and regional endpoints (if enabled) — continue to work unchanged for dual-stack registries; only IP-address-based allowlists need updating for IPv6. To learn more, see IPv6 dual-stack endpoints in Azure Container Registry (preview) and the ACR endpoint reference. If you have further questions about IPv6 dual-stack endpoints or dedicated data endpoints, reach out to us on the Azure Container Registry GitHub repository or file feedback through the Azure portal.315Views1like0CommentsAzure Copilot Observability Agent is generally available, with autonomous operations in preview
Complex cloud environments have outpaced manual operations. Agentic cloud operations connect people, tools, and data to streamline investigation workflows and move teams from scattered signals to evidence-backed next steps. With unified observability, teams can investigate Azure-monitored applications, Azure Kubernetes Service (AKS) environments, VMs, Foundry telemetry, infrastructure, and platform signals with greater context and control. Powered by Azure Monitor, the Azure Copilot Observability Agent is now generally available. It helps engineering, SRE, DevOps, and operations teams move from telemetry and alert noise to investigated issues, explainable reasoning, and recommended next steps that can reduce Time-To-Mitigate (TTM). Autonomous operations are also available in public preview. They help prepare context and reduce triage work while people remain responsible for mitigation decisions and any changes to the environment. From alert noise to investigated issues The Observability Agent helps teams reduce the effort required to understand operational problems. Instead of starting every investigation from a dashboard, query editor, or alert payload, teams can work with an AI companion that reasons across telemetry, Azure resource context, discovered topology, and custom instructions to identify what changed, what is correlated, and what evidence supports the conclusion. Teams can start with natural-language exploration and continue into deeper investigations when an issue requires more evidence. That light-to-deep workflow helps responders move from broad questions to a structured investigation without losing the reasoning trail. Here's what this looks like in practice: after a deployment, several alerts might fire across an app, database dependency, and compute resource. The Observability Agent can group those signals around the affected service, identify when the regression started, compare related dependencies and infrastructure metrics, and capture the findings in an Azure Monitor issue. The responder can then validate the evidence, add team context, route work to the right owner, and decide whether a rollback, configuration change, or code fix is appropriate. Explainable investigations across Azure-monitored signals Operations teams need more than a chatbot that answers questions. The Observability Agent follows an investigation workflow: it frames hypotheses, gathers evidence, compares signals by time, scope, and type, rules out weak explanations, and shows the reasoning path behind its findings. The Observability Agent can help teams: Investigate incidents and alerts across Azure-monitored applications, Azure Kubernetes Service (AKS) environments, VMs, Foundry telemetry, infrastructure, and platform signals Correlate related signals to reduce noise and surface higher-signal issues with context Explore telemetry using natural language while preserving transparency into the supporting data Compare signals by time, scope, and type to separate likely causes from coincidental changes Provide a reasoning trail that shows what the agent found, what it ruled out, and why Recommend next steps that engineers can review before deciding how to act This same investigation model applies to specialized skills and issue types, including customer's application, Azure Kubernetes Service (AKS), Foundry, VMs, and GenAI issues. When the relevant telemetry is available, the Observability Agent can correlate logs, metrics, traces, alerts, dependencies, resource graph, resource health, activity logs, Foundry telemetry, and changes. This helps teams investigate customer-visible issues with evidence, including latency, token spikes, tool-call failures, agent errors, hallucinations, deployments, API failures, performance regressions, infrastructure dependencies, and platform incidents. This explainability is central to the product. In production operations, trust is earned through evidence. The Observability agent is built to support human judgment, not bypass it. . Azure expertise, with context from your environment Context matters in every investigation. The same symptom can mean different things depending on application architecture, recent deployments, dependencies, historical incidents, and team practices. The Observability Agent brings Microsoft and Azure operational knowledge into the investigation experience. It can use discovered topology, Azure resource context, logs, metrics, traces, and custom instructions to ground investigations in signals that are more relevant to your environment. Native to Azure Monitor, with humans in control Because the Observability Agent is built into Azure Monitor, teams can use it close to the telemetry, alerts, and workflows they already rely on. Investigations can also be captured as Azure Monitor issues, creating a shared case file for humans and agents to collaborate on evidence, reasoning, and next steps. The Observability Agent is designed for governed AI operations inside Azure Monitor. Interactive chat and investigations use the signed-in user's identity and Azure role-based access control (RBAC). Prompts and responses are not used to train foundation models, and the agent doesn't restart resources, change configuration, or resolve issues on its own. Autonomous operations in public preview Alongside general availability, autonomous operations for the Observability Agent are available in public preview. When enabled, the agent can analyze alerts in the background, correlate related alerts when they likely represent the same incident, create Azure Monitor issues automatically, and run deep investigations on agent-created issues. This automatic triage helps reduce alert noise by turning streams of individual alerts into higher-signal issues with context, findings, and recommended next steps. Teams can review the issue, continue the investigation, and decide what action to take. Autonomous operations are designed to prepare context and reduce triage work, not to remove human control. Engineers remain responsible for decisions, approvals, and any changes to the environment. Next steps Check out our latest announcements and related blogs: Azure Blog and OMB Blog. Learn how to use the Observability Agent in Azure Copilot Observability Agent. Explore how investigations work in Deep investigations in the Azure Copilot Observability Agent. Learn more on how to Chat with your observability data Learn how teams preserve context in Azure Monitor issues. Review preview details in Autonomous operations in the Azure Copilot Observability Agent. Stay connected Follow this blog for ongoing deep dives, updates on current capabilities, and a preview of what's coming next. Live webinar - a walkthrough of real Observability Agent scenarios, best practices, and what's available today - along with a look at what's coming next, and live Q&A with the product team. Register for the Observability Agent webinar. We'd love your feedback The Observability agent continues to evolve based on real-world usage and operator feedback. Share your thoughts directly through the Give Feedback option in the experience, or reach us at enauerman@microsoft.com.10KViews6likes0Comments