azure networking
124 TopicsWhat’s new in Azure Firewall: Recent innovations
Organizations are modernizing their networks while managing a growing mix of applications, protocols, and security requirements. Recent Azure Firewall releases strengthen this journey with simpler traffic routing, expanded protocol support, more flexible application controls, improved health visibility, and higher intrusion-detection performance. This roundup highlights five recent capabilities now available in general availability or public preview. In this update: Explicit proxy, IPv6 support, HTTP header insertion, auto-learn SNAT routes, and IDPS performance improvements. Explicit proxy is now generally available Explicit proxy enables clients and applications to send HTTP and HTTPS traffic directly to Azure Firewall by configuring the firewall as their proxy. Instead of relying only on route-based traffic steering, organizations can use familiar proxy settings to centralize outbound web access through Azure Firewall. This release provides a simpler way to use Azure Firewall as a managed forward proxy. It can help teams consolidate web egress controls, support workloads that natively understand proxy configuration, and transition from traditional proxy appliances without introducing another infrastructure layer. Direct proxy configuration: Configure supported clients and applications to use Azure Firewall for HTTP and HTTPS traffic. Centralized web controls: Apply Azure Firewall policy and application rules to proxied traffic. Operational simplicity: Use a managed Azure service instead of deploying and maintaining separate proxy infrastructure. Migration flexibility: Support environments that already use proxy-aware applications or proxy auto-configuration workflows. For more details, visit - Azure Firewall explicit proxy | Microsoft Learn IPv6 support is now in public preview As address requirements grow and organizations adopt dual-stack architectures, IPv6 support is becoming an important part of cloud network design. Azure Firewall IPv6 support, now in public preview, extends centralized traffic filtering to IPv6 scenarios and helps customers protect applications and networks as they introduce IPv6 connectivity. With this preview, teams can evolve IPv4-only environments toward dual-stack networking while continuing to use Azure Firewall as a central enforcement point. This reduces the need to maintain separate security architectures for IPv4 and IPv6 traffic and provides a more consistent policy and operational model. Secure IPv6 traffic across hybrid environments: Filter east-west and hybrid IPv6 traffic across Azure and on-premises networks. Works seamlessly with IPv6-enabled Azure services: Integrate Azure Firewall into end-to-end IPv6 architectures alongside services like ExpressRoute and Virtual Network. Prepare for the future of networking: Build and secure dual-stack environments today while accelerating your IPv6 adoption journey. For more details, visit -Deploy Azure Firewall in dual stack mode (preview) | Microsoft Learn HTTP header insertion is now generally available HTTP header insertion enables Azure Firewall to add configured headers to HTTP requests that match application rules. This gives security and network teams an additional policy control for communicating trusted context to downstream services without requiring every client or application to add the header itself. The capability can support scenarios where applications use headers to enforce organization-specific access requirements, identify traffic handled by a trusted network path, or apply downstream controls. Because configuration is centralized in Azure Firewall policy, teams can apply the behavior consistently across matching traffic and reduce application-side changes. Tenant restriction enforcement: Organizations can inject tenant restriction headers into traffic destined for Microsoft Entra ID, helping prevent users from authenticating unauthorized tenants and strengthening identity governance controls. Secure egress for AVD and enterprise workloads: Support AVD, VDI, and enterprise egress scenarios where organizations need web traffic to carry approved tenant or context headers. Security and Compliance Enforcement: Administrators can add organization-specific headers to web traffic to support security policies, compliance requirements, and backend validation workflows. This helps ensure that only approved applications, tenants, or services are accessed through corporate environments. Operational Efficiency: Customers no longer need dedicated proxy devices solely for HTTP header injection. Azure Firewall can now perform header insertion natively as part of the application rule, reducing operational complexity and infrastructure costs. For more details, visit - Azure Firewall HTTP Header Insertion Configuration | Microsoft Learn Auto-learn SNAT routes is now generally available Source network address translation behavior depends on whether Azure Firewall treats a destination as private or public. In complex enterprise and hybrid networks, manually maintaining the private address ranges that should not be source-NATed can become time-consuming and error-prone as the environment changes. Auto-learn SNAT routes simplifies this process by dynamically learning relevant routes through Azure Route Server and using them to update the firewall’s private IP range configuration. This helps Azure Firewall preserve original source addresses for traffic destined to learned private networks while reducing ongoing configuration maintenance. Reduce manual SNAT management : Automatically learn private and registered routes through Azure Route Server, eliminating the need to manually maintain large No-SNAT prefix lists. Preserve source IPs across hybrid networks : Automatically apply learned routes as No-SNAT destinations, helping maintain source IP visibility and predictable routing for internal traffic. For more details, visit- Azure Firewall SNAT private IP address ranges | Microsoft Learn IDPS performance improvements are now generally available Azure Firewall Premium includes signature-based intrusion detection and prevention to identify and block malicious network activity. The latest IDPS performance improvements increase the amount of protected traffic that Azure Firewall Premium can process, helping organizations apply advanced inspection to demanding production workloads. Azure Firewall Premium now supports up to 22 Gbps with TLS inspection and IDPS in Deny mode, and up to 600 Mbps for a single TCP connection when IDPS is enabled in Alert or Deny mode. Actual performance depends on traffic characteristics, rule configuration, enabled features, and deployment conditions. Higher aggregate throughput: Protect larger traffic volumes while using advanced inspection capabilities. Improved single-flow performance: Better support applications that rely on high-throughput TCP connections. Strong prevention posture: Use IDPS Deny mode to actively block matching malicious traffic. Premium-scale security: Apply TLS inspection and IDPS to more bandwidth-intensive enterprise workloads For more details, visit - Azure Firewall performance | Microsoft Learn Building an advanced, more capable cloud firewall Together, these releases expand how Azure Firewall can protect modern networks. Explicit proxy and HTTP header insertion provide more flexible application-layer controls; IPv6 support helps customers evolve toward dual-stack architectures; Auto-learn SNAT routes reduces operational overhead in dynamic and hybrid environments; and IDPS performance improvements extend advanced threat prevention to higher-throughput workloads. Explore these capabilities in Azure Firewall and review the applicable Azure documentation for configuration requirements, supported scenarios, regional availability, and preview terms before enabling them in production environments.129Views1like0CommentsExpressRoute Gateway Microsoft initiated migration
Important: Microsoft initiated Gateway migrations are temporarily paused. You will be notified when migrations resume. Objective The backend migration process is an automated upgrade performed by Microsoft to ensure your ExpressRoute gateways use the Standard IP SKU. This migration enhances gateway reliability and availability while maintaining service continuity. You receive notifications about scheduled maintenance windows and have options to control the migration timeline. For guidance on upgrading Basic SKU public IP addresses for other networking services, see Upgrading Basic to Standard SKU. Important: As of September 30, 2025, Basic SKU public IPs are retired. For more information, see the official announcement. You can initiate the ExpressRoute gateway migration yourself at a time that best suits your business needs, before the Microsoft team performs the migration on your behalf. This gives you control over the migration timing. Please use the ExpressRoute Gateway Migration Tool to migrate your gateway Public IP to Standard SKU. This tool provides a guided workflow in the Azure portal and PowerShell, enabling a smooth migration with minimal service disruption. Backend migration overview The backend migration is scheduled during your preferred maintenance window. During this time, the Microsoft team performs the migration with minimal disruption. You don’t need to take any actions. The process includes the following steps: Deploy new gateway: Azure provisions a second virtual network gateway in the same GatewaySubnet alongside your existing gateway. Microsoft automatically assigns a new Standard SKU public IP address to this gateway. Transfer configuration: The process copies all existing configurations (connections, settings, routes) from the old gateway. Both gateways run in parallel during the transition to minimize downtime. You may experience brief connectivity interruptions may occur. Clean up resources: After migration completes successfully and passes validation, Azure removes the old gateway and its associated connections. The new gateway includes a tag CreatedBy: GatewayMigrationByService to indicate it was created through the automated backend migration Important: To ensure a smooth backend migration, avoid making non-critical changes to your gateway resources or connected circuits during the migration process. If modifications are absolutely required, you can choose (after the Migrate stage complete) to either commit or abort the migration and make your changes. Backend process details This section provides an overview of the Azure portal experience during backend migration for an existing ExpressRoute gateway. It explains what to expect at each stage and what you see in the Azure portal as the migration progresses. To reduce risk and ensure service continuity, the process performs validation checks before and after every phase. The backend migration follows four key stages: Validate: Checks that your gateway and connected resources meet all migration requirements for the Basic to Standard public IP migration. Prepare: Deploys the new gateway with Standard IP SKU alongside your existing gateway. Migrate: Cuts over traffic from the old gateway to the new gateway with a Standard public IP. Commit or abort: Finalizes the public IP SKU migration by removing the old gateway or reverts to the old gateway if needed. These stages mirror the Gateway migration tool process, ensuring consistency across both migration approaches. The Azure resource group RGA serves as a logical container that displays all associated resources as the process updates, creates, or removes them. Before the migration begins, RGA contains the following resources: This image uses an example ExpressRoute gateway named ERGW-A with two connections (Conn-A and LAconn) in the resource group RGA. Portal walkthrough Before the backend migration starts, a banner appears in the Overview blade of the ExpressRoute gateway. It notifies you that the gateway uses the deprecated Basic IP SKU and will undergo backend migration between March 7, 2026, and April 30, 2026: Validate stage Once you start the migration, the banner in your gateway’s Overview page updates to indicate that migration is currently in progress. In this initial stage, all resources are checked to ensure they are in a Passed state. If any prerequisites aren't met, validation fails and the Azure team doesn't proceed with the migration to avoid traffic disruptions. No resources are created or modified in this stage. After the validation phase completes successfully, a notification appears indicating that validation passed and the migration can proceed to the Prepare stage. Prepare stage In this stage, the backend process provisions a new virtual network gateway in the same region and SKU type as the existing gateway. Azure automatically assigns a new public IP address and re-establishes all connections. This preparation step typically takes up to 45 minutes. To indicate that the new gateway is created by migration, the backend mechanism appends _migrate to the original gateway name. During this phase, the existing gateway is locked to prevent configuration changes, but you retain the option to abort the migration, which deletes the newly created gateway and its connections. After the Prepare stage starts, a notification appears showing that new resources are being deployed to the resource group: Deployment status In the resource group RGA, under Settings → Deployments, you can view the status of all newly deployed resources as part of the backend migration process. In the resource group RGA under the Activity Log blade, you can see events related to the Prepare stage. These events are initiated by GatewayRP, which indicates they are part of the backend process: Deployment verification After the Prepare stage completes, you can verify the deployment details in the resource group RGA under Settings > Deployments. This section lists all components created as part of the backend migration workflow. The new gateway ERGW-A_migrate is deployed successfully along with its corresponding connections: Conn-A_migrate and LAconn_migrate. Gateway tag The newly created gateway ERGW-A_migrate includes the tag CreatedBy: GatewayMigrationByService, which indicates it was provisioned by the backend migration process. Migrate stage After the Prepare stage finishes, the backend process starts the Migrate stage. During this stage, the process switches traffic from the existing gateway ERGW-A to the new gateway ERGW-A_migrate. Gateway ERGW-A_migrate: Old gateway (ERGW-A) handles traffic: After the backend team initiates the traffic migration, the process switches traffic from the old gateway to the new gateway. This step can take up to 15 minutes and might cause brief connectivity interruptions. New gateway (ERGW-A_migrate) handles traffic: Commit stage After migration, the Azure team monitors connectivity for 15 days to ensure everything is functioning as expected. The banner automatically updates to indicate completion of migration: During this validation period, you can’t modify resources associated with both the old and new gateways. To resume normal CRUD operations without waiting 15 days, you have two options: Commit: Finalize the migration and unlock resources. Abort: Revert to the old gateway, which deletes the new gateway and its connections. To initiate Commit before the 15-day window ends, type yes and select Commit in the portal. When the commit is initiated from the backend, you will see “Committing migration. The operation may take some time to complete.” The old gateway and its connections are deleted. The event shows as initiated by GatewayRP in the activity logs. After old connections are deleted, the old gateway gets deleted. Finally, the resource group RGA contains only resources only related to the migrated gateway ERGW-A_migrate: The ExpressRoute Gateway migration from Basic to Standard Public IP SKU is now complete. Frequently asked questions How long will Microsoft team wait before committing to the new gateway? The Microsoft team waits around 15 days after migration to allow you time to validate connectivity and ensure all requirements are met. You can commit at any time during this 15-day period. What is the traffic impact during migration? Is there packet loss or routing disruption? Traffic is rerouted seamlessly during migration. Under normal conditions, no packet loss or routing disruption is expected. Brief connectivity interruptions (typically less than 1 minute) might occur during the traffic cutover phase. Can we make any changes to ExpressRoute Gateway deployment during the migration? Avoid making non-critical changes to the deployment (gateway resources, connected circuits, etc.). If modifications are absolutely required, you have the option (after the Migrate stage) to either commit or abort the migration.3.1KViews0likes2CommentsAzure Firewall explicit proxy is now generally available
We are excited to announce the general availability of explicit proxy in Azure Firewall. This capability brings the familiar explicit proxy configuration model natively to Azure Firewall, enabling applications and browsers to send outbound HTTP and HTTPS traffic directly to the firewall through standard proxy settings. Customers can now gain centralized policy enforcement and visibility while reducing their dependence on self-managed forward proxy appliances. Proxy the traffic you choose—not the entire subnet Traditional route-based steering can require all traffic from a subnet to traverse a firewall. Explicit proxy offers a more targeted model: configure selected applications or browsers to use the Azure Firewall private IP and proxy port, while other traffic continues on its existing route. This application-level control can simplify migrations, reduce unnecessary routing changes, and help teams apply inspection where it matters most. What’s new for general availability A simpler single-port experience: Serve both HTTP and HTTPS destinations through one HTTP explicit proxy endpoint port, reducing client configuration complexity. More secure PAC file access: Retrieve customer-owned proxy auto-configuration files from Azure Blob Storage using a managed identity. PAC files can be up to 256 KB. A streamlined portal workflow: Enable Explicit proxy when creating a Firewall Policy, with added guidance and validation to help prevent common configuration errors. Built for real-world modernization Modernize legacy proxy infrastructure. Use Azure Firewall as the forward proxy endpoint while retaining the explicit proxy model already configured in applications and browsers. This can help organizations consolidate infrastructure and reduce operational overhead associated with self-hosted proxy appliances. Apply selective application-level steering. Direct only the workloads that require centralized inspection through Azure Firewall, without forcing every flow on a subnet through the same path. Secure Azure Arc connectivity in hybrid environments. Organizations using ExpressRoute or VPN can configure Azure Firewall as a forward proxy for Azure Arc-enabled servers, providing an inspected, policy-controlled outbound path to required Azure services without opening direct internet access from corporate networks. How it works Enable explicit proxy in the Azure Firewall Policy and choose the HTTP proxy port used for both HTTP and HTTPS destinations. Configure applications manually with the firewall’s private IP and proxy port or enable proxy auto-configuration. If you use a PAC file, store it in Azure Blob Storage and configure the required managed identity and access permissions. Create an application rule in the Firewall Policy to allow the intended outbound destinations. Get started today Ready to simplify outbound web traffic steering? Explore the Azure Firewall explicit proxy documentation for prerequisites, portal and API configuration steps, PAC file guidance, and supported scenarios. With explicit proxy now generally available, Azure Firewall gives organizations another flexible path to modernize secure egress—combining a familiar proxy model with the simplicity, scale, and centralized management of a cloud-native service.251Views0likes0CommentsSimpler, private connectivity between Azure and AWS with Azure Multicloud Interconnect
Multicloud is no longer the exception — it is how most enterprises operate. Teams run analytics in one cloud and applications in another, place workloads to meet data-residency requirements, and increasingly move large volumes of data between clouds to train and serve AI models. Yet the network that connects these environments has remained one of the hardest parts of a multicloud strategy to get right. Connecting Azure and AWS privately has traditionally meant stitching together Azure ExpressRoute, AWS Direct Connect, a connectivity provider or colocation footprint, customer-managed routers, BGP sessions, and link-layer encryption — then owning the resiliency design and day-to-day operations across all of it. The result is often slow to deliver, difficult to troubleshoot, and inconsistent in performance. Today, Microsoft and AWS are together introducing Azure Multicloud Interconnect, a jointly engineered, fully managed service that delivers private, high-throughput connectivity between Azure and AWS through a single logical resource. Both clouds coordinate provisioning, resiliency, encryption, and lifecycle operations, so the connection between your environments simply works — end to end. The challenge with connecting clouds today For most organizations, cross-cloud connectivity has been a build-it-yourself exercise. Establishing a private link between Azure and AWS typically requires provisioning an ExpressRoute circuit on one side and a Direct Connect circuit on the other, engaging a connectivity provider or securing space in a colocation facility to bridge them, and then configuring and maintaining the routers, BGP peering, and encryption that tie the two clouds together. Each of those pieces is owned by a different team or vendor, which makes the end-to-end path only as reliable as its least-managed component. When something breaks, isolating the root cause means coordinating across Microsoft, AWS, a network provider, and your own operations team — and mean time to resolution suffers as a result. Capacity planning is equally difficult: bandwidth is often provisioned for peak demand and left underused, while scaling up to meet a new AI or data initiative can take weeks. The outcome is a connectivity layer that is slow to stand up, expensive to operate, and hard to reason about — exactly the opposite of what a multicloud strategy is meant to deliver. Azure Multicloud Interconnect removes that complexity by turning cross-cloud connectivity into a managed service. What is Azure Multicloud Interconnect? Azure Multicloud Interconnect is a provider-managed, private intercloud connectivity service built on the proven foundations of Azure ExpressRoute and AWS Direct Connect. Instead of assembling and operating the underlying components yourself, you establish one interconnect resource and consume a private, dedicated path between your Azure and AWS environments. Under the covers, Microsoft and AWS coordinate the circuits, routing, and encryption on your behalf. You get the outcome you want — a reliable private connection between clouds — without becoming the systems integrator for it. Because the service builds on ExpressRoute and Direct Connect, it fits naturally into the connectivity models, tooling, and operational practices your teams already use in each cloud. One managed resource replaces a stack of circuits, routers, BGP sessions, and encryption you would otherwise build and operate yourself. Figure 1. Azure Multicloud Interconnect provides a single, managed private path between Azure and AWS, built on ExpressRoute and Direct Connect with a quad-redundant, 400G-class design and MACsec encryption. How it works Azure Multicloud Interconnect is engineered for the performance and availability that production and AI-scale workloads demand: A single logical connection. You provision and manage one interconnect resource. There are no customer-owned routers to rack, configure, or patch between the clouds. Quad-redundant, multi-site design. The data path is built across redundant devices and diverse sites — four Azure Microsoft Enterprise Edge (MSEE) routers and four AWS routers — so there is no single point of failure. High bandwidth backbone. A high-capacity LAG-based design provides the headroom needed for large-scale data movement and distributed AI traffic. Elastic bandwidth. Scale capacity up or down as demand changes, rather than provisioning for peak and paying for it year-round. Encryption by default. MACsec link-layer encryption is enabled automatically, so traffic between clouds is protected without extra configuration. Enterprise-ready networking. The service is IPv6-ready, supports APIPA addressing, and targets a 99.99% availability SLA. Routing and provisioning are coordinated by the platform, which removes the most common sources of cross-cloud connectivity errors — mismatched BGP configuration, inconsistent encryption settings, and asymmetric or fragile failover paths. A foundation you can trust Azure Multicloud Interconnect is not a new, unproven path — it extends two connectivity services that enterprises already rely on: Azure ExpressRoute and AWS Direct Connect. Both are private connectivity backbones trusted for mission-critical hybrid and cloud workloads, and the interconnect inherits their carrier-grade capacity, global edge presence, and operational maturity from day one. It also means your teams do not have to learn a separate model. The interconnect appears as a first-class resource governed by the identity, access, and monitoring controls each cloud already provides, and it can be automated with the tooling you use today. Extending trusted services, rather than introducing a parallel one, is what allows Microsoft and AWS to offer intercloud connectivity as a managed experience with confidence. Why this matters for your team Azure Multicloud Interconnect is designed to change the economics and the experience of running across clouds. Taken together, these benefits shift cross-cloud networking from a specialized, high-effort project to a repeatable, on-demand capability. Instead of dedicating senior engineers to build and babysit intercloud links, your team can direct that expertise toward the applications and data platforms that differentiate your business — while trusting that the connective tissue between clouds is reliable, secure, and ready to scale. Faster time to value. Turn up private Azure–AWS connectivity through a guided, managed workflow instead of a multi-week integration project spanning several vendors. Lower operational burden. Microsoft manages the infrastructure, resiliency, and lifecycle, so your networking team is freed from patching routers and diagnosing cross-cloud faults. Predictable, high performance. A dedicated, private path with high capacity that delivers the consistent throughput and low latency that public-internet or VPN paths cannot guarantee. Security you don’t have to assemble. Private connectivity plus default MACsec encryption keeps intercloud traffic off the public internet and protected in transit. A consistent experience in both clouds. Provision, monitor, and manage the interconnect using the native constructs and tooling your teams already rely on in Azure and AWS alike. Built for the way enterprises use multicloud “I've heard a very consistent message across many years of customer engagements - we are multi-cloud enterprise by design. With this announcement we are taking a burden on connecting clouds away from the customers." Igor Sakhnov, CVP Azure Networking Customers have told us where dedicated, managed intercloud connectivity makes the biggest difference. Azure Multicloud Interconnect is designed for scenarios such as: Distributed AI workloads. Move training data and model outputs between clouds at high throughput to feed pipelines wherever the compute lives. Large-scale data movement. Replicate datasets, back up across clouds, and support analytics that span Azure and AWS. Cross-cloud disaster recovery. Use a second cloud as a resilient recovery target over a private, reliable link. Hybrid and best-of-breed architectures. Run each application on the cloud that suits it best while keeping the connection between them private and performant. Regulated and sovereign workloads. Keep intercloud traffic on a private path to help meet data-residency and compliance requirements. Workload migration. Rehost or rebalance workloads between clouds without re-engineering connectivity for each move. Availability and what’s next Azure Multicloud Interconnect is launching first for Azure and AWS connectivity, the pairing customers ask about most often. Microsoft and AWS are starting with a preview so that you can validate the experience against your own architectures, with general availability to follow. This is the beginning of a broader journey. Building on existing multicloud connectivity capabilities, we plan to extend the managed interconnect model to additional clouds — including Google Cloud (coming soon)— so that a consistent, provider-managed experience can span your entire multicloud estate. As always, the roadmap will be guided by customer demand and real-world use. Our shared goal is simple: make the network the easiest part of your multicloud strategy, not the hardest. By delivering intercloud connectivity as a managed service, Microsoft and AWS want every organization — from a team moving its first dataset between clouds to an enterprise operating at AI scale — to connect Azure and AWS privately, securely, and with confidence, and to do it in minutes rather than months. Get started Learn more about Azure Multicloud Interconnect, Azure ExpressRoute, and AWS Direct Connect from Microsoft and AWS documentation, talk with your Microsoft or AWS account team about joining the preview, and tell us which cloud pairings and scenarios matter most for your organization. We are building this with your feedback. Azure Multicloud Interconnect is jointly delivered by Microsoft and AWS, built on Azure ExpressRoute and AWS Direct Connect. Feature availability, performance targets, and timelines may evolve as the service moves from preview to general availability.3.4KViews2likes0CommentsLessons Learned #551: Azure SQL Connection Timeouts: Three Things to Check
An application starts reporting intermittent timeouts when connecting to Azure SQL Database. Some requests succeed, others fail, and a test from a developer’s laptop works perfectly. The database appears online, no recent deployment seems related, and the natural reaction is to ask: Is Azure SQL unavailable? Is the firewall blocking the connection? Should we increase the connection timeout? Should we change the driver or scale the database? Those are reasonable questions, but they may lead the investigation in the wrong direction. The most important lesson is simple: A timeout tells us how long the application waited. It does not tell us what the application was waiting for. Not every “SQL timeout” happens inside Azure SQL From the application’s point of view, opening a database connection may involve several operations: Resolving the server name. Reaching the SQL endpoint. Obtaining a Microsoft Entra access token. Waiting for an available pooled connection. Completing the SQL login. Executing the first command. When all these operations are reported through the same application method or log entry, it can look as though Azure SQL took thirty seconds to accept the connection. In reality, only part of that time may have been spent connecting to the database. In one anonymized support scenario, the application experienced problems mainly on its first connection. Network tests were successful and no corresponding SQL connection failure was identified. The investigation eventually showed that access-token acquisition was consuming a significant part of the available time. Increasing the SQL timeout or changing the firewall would not have addressed the real delay. Check 1: Capture the complete error and the exact time A screenshot containing only “Connection Timeout Expired” is rarely enough. Capture: The complete exception and inner exception. The operation being performed. The driver and version. The authentication method. The exact timestamp in UTC. Whether the issue affects every connection or only some of them. The wording around the timeout matters. For example, a timeout while obtaining a connection from the pool points toward the application’s pooling and concurrency behavior. A pre-login or TLS error belongs to a different investigation. A command timeout after the connection was established is usually a query-performance problem rather than a connection problem. Check 2: Measure the application timeline The application should record important operations separately. A simple timeline can completely change the investigation: 10:14:20.100 Token acquisition started 10:14:28.400 Token acquired 10:14:28.405 SQL connection started 10:14:29.050 SQL connection established The complete operation took almost nine seconds, but Azure SQL connection establishment took less than one second. Useful measurements include: Token-acquisition duration. Time waiting for a pooled connection. SQL connection-open duration. SQL command duration. Number of retry attempts. Applications using Microsoft Entra authentication must obtain an access token before authenticating to Azure SQL. Measuring that operation separately helps distinguish an identity delay from a database connectivity problem. This is particularly useful when the issue appears: On the first connection after startup. After a token expires. Only with Managed Identity or Workload Identity. Intermittently, while SQL authentication connections remain unaffected. Check 3: Test from the application environment A successful connection from a laptop does not validate the path used by an application running in: Azure App Service. Azure Functions. Azure Kubernetes Service. A virtual machine. An on-premises application server. A container or integration runtime. The laptop and the application may use different DNS servers, routes, firewalls, proxies and identities. Connectivity and DNS tests should therefore be performed from the environment that is actually failing. This becomes especially important when Private Endpoint is used. The application should continue connecting with: <server>.database.windows.net It should not use the Private Endpoint IP address or the privatelink.database.windows.net hostname directly. Direct login attempts using the private IP or the private-link FQDN fail; the normal logical-server FQDN must remain in the connection string. From the affected environment, confirm that: The expected DNS server answers the request. The server FQDN resolves to the expected private IP. The Private Endpoint connection is approved. The Private DNS zone is linked correctly. The resolved address is reachable through the intended route. A test from an unrelated machine is still useful for comparison, but it does not prove that the application path is healthy. Observed symptom Likely investigation area Timeout while obtaining a connection from the pool Application connection pooling Server name cannot be resolved DNS TCP connection to the endpoint cannot be established Network path, firewall or routing Error during the pre-login handshake TLS, driver, network interruption or pre-login processing Authentication or access-token error Microsoft Entra authentication, identity or token acquisition Timeout during the post-login phase Login completion, session initialization or server-side processing Execution or command timeout after connecting Query execution and database performance Avoid changing several things at once During a production incident, it is tempting to: Increase the timeout. Add firewall rules. Change the connection policy. Upgrade the driver. Restart the application. Clear connection pools. Applying several changes together makes it difficult to determine which one helped, and some may only hide the symptom. A better approach is to define one hypothesis: We believe DNS in the application environment is resolving the public endpoint instead of the Private Endpoint. Then define: The evidence supporting the hypothesis. One controlled change. The expected result. How the result will be measured. How the change will be reverted. Azure SQL supports Proxy and Redirect connection policies, which determine how traffic flows after reaching the Azure SQL gateway. The policy is configured for the logical server, so it should be verified before making firewall assumptions or changes. What should we collect before opening a support request? A small but precise evidence package can avoid several rounds of questions: Complete error and inner exception. Exact UTC timestamps. Application platform and location. Public or Private Endpoint. Server FQDN used by the application. Driver and version. Authentication method. Token, pool, connection and command durations. DNS result from the affected environment. Whether the issue is constant, intermittent or limited to the first connection. Recent application, network, identity or configuration changes.234Views0likes0CommentsAdvertised gateway prefixes in Azure
Introduction Large Azure hub-and-spoke environments can advertise a significant number of routes toward on-premises networks. By default, Azure VPN Gateway and ExpressRoute Gateway advertise the address spaces of the hub virtual network and the address spaces of peered spoke virtual networks that use gateway transit. As the number of spokes and address spaces grows, the border gateway protocol (BGP) route table grows as well. Advertised gateway prefixes provide a native way to summarize those Azure-side routes. The feature is configured on the hub virtual network through the summarizedGatewayPrefixes property, which is exposed in the Azure portal as Advertised gateway prefixes. Fewer BGP prefixes Replace many individual networks with one or more aggregated CIDRs. Better scale management Help large hub-and-spoke designs stay within advertised-prefix limits. Cleaner route visibility Make the intended Azure address plan easier to recognize on provider and on-premises route views. In this blog we will demonstrate leveraging Advertised gateway prefixes to summarize many Azure advertised prefixes into one. What are advertised gateway prefixes? Advertised gateway prefixes are summarized CIDR blocks that an Azure hybrid gateway advertises toward on-premises instead of advertising every covered hub and spoke address space individually. The configuration belongs on the gateway virtual network, usually the hub VNet that contains GatewaySubnet and the ExpressRoute or VPN gateway. How the route advertisement changes Default behavior With advertised gateway prefixes Hub: 10.27.0.0/24 Spoke 1: 10.27.1.0/24 Spoke 2: 10.27.2.0/24 Spoke 3: 10.27.3.0/24 Summary: 10.27.0.0/22 Four individual prefixes are advertised. One summary prefix is advertised. A spoke outside the summary is still advertised separately. Example: 172.16.1.0/24 remains visible if it is not covered by 10.27.0.0/22. When should you use it? You use a hub-and-spoke topology with gateway transit and many spoke address spaces. You want to advertise a covering prefix, such as a /16, instead of many smaller prefixes, such as multiple /24 networks. You are approaching ExpressRoute advertised-prefix limits or want to reduce route-table growth before scale becomes a problem. Your Azure address plan is sufficiently structured to create safe, intentional summary ranges. Note: ExpressRoute scale context: Microsoft documentation lists a maximum of 1,000 IPv4 prefixes and 100 IPv6 prefixes advertised from a virtual network to on-premises on a single ExpressRoute connection through private peering. Exceeding the connection prefix limit can cause the connection between the circuit and gateway to disconnect until the prefix count is reduced. Prerequisites A hub virtual network with GatewaySubnet. An ExpressRoute gateway or VPN gateway deployed in the hub virtual network. One or more peered spoke virtual networks if you want to demonstrate route summarization across spokes. A planned IPv4 and, when applicable, IPv6 summary that covers the intended hub and spoke address spaces. Access to the provider or on-premises BGP route view for validation. In this walkthrough, Megaport is used as the external verification point. Configure advertised gateway prefixes in the Azure portal The portal configuration is performed on the hub VNet, not on the ExpressRoute circuit and not on each spoke VNet. In the Azure portal, search for Virtual networks and select the hub VNet that contains GatewaySubnet. Open the hub virtual network Under Settings, select Address space. Open Address space In Advertised gateway prefixes, select + Add prefix. Enter the covering CIDR, such as 10.0.0.0/22 for four contiguous /24 networks. Add the summarized prefix For a dual-stack design, add IPv4 and IPv6 summarized prefixes explicitly. Add additional address families if needed Select Save and confirm the summarized prefixes remain listed in the Advertised gateway prefixes section. Save the configuration Validate the summarized route on routers or provider After Azure applies the change and BGP converges, I’m using ExpressRoute and Megaport here for my example, you can either log in to your router directly or view the incoming BGP routes via your provider’s interface to confirm what routes are being received from Azure. We want to capture a view that displays the BGP prefixes learned from the Azure side. Capture a baseline screenshot before enabling the feature, showing the individual hub and spoke prefixes. After configuration, refresh the route view and locate the new summarized prefix. Confirm that covered hub and spoke prefixes are no longer advertised individually. Confirm that any address space outside the configured summary remains advertised separately. Before: After: In my demo environment I have a hub and 20 spokes all within the 10.27.0.0/16 address prefix. We can see a successful implementation as before enabling the feature, I had a total of 21 prefixes being received at Megaport, after enabling the feature I have 1. Important design considerations Configure the hub, not the spokes: Only the virtual network containing the gateway subnet uses the summarizedGatewayPrefixes property for this behavior. A value placed on a spoke VNet is ignored. Avoid overlap within the prefix list: Do not configure overlapping entries in the advertised gateway prefixes list. Expect uncovered networks to remain visible: If a hub or spoke address space is not covered by a configured summary, the gateway continues to advertise it individually. Plan dual-stack explicitly: IPv4 and IPv6 summaries must be added separately. Removing all entries restores default behavior: When every advertised gateway prefix is removed, Azure returns to advertising hub and peered spoke address spaces individually. Protect the on-premises edge: Use appropriate route policies so only expected prefixes are accepted. Summarization simplifies advertisements, but it does not replace routing governance. Conclusion Advertised gateway prefixes give Azure networking teams a straightforward, native method to control route scale from a gateway-enabled hub VNet. Instead of sending every hub and spoke prefix across ExpressRoute or VPN connections, the gateway can advertise a concise list of summarized prefixes. The result is a smaller and more intentional route advertisement, while uncovered address spaces remain visible for compatibility. References Advertised gateway prefixes in Azure virtual networks Configure advertised gateway prefixes using the Azure portal Azure ExpressRoute FAQ About ExpressRoute virtual network gateways410Views0likes0CommentsInter-Hub Connectivity Using Azure Route Server
By Mays_Algebary shruthi_nair As your Azure footprint grows with a hub-and-spoke topology, managing User-Defined Routes (UDRs) for inter-hub connectivity can quickly become complex and error-prone. In this article, we’ll explore how Azure Route Server (ARS) can help streamline inter-hub routing by dynamically learning and advertising routes between hubs, reducing manual overhead and improving scalability. Baseline Architecture The baseline architecture includes two Hub VNets, each peered with their respective local spoke VNets as well as with the other Hub VNet for inter-hub connectivity. Both hubs are connected to local and remote ExpressRoute circuits in a bowtie configuration to ensure high availability and redundancy, with Weight used to prefer the local ExpressRoute circuit over the remote one. To maintain predictable routing behavior, the VNet-to-VNet configuration on the ExpressRoute Gateway should be disabled. Note: Adding ARS to an existing Hub with Virtual Network Gateway will cause downtime that expect to last 10 minutes. Scenario 1: ARS and NVA Coexist in the Hub Option A: Full Traffic Inspection ARS and NVA Coexist in the Hub In this scenario, ARS is deployed in each Hub VNet, alongside the Network Virtual Appliances (NVAs). NVA1 in Region1 establishes BGP peering with both the local ARS (ARS1) and the remote ARS (ARS2). Similarly, NVA2 in Region2 peers with both ARS2 (local) and ARS1 (remote). Let’s break down what each BGP peering relationship accomplishes. For clarity, we’ll focus on Region1, though the same logic applies to Region2: NVA1 Peering with Local ARS1 Through BGP peering with ARS1, NVA1 dynamically learns the prefixes of Spoke1 and Spoke2 at the OS level, eliminating the need to manually configure these routes. The same applies for NVA2 learning Spoke3 and Spoke4 prefixes via its BGP peering with ARS2. NVA1 Peering with Remote ARS2 When NVA1 peers with ARS2, the Spoke1 and Spoke2 prefixes are propagated to ARS2. ARS2 then injects these prefixes into NVA2 at both the NIC level with NVA1 as the next hop, and at the OS level. This mechanism removes the need for UDRs on the NVA subnets to enable inter-hub routing. Additionally, ARS2 advertises the Spoke1 and Spoke2 prefixes to both ExpressRoute circuits (EXR2 and EXR1 due to bowtie configuration) via GW2. 👉Important: To ensure that ARS2 accepts and propagates Spoke1/Spoke2 prefixes received via NVA1, AS-Override must be enabled. Without AS-Override, BGP loop prevention will block these routes at ARS2, since both ARS1 and ARS2 use the default ASN 65515, and ARS2 will consider the route as already originated locally. The same principle applies in reverse for Spoke3 and Spoke4 prefixes being advertised from NVA2 to ARS1. Traffic Flow Inter-Hub Traffic: Spoke VNets are configured with UDRs that contain only a default route (0.0.0.0/0) pointing to the local NVA as the next hop. Additionally, the “Propagate Gateway Routes” setting should be set to False to ensure all traffic, whether East-West (intra-hub/inter-hub) or North-South (to/from internet), is forced through the local NVA for inspection. Local NVAs will have the next hop to the other region spokes injected at the NIC level by local ARS, pointing to the other region NVA, for example NVA2 will have next hop to Spoke1 and Spoke2 as NVA1 (10.0.1.4) and vice versa. Why are UDRs still needed on spokes if ARS handles dynamic routing? Even with ARS in place, UDRs are required to maintain control of the next hop for traffic inspection. For instance, if Spoke1 and Spoke2 do not have UDRs, they will learn the remote spoke prefixes (e.g., Spoke3/Spoke4) injected via ARS1, which received them from NVA2. This results in Spoke1/Spoke2 attempting to route traffic directly to NVA2, a path that is invalid, since the spokes don’t have the path to NVA2. The UDR ensures traffic correctly routes through NVA1 instead. On-Premises Traffic: To explain the on-premises traffic flow, we'll break it down into two directions: Azure to on-premises, and on-premises to Azure. Azure to On-Premises Traffic Flow: As previously noted, Spokes send all traffic, including traffic to on-premises, via NVA1 due to the default route in the UDR. NVA1 then routes traffic to the local ExpressRoute circuit, using Weight to prefer the local path over the remote. Note: While NVA1 learns on-premises prefixes from both local and remote ARSs at the OS level, this doesn’t affect routing decisions. The actual NIC-level route injection determines the next hop, ensuring traffic is sent via the correct path—even if the OS selects a different “best” route internally. The screenshot below from NVA1 shows four next hops to the on-premises network 10.2.0.0/16. These include the local ARS (ARS1: 10.0.2.5 and 10.0.2.4) and the remote ARS (ARS2: 10.1.2.5 and 10.1.2.4). On-Premises to Azure Traffic Flow In a bowtie ExpressRoute configuration, Azure VNet prefixes are advertised to on-premises through both local and remote ExpressRoute circuits. Because of this dual advertisement, the on-premises network must ensure optimal path selection when routing traffic to Azure. From Azure side, to maintain traffic symmetry, add UDRs at the GatewaySubnet (GW1 and GW2) with specific routes to the local Spoke VNets, using the local NVA as the next hop. This ensures return traffic flows back through the same path it entered. 👉How Does the ExpressRoute Edge Router Select the Optimal Path? You might ask: If Spoke prefixes are advertised by both GW1 and GW2, how does the ExpressRoute edge router choose the best path? (e.g., diagram below shows EXR1 learns Region1 prefixes from GW1 and GW2) Here’s how: Edge routers (like EXR1) receive the same Spoke prefixes from both gateways. However, these routes have different AS-Path lengths: - Routes from the local gateway (GW1) have a shorter AS-Path. - Routes from the remote gateway (GW2) have a longer AS-Path because NVA1’s ASN (e.g., 65001) is prepended twice as part of the AS-Override mechanism. As a result, the edge router (EXR1) will prefer the local path from GW1, ensuring efficient and predictable routing. For example: EXR1 receives Spoke1, Spoke2, and Hub1-VNet prefixes from both GW1 and GW2. But because the path via GW1 has a shorter AS-Path, EXR1 will select that as the best route. (Refer to the diagram below for a visual of the AS-Path difference). Final Traffic Flow: Option-A Insights: This design simplifies UDR configuration for inter-hub routing, especially useful when dealing with non-contiguous prefixes or operating across multiple hubs. For simplicity, we used a single NVA in each Hub-VNet while explaining the setup and traffic flow throughout this article. However, a high available (HA) NVA deployment is recommended. To maintain traffic symmetry in an HA setup, you’ll need to enable the next-hop IP feature when peering with Azure Route Server (ARS). When on-premises traffic inspection is required, the UDR setup in the GatewaySubnet becomes more complex as the number of Spokes increases. As your Azure network scales, keep in mind that Azure Route Server supports a maximum of 16 BGP peers per instance (as of the time writing this article). This limit can impact architectures involving multiple NVAs or hubs. Option B: Bypass On-Premises Inspection If on-premises traffic inspection is not required, NVAs can advertise a supernet prefix summarizing the local Spoke VNets to the remote ARS. This approach provides granular control over which traffic is routed through the NVA and eliminates the need for BGP peering between the local NVA and local ARS. All other aspects of the architecture remain the same as described in Option A. For example, NVA2 can advertise the supernet 192.168.2.0/23 (supernet of Spoke3 and Spoke4) to ARS1. As a result, Spoke1 and Spoke2 will learn this route with NVA2 as the next hop. To ensure proper routing (as discussed earlier) and inter-hub inspection, you need apply a UDR in Spoke1 and Spoke2 that overrides this exact supernet prefix, redirecting traffic to NVA1 as the next hop. At the same time, traffic destined for on-premises will follow the system route through the local ExpressRoute gateway, bypassing NVA1 altogether. In this setup: UDRs on the Spokes should have "Propagate Gateway Routes" set to True. No UDRs are needed in the GatewaySubnet. 👉Can NVA2 Still Advertise Specific Spoke Prefixes? You might wonder: Can NVA2 still advertise specific prefixes (e.g., Spoke3 and Spoke4) learned from ARS2 to ARS1 instead of a supernet? Yes, this is technically possible, but it requires maintaining BGP peering between NVA2 and ARS2. However, this introduces UDR complexity in Spoke1 and Spoke2, as you'd need to manually override each specific prefix. This also defeats the purpose of using ARS for simplified route propagation, undermining the efficiency and scalability of the design. Bypass On-Premises Inspection Final Traffic Flow: Option B: Bypass on-premises inspection traffic flow Option-B Insights: This approach reduces the number of BGP peerings per ARS. Instead of maintaining two BGP sessions (local NVA and remote NVA) per Hub, you can limit it to just one, preserving capacity within ARS’s 8-peer limit for additional inter-hub NVA peerings. Each NVA should advertise a supernet prefix to the remote ARS. This can be challenging if your Spokes don’t use contiguous IP address spaces, as described in Option B. Scenario 2: ARS in the Hub and the NVA in Transit VNet In Scenario 1, we highlighted that when on-premises inspection is required, managing UDRs at the GatewaySubnet becomes increasingly complex as the number of Spoke VNets grows. This is due to the need for UDRs to include specific prefixes for each Spoke VNet. In this scenario, we eliminate the need to apply UDRs at the GatewaySubnet altogether. A in Transit VNet In this design, the NVA will be deployed in Transit VNet, where: Transit-VNet will be peered with local Spoke VNets and with the local Hub-VNet to enable intra-Hub and on-premises connectivity. Transit-VNet also peered with remote Transit VNets (e.g., Transit-VNet1 peered with Transit-VNet2) to handle inter-Hub connectivity through the NVAs. Additionally, Transit-VNets are peered with remote Hub-VNets, to establish BGP peering with the remote ARS. NVAs OS will need to add static routes for the local Spoke VNets prefixes, it can be specific or it can supernet prefix, which will later be advertised to ARSs over BGP Peering, then ARS will advertise it to on-premises via ExpressRoute. NVAs will BGP peer with local ARS and also with the remote ARS. To understand the reasoning behind this design, let’s take a closer look at the setup in Region1, focusing on how ARS and NVA are configured to connect to Region2. This will help illustrate both inter-hub and on-premises connectivity. The same concept applies in reverse from Region2 to Region1. Inetr-Hub: To enable NVA1 in Region1 to learn prefixes from Region2, NVA2 will configure static routes at the OS level for Spoke3 and Spoke4 (or their supernet prefix) and advertise them to ARS1 via remote BGP peering. As a result, these prefixes will be received by NVA1, both at the NIC level, with NVA2 as the next hop, and at the OS level for proper routing. Spoke1 and Spoke2 will have a UDR with a default route pointing to NVA1 as the next hop. For instance, when Spoke1 needs to communicate with Spoke3, the traffic will first route through NVA1. NVA1 will then forward the traffic to NVA2 using VNet peering between the two Hubs. A similar configuration will be applied in Region2, where NVA1 will configure static routes at the OS level for Spoke1 and Spoke2 (or their supernet prefix) and advertise them to ARS2 via remote BGP peering, as a result, these prefixes will be received by NVA2, both at the NIC level (injected by ARS2), with NVA1 as the next hop, and at the OS level for proper routing. Note: At the OS level, NVA1 learns Spoke3 and Spoke4 prefixes from both local and remote ARSs. However, the NIC-level route injection determines the actual next hop, so even if the OS selects a different best route, it won’t affect forwarding behavior. same applies to NVA2. On-Premises Traffic: To explain the on-premises traffic flow, we'll break it down into two directions: Azure to on-premises, and on-premises to Azure. Azure to On-Premises Traffic Flow: Spokes in Region1 route all traffic through NVA1 via a default route defined in their UDRs. Because of BGP peering between NVA1 and ARS1, ARS1 advertises the Spoke1 and Spoke2 (or their supernet prefix) to on-premises through ExpressRoute (EXR1). The Transit-VNet1 (hosting NVA1) is peered with Hub1-VNet, with “Use Remote Gateway” enabled. This allows NVA1 to learn on-premises prefixes from the local ExpressRoute gateway (GW1), and traffic to on-premises is routed through the local ExpressRoute circuit (EXR1) due to higher BGP Weight configuration. Note: At the OS level, NVA1 learns on-prem prefixes from both local and remote ARSs. However, the NIC-level route injection determines the actual next hop, so even if the OS selects a different best route, it won’t affect forwarding behavior. same applies to NVA2. On-Premises to Azure Traffic Flow: Through BGP peering with ARS1, NVA1 enables ARS1 to advertise Spoke1 and Spoke2 (or their supernet prefix) to both EXR1 and EXR2 circuits (due to the ExpressRoute bowtie setup). Additionally, due to BGP peering between NVA1 and ARS2, ARS2 also advertises Spoke1 and Spoke2 (or their supernet prefix) to EXR2 and EXR1 circuits. As a result, both ExpressRoute edge routers in Region1 and Region2 learn the same Spoke prefixes (or their supernet prefix) from both GW1 and GW2, with identical AS-Path lengths, as shown below. EXR1 learns Region1 Spokes's supernet prefixes from GW1 and GW2 This causes non-optimal inbound routing, where traffic from on-premises destined to Region1 Spokes may first land in Region2’s Hub2-VNet before traversing to NVA1 in Region1. However, return traffic from Spoke1 and Spoke2 will always exit through Hub1-VNet. To prevent suboptimal routing, configure NVA1 to prepend the AS path for Spoke1 and Spoke2 (or their supernet prefix) when advertising them to the remote ARS2. Likewise, ensure NVA2 prepends the AS path for Spoke3 and Spoke4 (or their supernet prefix) when advertising to ARS1. This approach helps maintain optimal routing under normal conditions and during ExpressRoute failover scenarios. Below diagram shows NVA1 is setting AS-Prepend for Spoke1 and Spoke2 supernet prefix when BGP peer with remote ARS (ARS1), same will apply for NVA2 when advertising Spoke3 and Spoke4 prefixes to ARS1. Final Traffic Flow: Full Inspection: Traffic flow when NVA in Transit-VNet Insights: This solution is ideal when full traffic inspection is required. Unlike Scenario 1 - Option A, it eliminates the need for UDRs in the GatewaySubnet. When ARS is deployed in a VNet (typically in Hub VNets), the VNet will be limited to 500 VNet peerings (as of the time writing this article). However, in this design, Spokes peer with the Transit-VNet instead of directly with the ARS VNet, allowing you to scale beyond the 500-peer limit by leveraging Azure Virtual Network Manager (AVNM) or submitting a support request. Some enterprise customers may encounter the 1,000-route advertisement limit on the ExpressRoute circuit from the ExpressRoute gateway. With Summarized Gateway Prefixes now generally available, Azure provides native control to summarize VNet address spaces before advertising them to ExpressRoute. This helps reduce the advertised route count without relying solely on NVAs for summarization. For simplicity, we used a single NVA in each Hub-VNet while explaining the setup and traffic flow throughout this article. However, a high available (HA) NVA deployment is recommended. To maintain traffic symmetry in an HA setup, you’ll need to enable the next-hop IP feature when peering with Azure Route Server (ARS). This design does require additional VNet peerings, including: Between Transit-VNets (inter-region), Between Transit-VNets and local Spokes, and Between Transit-VNets and both local and remote Hub-VNets.3.5KViews5likes2CommentsAzure Incident Retrospective - Please register! Session 2 - Tracking ID: ZJV6-SGG-BW8
Join our upcoming live webcast for a transparent discussion about this recent Azure service incident led by our engineering teams. Network connectivity issues in West US Tracking ID: ZJV6-SGG | Impacted: 23 July 2026 Same content presented in both sessions: pick the one that works best for your timezone! What to expect 📚 Understand What happened, how we responded, and what we learned 💬 Ask Live Q&A with our engineering experts throughout the session 🛠 Learn The fixes we've put in place and guidance for workload resiliency Choose your session Same content presented at both times: pick the one that works best for your timezone: Session 1 17:30 UTC Thursday, 27 Aug 2026 Register now → Session 2 05:30 UTC Friday, 28 Aug 2026 Register now → 10:30 AM US Pacific (PDT) 1:30 PM US Eastern (EDT) 6:30 PM London (BST) 1:30 AM +1 Beijing (CST) 3:30 AM +1 Sydney (AEDT) 5:30 AM +1 Auckland (NZDT) 10:30 PM -1 US Pacific (PDT) 1:30 AM US Eastern (EDT) 6:30 AM London (BST) 1:30 PM Beijing (CST) 3:30 PM Sydney (AEDT) 5:30 PM Auckland (NZDT) Our engineering leaders Jamie Gaudette Vice President Azure Networking Cloud+AI Engineering LinkedIn ↗ ⚠️ Prepare before the livestream Read the Post Incident Review (PIR) ahead of time so you can ask any follow up questions during the live Q&A Helpful resources 🔔 Azure Service Health Alerts Get alerts for relevant incidents by setting up notifications via email, SMS, or webhook 🎥 Past Retrospective Recordings Watch recordings of previous retrospective livestreams 📄 Azure Post Incident Reviews Learn more about PIRs and the retrospective program149Views0likes0CommentsAzure Incident Retrospective - Please register! Session 1 - Tracking ID: ZJV6-SGG-BW8
Join our upcoming live webcast for a transparent discussion about this recent Azure service incident led by our engineering teams. Network connectivity issues in West US Tracking ID: ZJV6-SGG | Impacted: 23 July 2026 Same content presented in both sessions: pick the one that works best for your timezone! What to expect 📚 Understand What happened, how we responded, and what we learned 💬 Ask Live Q&A with our engineering experts throughout the session 🛠 Learn The fixes we've put in place and guidance for workload resiliency Choose your session Same content presented at both times: pick the one that works best for your timezone: Session 1 17:30 UTC Thursday, 27 Aug 2026 Register now → Session 2 05:30 UTC Friday, 28 Aug 2026 Register now → 10:30 AM US Pacific (PDT) 1:30 PM US Eastern (EDT) 6:30 PM London (BST) 1:30 AM +1 Beijing (CST) 3:30 AM +1 Sydney (AEDT) 5:30 AM +1 Auckland (NZDT) 10:30 PM -1 US Pacific (PDT) 1:30 AM US Eastern (EDT) 6:30 AM London (BST) 1:30 PM Beijing (CST) 3:30 PM Sydney (AEDT) 5:30 PM Auckland (NZDT) Our engineering leaders Jamie Gaudette Vice President Azure Networking Cloud+AI Engineering LinkedIn ↗ ⚠️ Prepare before the livestream Read the Post Incident Review (PIR) ahead of time so you can ask any follow up questions during the live Q&A Helpful resources 🔔 Azure Service Health Alerts Get alerts for relevant incidents by setting up notifications via email, SMS, or webhook 🎥 Past Retrospective Recordings Watch recordings of previous retrospective livestreams 📄 Azure Post Incident Reviews Learn more about PIRs and the retrospective program161Views1like0CommentsSecure Native Access to Azure Kubernetes Service (AKS) Private Clusters with Azure Bastion
Written in Collaboration with AvirupChat ShabazShaik Mohit_Kumar YuriDiogenes Introduction: As organizations move toward containerized workloads such as Azure Kubernetes Service (AKS), secure access to cluster resources becomes critical. Making the Kubernetes API server private, therefore, significantly reduces the attack surface. However, it also changes how engineers perform routine management tasks. Operations as simple as running kubectl logs, describing resources, or troubleshooting workloads require connectivity to the virtual network hosting the cluster. This challenge becomes even more apparent during day-to-day operations. An engineer may have the necessary Kubernetes credentials and Azure RBAC permissions yet still be unable to access the cluster because the API server is reachable only from within the private network. Establishing VPN connectivity, using jump hosts, or deploying dedicated management workstations often becomes part of the operational workflow, adding complexity to what should be straightforward administrative tasks. This balance between maintaining strong network isolation and enabling efficient cluster management is exactly what Azure Bastion's native client tunneling support for private AKS clusters is designed to address. Many Azure customers already use Azure Bastion to securely access virtual machines without exposing them to the public internet. Native client tunneling now brings that same secure access model to AKS private clusters. This capability is in public preview currently. Before discussing how it works, it is worth understanding the operational challenges it is designed to solve and how it's done. Bastion for Azure Kubernetes Service (AKS) private clusters: At its core, Azure Bastion is designed to simplify secure access to private Azure resources. Since its launch, it has provided browser-based and native client RDP and SSH connectivity to Azure VMs without exposing management ports to the internet. No public IPs on your VMs. No inbound NSG rules for port 22 or 3389. Just TLS over port 443, routed through a fully managed service that Microsoft patches, scales, and secures on your behalf. When you establish a Bastion tunnel to your private AKS cluster, you are not signing in to an intermediary machine and running kubectl from there. You are opening an encrypted tunnel from your local machine, through Bastion, directly to the private API server endpoint. Your kubeconfig points to localhost on a dynamically assigned port, so local kubectl, helm, and scripts continue to work as they would against a public cluster. You get the developer experience of a public cluster with the security posture of a private one as shown in Figure 1. The high-level flow has three parts: The engineer authenticates to Azure, retrieves cluster credentials, and uses Azure Bastion to establish a managed tunnel into the private network. Local Kubernetes tools then use that tunnel to reach the private API server. This keeps the workflow familiar while removing the need for jump hosts or VPN-based access paths. However, this raises an important question - Does making access easier also make it easier for an attacker to connect? The short answer is no. In fact, Bastion tunneling strengthens the security posture in ways that are worth unpacking. First, the API server itself remains private. There is no public endpoint. There is no public IP address discoverable by scanners. The only network path to the API server runs through Azure's managed infrastructure. Authentication and authorization then depend on the AKS cluster configuration. Second, Bastion eliminates an entire class of infrastructure that itself becomes a security target. Jump box VMs, when not meticulously maintained, accumulate credentials, kubeconfig files, and browser sessions. They are machines that admins log into, which means they are machines that can be compromised. Bastion is not a machine you log into. It is an agentless, managed tunnel you pass through. And for public clusters, there is a related but distinct benefit. Many teams use API server authorized IP ranges to restrict which source IPs can reach their public API endpoint. This is a good practice, but it breaks down quickly when your team works remotely, uses dynamic IPs, or includes contractors. Adding Bastion's stable public IP to the authorized range gives you a consistent, managed access path without the operational toil of constantly updating IP allow lists. With the network path established through Bastion, the next question is how users are authenticated and authorized once they reach the AKS API server. That distinction matters: Bastion provides secure connectivity, while AKS access is governed by the authentication and authorization model configured on the cluster. Authentication and authorization options: When creating an AKS cluster, you can choose from three authentication and authorization modes, as shown in Figure below. That choice shapes everything about how access is granted, audited, and governed over the lifetime of the cluster. ⚠️Local accounts with Kubernetes RBAC is the default if you change nothing. It uses a static certificate that never expires, is shared across all administrators, and has no connection to your corporate identity provider. There is no MFA. No Conditional Access. Note: For most production environments, Microsoft recommends using Microsoft Entra ID integrated authentication rather than local accounts to enable centralized identity, MFA, and auditing. ☑️ Microsoft Entra ID with Kubernetes RBAC moves authentication to Entra ID, which means users sign in with their corporate identity, MFA is enforced, and Conditional Access policies apply. Authorization is handled through native Kubernetes Role and ClusterRole bindings, which reference Entra ID users and groups as subjects. This works well for teams that manage cluster configuration through GitOps, because RBAC manifests live alongside other cluster YAML. The limitation is that these permissions are invisible in Azure IAM — they live only inside the cluster. ✅ Microsoft Entra ID with Azure RBAC is the model we recommend for most production environments. Authentication still flows through Entra ID with full MFA and Conditional Access support. But authorization is handled by Azure RBAC role assignments on the AKS resource itself. Permissions are visible in Azure IAM, participate in access reviews, and integrate with Privileged Identity Management for just-in-time elevation. You can assign the built-in Azure Kubernetes Service RBAC Cluster Admin, Admin, Writer, or Reader roles at either cluster scope or namespace scope. A single subscription-level role assignment can grant access to every cluster in the subscription. The real value is the combination. Bastion provides the encrypted network tunnel. Entra ID provides the identity, MFA, and Conditional Access. Azure RBAC provides centralized, auditable authorization that your security team can review alongside every other Azure resource. For the exact steps to configure Entra ID authentication and Azure RBAC on your AKS cluster, see the AKS identity documentation. End-to-end flow for connectivity to AKS via Bastion: az account set --subscription <subscription ID> Retrieve credentials to your AKS private cluster using the commands below: az aks get-credentials --name <AKSClusterName> --resource-group <ResourceGroupName> Open the tunnel to your target AKS Cluster with the following command: az aks bastion --name <aksClusterName> --resource-group <aksClusterResourceGroup> --bastion <bastionResourceId> Now the default authentication method is Device code authentication i.e., this authentication method prompts the device code for the user to sign in from a browser session. If you want CLI only authentication, you can run the following command next. kubelogin convert-kubeconfig -l azurecli Then go on with your AKS connectivity: kubectl get nodes How to Try It: If you are already running private AKS clusters, getting started is straightforward. You need a Standard or Premium Azure Bastion host deployed in the same VNet as your cluster (or in a peered VNet), with native client support enabled. The aks-preview and bastion CLI extensions handle the rest. The detailed connection steps are documented on Microsoft Learn. We Want Your Feedback Write a comment in this blog or open an issue on the AKS GitHub repository or leave feedback directly on the Microsoft Learn documentation page. We read everything, and it genuinely influences what we prioritize.698Views2likes2Comments