Deploy an Azure Container Apps (ACA) environment with a dedicated D4 workload profile in West US, run simple apps, and observe how replica count and reserved CPU/memory drive compute consumption. Learn to monitor with Azure Monitor metrics and Grafana, alert on over-provisioned (low-utilization) apps, build dashboards, and understand why right-sizing CPU/memory matters far more in ACA than in AKS.
Welcome to Part 2 of our three-part Azure Container Apps series. This time, we’ll explore how to monitor resource consumption, detect over-provisioned applications, and right-size ACA workloads for better efficiency and cost control.
Imagine a small API with quiet periods, short traffic spikes, and a surprisingly high monthly bill. Its replicas use little CPU most of the day, but they reserve enough capacity to affect how workloads pack onto Dedicated nodes. This blog shows how to distinguish application allocation from billed node capacity, investigate right-sizing opportunities, and verify changes without compromising peak performance.
1. The key idea: ACA reserves — it does not “request”
This is the most important concept in this series.
In Kubernetes/AKS a Pod container has two independent knobs:
|
Knob |
Meaning |
Effect |
|
requests.cpu/memory |
Guaranteed floor used for scheduling and reserved capacity |
You can set this low so many Pods pack onto a node (oversubscription) |
|
limits.cpu/memory |
Hard ceiling a container may burst to |
You can set this high so a bursty app gets headroom |
So in AKS you can say “request 100m CPU, but allow bursting up to 2 cores.” The node is oversubscribed; you only pay for the node, and idle apps cost you almost nothing at the request level.
In Azure Container Apps there is only one knob per container: cpu and memory. That single value is both the request and the limit; it is fully reserved for every replica, for the entire time the replica runs.
One reserved knob in ACA vs. request/limit in AKS
Consequences you will demonstrate in this exercise:
- There is no way to set a low floor and a high ceiling in ACA. If you allocate `cpu = 1.0`, each running replica reserves 1 vCPU of node capacity even if the app idles at 2% CPU. This reduces packing density and can cause the dedicated profile to add another node, directly increasing node-based billing.
- Right-sizing the cpu/memory allocation is the single biggest cost lever — much more so than in AKS, where a too-large limit is “free” until the app actually bursts.
- On a dedicated profile, unused reserved capacity on a node is money already spent — the node is billed regardless of app utilization.
Allocation rules
|
Plan |
CPU:Memory rule |
Range |
|
Consumption profile |
Memory (GiB) must be exactly 2× vCPU |
0.25 vCPU / 0.5 GiB … up to 4 vCPU / 8 GiB |
|
Dedicated profile (D4, D8, …) |
Flexible ratio, any combo up to the profile node size |
Up to node size minus a small system reservation |
A D4 node = 4 vCPU / 16 GiB. A small amount (~0.5 vCPU / ~1 GiB) is reserved by the platform, so usable capacity per node is roughly 3.5 vCPU / ~15 GiB for your replicas combined.
2. Topology
ACA environment topology with dedicated D4 profile and observability sinks
3. Prerequisites & variables
# Azure CLI + containerapp extension
az extension add --name containerapp --upgrade
az provider register --namespace Microsoft.App
az provider register --namespace Microsoft.OperationalInsights
# Workshop variables
export RG="aca-obs-rg"
export LOC="westus"
export ENV="aca-obs-env"
export PROFILE="d4prof" # dedicated workload profile name
export APP="cpu-demo"
export ACR="srinmantest" # existing registry in this repo - use your own ACR for your test
4. Step 1 — Create the environment with a dedicated D4 profile
az group create --name "$RG" --location "$LOC"
# Environment with workload profiles enabled (starts with a default Consumption profile)
az containerapp env create \
--name "$ENV" \
--resource-group "$RG" \
--location "$LOC" \
--enable-workload-profiles
# Add a DEDICATED D4 profile (4 vCPU / 16 GiB per node)
az containerapp env workload-profile add \
--name "$ENV" \
--resource-group "$RG" \
--workload-profile-name "$PROFILE" \
--workload-profile-type D4 \
--min-nodes 1 \
--max-nodes 3
Verify the profile and note the billing model:
az containerapp env workload-profile list \
--name "$ENV" --resource-group "$RG" \
--query "[].{name:name, type:properties.workloadProfileType, min:properties.minimumCount, max:properties.maximumCount}" -o table
Billing note (dedicated plan): you pay for the vCPU/GiB of the profile nodes that are running (between min-nodes and max-nodes), plus a per-environment management fee — not for actual app utilization. A min-nodes 1 D4 profile bills for one full 4 vCPU / 16 GiB node even if it is nearly idle. See the ACA pricing page (Dedicated plan).
5. Step 2 — Deploy a simple app with an explicit CPU/memory reservation
We intentionally over-allocate (1 vCPU / 2 GiB) to a lightweight app so the over-provisioning story is visible later.
az containerapp create \
--name "$APP" \
--resource-group "$RG" \
--environment "$ENV" \
--workload-profile-name "$PROFILE" \
--image mcr.microsoft.com/azuredocs/containerapps-helloworld:latest \
--target-port 80 \
--ingress external \
--cpu 1.0 --memory 2.0Gi \
--min-replicas 1 --max-replicas 1 \
--query properties.configuration.ingress.fqdn -o tsv
Key points: - --cpu 1.0 --memory 2.0Gi is reserved per replica. The helloworld app uses a tiny fraction of that. On the dedicated D4 node, this replica consumes 1 of ~3.5 usable vCPU. You could fit ~3 such replicas before a second node is added.
Capture the app resource ID for later metric/alert commands:
export APP_ID=$(az containerapp show -n "$APP" -g "$RG" --query id -o tsv)
export ENV_ID=$(az containerapp env show -n "$ENV" -g "$RG" --query id -o tsv)
echo "$APP_ID"
6. Step 3 — Demonstrate: more replicas = more reserved compute
Force replicas up and watch reserved cores and node count grow. Pin the replica count so scaling is deterministic:
First, capture the baseline before scaling — how many profile nodes are running and how much is reserved vs. actually used:
The baseline uses five metrics — the first is environment-scoped, the rest are app-scoped:
|
Metric |
Scope |
Aggregation |
What it tells you |
|
NodeCount |
environment ($ENV_ID) |
Maximum |
Number of D4 nodes the workload profile is currently running — each node is billed |
|
CoresQuotaUsed |
app ($APP_ID) |
Maximum |
vCPU cores reserved by the app = replicas × cpu allocation |
|
Replicas |
app ($APP_ID) |
Maximum |
Active replica count driving the reservation |
|
UsageNanoCores |
app ($APP_ID) |
Average |
Actual CPU consumed (1e9 nanocores = 1 core) — usually a tiny fraction of the reservation |
|
CpuPercentage (preview) |
app ($APP_ID) |
Average |
Actual CPU as a % of the reserved limit — the gap between this and 100% is pre-paid waste |
CoresQuotaUsed/Replicas use Maximum (a reservation is a step value — you want its peak in the interval). UsageNanoCores/CpuPercentage use Average (usage fluctuates — the mean over the window is more representative). NodeCount is a preview metric emitted per minute.
# How many D4 nodes the workload profile is currently running (baseline should be 1)
az monitor metrics list --resource "$ENV_ID" \
--metric NodeCount --aggregation Maximum --interval PT1M -o table
# Reserved cores right now
az monitor metrics list --resource "$APP_ID" \
--metric CoresQuotaUsed --aggregation Maximum --interval PT1M -o table
# Active replica count driving that reservation
az monitor metrics list --resource "$APP_ID" \
--metric Replicas --aggregation Maximum --interval PT1M -o table
# Actual CPU consumed (absolute nanocores)
az monitor metrics list --resource "$APP_ID" \
--metric UsageNanoCores --aggregation Average --interval PT5M -o table
# Actual CPU as a % of the reserved limit
az monitor metrics list --resource "$APP_ID" \
--metric CpuPercentage --aggregation Average --interval PT5M -o table
Note the current node count and reserved cores — you’ll compare against these after scaling out.
# Force 4 replicas (4 × 1 vCPU = 4 reserved cores → exceeds one D4 node's usable ~3.5 → 2nd node added)
az containerapp update \
--name "$APP" --resource-group "$RG" \
--min-replicas 4 --max-replicas 4
Observe the effect on reserved compute (three complementary metrics):
# Reserved cores for the app (grows linearly with replica count)
az monitor metrics list --resource "$APP_ID" \
--metric CoresQuotaUsed --aggregation Average --interval PT1M -o table
# Replica count
az monitor metrics list --resource "$APP_ID" \
--metric Replicas --aggregation Maximum --interval PT1M -o table
# Node count of the workload profile (environment-level; more replicas → more billed nodes)
az monitor metrics list --resource "$ENV_ID" \
--metric NodeCount --aggregation Maximum --interval PT1M -o table
Replica scaling increases reserved cores
Takeaway: every replica reserves its full cpu/memory allocation. Scaling out multiplies reserved compute regardless of whether the app is busy. On a dedicated profile, crossing a node’s usable capacity provisions — and bills — another full D4 node.
Return to 1 replica when done:
az containerapp update --name "$APP" --resource-group "$RG" --min-replicas 1 --max-replicas 1
7. Step 4 — Check an app’s CPU & memory with Azure Monitor metrics
ACA emits platform metrics automatically to the Azure Monitor metrics store — no agent, no Azure Monitor workspace, no configuration required. Namespace: Microsoft.App/containerapps.
Metrics that matter for right-sizing
|
Metric ID |
Title |
Use |
|
UsageNanoCores |
CPU Usage |
Absolute CPU used (1e9 nanocores = 1 core) |
|
WorkingSetBytes |
Memory Working Set Bytes |
Absolute memory used |
|
CpuPercentage (preview) |
CPU Usage Percentage |
% of the reserved CPU limit used — the right-sizing signal |
|
MemoryPercentage (preview) |
Memory Percentage |
% of the reserved memory limit used |
|
CoresQuotaUsed / TotalCoresQuotaUsed |
Reserved Cores |
Cores reserved, independent of usage |
|
Replicas |
Replica count |
Active replicas |
|
Requests |
Requests |
HTTP requests (split by status code) |
The *Percentage metrics are the key ones for this blog: they compare actual usage against the reservation. A low percentage with a large allocation = wasted, pre-paid compute.
Portal
- Open the container app → Monitoring → Metrics.
- Scope is the app; namespace Container Apps.
- Add CPU Usage Percentage and Memory Percentage, aggregation Avg.
- Use Apply splitting → Replica to see per-replica usage.
CLI
# Absolute CPU (nanocores) and memory (bytes)
az monitor metrics list --resource "$APP_ID" \
--metric UsageNanoCores WorkingSetBytes \
--aggregation Average --interval PT1M -o table
# Utilization vs. the reservation (the numbers that reveal over-provisioning)
az monitor metrics list --resource "$APP_ID" \
--metric CpuPercentage MemoryPercentage \
--aggregation Average --interval PT5M -o table
Where to view metrics in the UI
You can inspect ACA metrics through three complementary browser experiences. They read from the same Azure Monitor platform metrics, but each is optimized for a different workflow:
|
UI |
How to open it |
Best for |
|
Container Apps portal |
Sign in at containerapps.azure.com, select the app, and open its monitoring experience |
ACA-focused operations: move quickly between apps, revisions, replicas, live logs, and app health without navigating the full Azure portal |
|
Azure Portal ACA Dashboards with Grafana |
In portal.azure.com, open the Container Apps environment or app and use its Monitoring dashboards / Grafana integration |
Curated ACA views that correlate app and environment signals, including replicas, requests, CPU, memory, and workload-profile capacity |
|
Azure Managed Grafana portal |
Open the Azure Managed Grafana resource in the Azure portal, select its Endpoint, then use Dashboards in Grafana |
Persistent, shared, multi-app dashboards; longer time ranges; alerting; and correlation with other Azure Monitor or Prometheus data |
Use the Container Apps portal for fast app-level troubleshooting, the Azure portal ACA dashboards for built-in service views, and the Azure Managed Grafana portal when you need a reusable cost and operations dashboard. For right-sizing, show CpuPercentage, MemoryPercentage, CoresQuotaUsed, and Replicas together so utilization is visible next to reserved capacity.
Data-source note: built-in ACA metrics are stored in Azure Monitor. Viewing them in any of these UIs does not require an Azure Monitor workspace. A workspace is needed only for custom Prometheus metrics or log-based queries.
Reading UsageNanoCores
UsageNanoCores is reported in nanocores, where 1 vCPU = 1,000,000,000 (1e9) nanocores. Divide by 1e9 to get cores:
|
Raw value |
÷ 1e9 = cores |
× 1000 = millicores |
vs. 1.0 vCPU reservation |
|
152743.5 |
0.00015 cores |
0.15 millicores |
≈ 0.015 % used |
|
250000000 |
0.25 cores |
250 millicores |
25 % used |
|
1000000000 |
1.0 core |
1000 millicores |
100 % used (saturated) |
So a reading like 152743.5 means the container is essentially idle — using ~0.15 millicores while you reserve a full vCPU. CpuPercentage shows the same story as a ready-made percentage (~0 %), so you don’t have to do the nanocore math for right-sizing decisions.
Memory works the same way: WorkingSetBytes is raw bytes — divide by 1024³ (≈1.074e9) for GiB and compare against the --memory reservation, or just read MemoryPercentage.
(Optional) Generate load to see CPU move
The helloworld image stays near-idle (great for the over-provisioning demo). To see CPU climb, hit the endpoint in a loop or deploy a CPU-busy image:
FQDN=$(az containerapp show -n "$APP" -g "$RG" --query properties.configuration.ingress.fqdn -o tsv)
# simple load
for i in $(seq 1 5000); do curl -s "https://$FQDN" > /dev/null; done
8. Step 5 — Monitor with Grafana
Grafana can chart ACA platform metrics directly — you do NOT need an Azure Monitor workspace or Prometheus for the built-in metrics. Azure Managed Grafana ships with a pre-authenticated Azure Monitor data source that queries the same metrics store as Metrics Explorer.
Grafana reads ACA metrics via the Azure Monitor data source
|
Path |
Data source |
Needs Azure Monitor workspace? |
|
Built-in ACA metrics (CPU%, Mem%, Replicas…) |
Azure Monitor |
❌ No |
|
Custom app Prometheus metrics |
Prometheus (scrape → remote-write) |
✅ Yes |
Use the existing Grafana instance and grant it access
This blog uses the existing amg-west Azure Managed Grafana instance in resource group infrarg (no new Grafana is created). Grant its managed identity read access to the ACA metrics resource group:
az extension add --name amg --upgrade
export GRAFANA="amg-west"
export GRAFANA_RG="infrarg"
# Grant the existing Grafana's managed identity read access to metrics on the ACA RG
GRAFANA_MI=$(az grafana show -n "$GRAFANA" -g "$GRAFANA_RG" --query identity.principalId -o tsv)
RG_ID=$(az group show -n "$RG" --query id -o tsv)
az role assignment create \
--assignee "$GRAFANA_MI" \
--role "Monitoring Reader" \
--scope "$RG_ID"
Build the panel
- Open the Grafana endpoint (az grafana show -n "$GRAFANA" -g "$GRAFANA_RG" --query properties.endpoint -o tsv).
- Dashboards → New → Add visualization → Azure Monitor data source.
- Service: Metrics; Resource: your container app; Namespace: Microsoft.App/containerapps.
- Metric: CPU Usage Percentage; add a second query for Memory Percentage and Replica count.
- Save the dashboard as “ACA — Reservation vs. Utilization.”
Tip: a panel that overlays CpuPercentage (usage) against a constant 100% line makes over-provisioning obvious — a flat 3% line under a 1 vCPU reservation is pre-paid waste.
9. Step 6 — Alert on over-provisioned (low-utilization) apps
Goal: Identify low-utilized apps and investigate right-sizing opportunity. Let's fire an alert when an app consistently uses low CPU/memory while a high allocation is reserved — i.e., a right-sizing candidate. Use the *Percentage metrics with a long evaluation window so short idle dips don’t trigger it.
# Action group (email) to receive the alert
az monitor action-group create \
--name aca-rightsizing-ag \
--resource-group "$RG" \
--short-name acaRS \
--action email owner you@example.com
# Alert: CPU utilization stays below 20% of the reservation for a long time
az monitor metrics alert create \
--name "aca-cpu-underutilized" \
--resource-group "$RG" \
--scopes "$APP_ID" \
--description "App reserves CPU it does not use (right-size candidate)" \
--condition "avg CpuPercentage < 20" \
--window-size 1h \
--evaluation-frequency 15m \
--severity 3 \
--action aca-rightsizing-ag
# Alert: memory utilization stays below 30% of the reservation
az monitor metrics alert create \
--name "aca-mem-underutilized" \
--resource-group "$RG" \
--scopes "$APP_ID" \
--description "App reserves memory it does not use (right-size candidate)" \
--condition "avg MemoryPercentage < 30" \
--window-size 1h \
--evaluation-frequency 15m \
--severity 3 \
--action aca-rightsizing-ag
CpuPercentage and MemoryPercentage measure utilization relative to configured resource limits, making them the preferred metrics for rightsizing decisions. Sustained low percentages often indicate excess allocated capacity, while sustained high percentages (for example, CPU > 85% for 15 minutes) can signal resource saturation and the need for scaling.
This sizing exercise is very similar to rightsizing CPU and memory requests in a Kubernetes deployment. In both cases, the goal is to align provisioned resources with actual workload demand, reducing wasted capacity.
Window size, daily peaks, and the 24-hour maximum
A metric alert collapses the entire window into a single aggregated value before comparing it to the threshold — so the aggregation you choose matters as much as the window length:
- Yes, --window-size 24h is valid — 24 hours (1 day) is the maximum window for a metric alert. --evaluation-frequency must be ≤ the window (for a 24h window, use 1h).
- But avg over 24h hides daily peaks. An app that idles most of the day yet spikes hard for an hour still shows a low 24h average, so avg CpuPercentage < 20 would flag it as a right-size candidate even though you need that reservation for the peak.
- If you care about peaks, alert on the daily Maximum instead. Then the alert fires only when even the busiest moment of the day is under-utilized — a much safer “shrink it” signal:
# Fires only when the day's PEAK CPU is still low (respects intraday spikes)
az monitor metrics alert create \
--name "aca-cpu-underutilized-24h" \
--resource-group "$RG" \
--scopes "$APP_ID" \
--description "Even the daily peak CPU stays low (safe right-size candidate)" \
--condition "max CpuPercentage < 50" \
--window-size 24h \
--evaluation-frequency 1h \
--severity 3 \
--action aca-rightsizing-ag
For patterns longer than 24 h (e.g. “low for 7 days straight” or percentiles across weeks), a single metric alert can’t span that window — use a log (scheduled query) alert over the metrics in Log Analytics, or review a workbook/Grafana trend instead.
9.1 Hands-on: fire an under-utilization alert from a 12-hour idle app
This walkthrough proves the whole loop end to end: deploy a deliberately over-allocated, idle app, attach a CPU and memory < 50 % alert with a 12-hour window, and receive an email when it fires. Because the app never does any work, both utilization metrics sit near 0 %, so the alert is expected to fire once the first full window has been evaluated.
9.1.1 — Deploy a new idle app (over-allocated on purpose)
export IDLE_APP="idle-alert-demo"
# Idle helloworld app with a generous 1 vCPU / 2 GiB reservation it will never use
az containerapp create \
--name "$IDLE_APP" \
--resource-group "$RG" \
--environment "$ENV" \
--workload-profile-name "$PROFILE" \
--image mcr.microsoft.com/azuredocs/containerapps-helloworld:latest \
--target-port 80 --ingress external \
--cpu 1.0 --memory 2.0Gi \
--min-replicas 1 --max-replicas 1 \
--query properties.configuration.ingress.fqdn -o tsv
# Capture the resource ID for the alert scope
export IDLE_APP_ID=$(az containerapp show -n "$IDLE_APP" -g "$RG" --query id -o tsv)
echo "$IDLE_APP_ID"
Do not send traffic to this app. Leaving it idle keeps CpuPercentage/MemoryPercentage near 0 %, which is exactly the under-utilization condition the alert detects.
9.1.2 — Create the email action group
az monitor action-group create \
--name aca-idle-alert-ag \
--resource-group "$RG" \
--short-name acaIdle \
--action email owner you@example.com
The recipient receives a one-time Azure confirmation email the first time this action group is created — no need to act on it for the alert to work.
9.1.3 — Create a combined CPU + memory alert (12-hour window)
Passing two --condition flags creates a single alert with both criteria AND-ed together — it fires only when the app is under-utilized on both CPU and memory over the window. A 12-hour window is valid (≤ 24 h max); the evaluation frequency must be ≤ the window, so we use 1h.
az monitor metrics alert create \
--name "aca-idle-underutilized-12h" \
--resource-group "$RG" \
--scopes "$IDLE_APP_ID" \
--description "Idle app: CPU and memory PEAK both under 50% of reservation for 12h (right-size candidate)" \
--condition "max CpuPercentage < 50" \
--condition "max MemoryPercentage < 50" \
--window-size 12h \
--evaluation-frequency 1h \
--severity 3 \
--action aca-idle-alert-ag
Why max, not avg: using the Maximum aggregation means the alert fires only when even the peak CPU and memory over the entire 12-hour window stay under 50 %. This is the stricter, safer signal — it surfaces apps that are truly unused or barely utilized, and won’t be fooled by a low average that hides a real spike. An avg would flag an app whose mean is low even if it needs the reservation for occasional bursts.
9.1.4 — Check for the fired alert
Review Azure portal for alerts fired.
Keep the app idle and check Azure Monitor > Alerts. The rule evaluates every hour over a rolling 12-hour window. Check your email for the notification, but don't expect it at a fixed time after creating the app.
9.1.5 — Clean up the test resources
az monitor metrics alert delete --name "aca-idle-underutilized-12h" -g "$RG"
az monitor action-group delete --name aca-idle-alert-ag -g "$RG"
az containerapp delete --name "$IDLE_APP" -g "$RG" --yes
10. Step 7 — Dashboards
Three ways to persist the view, from quickest to richest:
a) Pin to an Azure portal dashboard
Container app → Metrics → build a chart (CPU%, Mem%, Replicas) → Save to dashboard → Pin. Fast, shareable, no extra resources.
b) Azure Monitor Workbook
Azure Monitor → Workbooks → New → add Metric parameters/tiles scoped to your app(s). Workbooks support multiple apps, parameters (revision/replica), and a mix of metrics + Log Analytics (KQL) tiles in one report. Best for a reusable “fleet right-sizing” report across many apps.
c) Grafana dashboard (from Step 5)
Best for real-time NOC-style views and mixing ACA metrics with other Azure or Prometheus sources. Add rows for Reservation vs. Utilization, Replica count over time, and Reserved cores (CoresQuotaUsed) to visualize capacity.
Recommended dashboard rows for a cost-optimization: 1. CpuPercentage & MemoryPercentage (utilization vs. reservation) 2. CoresQuotaUsed / TotalCoresQuotaUsed (reserved cores) 3. Replicas and environment NodeCount (scale → billed nodes) 4. Requests (demand driving the scale)
11. Active vs. inactive revisions and compute usage
Every containerapp update that changes the container (image, env vars, resources, scale) creates a new revision. Understanding revision state is essential for compute/cost control.
The key point the diagram makes — and that a simple “active/inactive” view misses — is that the revision mode decides how many revisions stay active at once, and every active revision with running replicas reserves its full cpu/memory on the node. So the mode is what actually drives the total reserved capacity:
|
Revision mode |
Revision state |
Replicas reserve node compute? |
Billed? |
|
Single (default) |
New/current active revision |
✅ Yes — its replicas reserve cpu/memory |
✅ Yes - Indirectly billed as capacity is reserved and not available for other workloads |
|
Single (default) |
Previous revision — auto-deactivated on each update (old/new revisions can overlap during deployment of the new revision from a capacity usage standpoint) |
❌ No — replicas are torn down |
❌ No |
|
Multiple |
Every active revision (old + new) with min-replicas > 0 |
✅ Yes — each revision’s replicas reserve cpu/memory in parallel |
✅ Yes — summed across all active revisions. Indirectly billed as stated before |
|
Multiple |
Active revision scaled to 0 (Consumption, no traffic) |
⚠️ Only while replicas are up |
⚠️ Only while replicas are up. Indirectly billed as stated before |
|
Single or Multiple |
Inactive (deactivated) |
❌ No — config only, no replicas |
❌ No |
Active vs inactive revisions and their compute cost
Why it matters for compute:
In multiple-revision mode, old active revisions with min-replicas > 0 keep reserving compute on the node even after you ship a new revision. Two active revisions at 1 replica × 1 vCPU = 2 reserved vCPU on the node, not 1 — the reservation stacks per active revision. Ship a third and you’re at 3, and so on until they’re deactivated.
Single-revision mode (default) automatically deactivates the previous revision on every update, tearing down its replicas and releasing its reserved compute — so the node only ever holds the current revision’s reservation except during deployment when it holds both for a short period of time.
Inactive revisions cost nothing in either mode — they hold configuration only, no replicas, so they reserve nothing on the node.
Inspect and clean up:
# List revisions and their state / replica counts
az containerapp revision list -n "$APP" -g "$RG" \
--query "[].{name:name, active:properties.active, replicas:properties.replicas, traffic:properties.trafficWeight}" -o table
# Deactivate an old active revision to release reserved compute
az containerapp revision deactivate -n "$APP" -g "$RG" --revision <old-revision-name>
# Prefer single-revision mode unless you need blue/green or traffic splitting
az containerapp revision set-mode -n "$APP" -g "$RG" --mode single
Cost rule of thumb: reserved cores = Σ (active revisions × their replicas × cpu allocation). Deactivating stale revisions and using single-revision mode are the cheapest wins.
11.1 Demo — watch multiple-revision mode stack reserved compute
This demo makes the “old active revisions keep reserving compute” story concrete. You’ll deploy a dedicated app, flip it to multiple-revision mode, ship two more revisions, and watch TotalCoresQuotaUsed climb with each active revision — even though only one revision serves traffic. The final step restores the app to its original single-revision, single-revision state and scales it to 0, so you can restart it from the portal and rerun the demo without recreating it.
Why a dedicated profile: on d4prof the reserved cores are pinned to the node and clearly visible via TotalCoresQuotaUsed/NodeCount. CoresQuotaUsed is revision-scoped (dimension: revisionName), while TotalCoresQuotaUsed sums all active revisions for the app. Each active revision at 1 vCPU adds a full reserved core.
11.1.1 — Create the demo app (single revision, 1 replica)
export REV_APP="revision-demo"
az containerapp create \
--name "$REV_APP" \
--resource-group "$RG" \
--environment "$ENV" \
--workload-profile-name "$PROFILE" \
--image mcr.microsoft.com/azuredocs/containerapps-helloworld:latest \
--revision-suffix v1 \
--target-port 80 --ingress external \
--cpu 1.0 --memory 2.0Gi \
--min-replicas 1 --max-replicas 1 \
--query properties.configuration.ingress.fqdn -o tsv
export REV_APP_ID=$(az containerapp show -n "$REV_APP" -g "$RG" --query id -o tsv)
Baseline the reserved cores — expect 1 (one active revision × 1 replica × 1 vCPU):
az monitor metrics list --resource "$REV_APP_ID" \
--metric TotalCoresQuotaUsed --aggregation Maximum --interval PT1M -o table
11.1.2 — Switch to multiple-revision mode
In this mode ACA no longer deactivates the previous revision when you ship a new one — old revisions stay active and keep their reserved compute:
az containerapp revision set-mode -n "$REV_APP" -g "$RG" --mode multiple
11.1.3 — Ship two more revisions (each stays active and reserves compute)
Each update with a new suffix creates a new active revision. Because we’re in multiple-revision mode with min-replicas 1, all three revisions run replicas in parallel:
# Revision 2
az containerapp update -n "$REV_APP" -g "$RG" \
--revision-suffix v2 --set-env-vars "DEMO_VER=2"
# Revision 3
az containerapp update -n "$REV_APP" -g "$RG" \
--revision-suffix v3 --set-env-vars "DEMO_VER=3"
Confirm all three revisions are active, each with a running replica:
az containerapp revision list -n "$REV_APP" -g "$RG" \
--query "[].{name:name, active:properties.active, replicas:properties.replicas, traffic:properties.trafficWeight}" -o table
Now watch reserved cores — expect 3 (three active revisions × 1 replica × 1 vCPU), even though only the latest revision receives traffic:
# Reserved cores — should now be ~3, up from the baseline of 1
az monitor metrics list --resource "$REV_APP_ID" \
--metric TotalCoresQuotaUsed --aggregation Maximum --interval PT1M -o table
# Node count may rise too, since 3 reserved vCPU still fits one D4 but headroom shrinks
az monitor metrics list --resource "$ENV_ID" \
--metric NodeCount --aggregation Maximum --interval PT1M -o table
Takeaway: multiple-revision mode stacks the reservation — every stale active revision is a full, ongoing reservation. If you use this mode for blue/green or traffic splitting, you must deactivate old revisions once they’re no longer needed, or reserved cores grow with every deploy.
11.1.4 — Deactivate the stale revisions to release compute
# Deactivate v1 and v2, keeping only the latest (v3) active
az containerapp revision deactivate -n "$REV_APP" -g "$RG" \
--revision "$REV_APP--v1"
az containerapp revision deactivate -n "$REV_APP" -g "$RG" \
--revision "$REV_APP--v2"
# Reserved cores should drop back toward 1
az monitor metrics list --resource "$REV_APP_ID" \
--metric TotalCoresQuotaUsed --aggregation Maximum --interval PT1M -o table
The app-level metrics show total reserved cores rising from 1 to 3 while revisions v2 and v3 are active, then returning to 1 after v1 and v2 are deactivated:
Total reserved cores rise with three active revisions and fall after stale revisions are deactivated
The environment-level metric shows the dedicated workload profile scaling from one node to two to accommodate the additional reservations, then returning to one node after deactivation:
Workload profile node count rises from one to two during the multi-revision demo and returns to one
11.1.5 — Restore the app to its original state (idle, ready to restart)
Bring the app back to single-revision mode and scale it to 0 so it reserves nothing while parked. Next time you can start it from the portal (set min-replicas back to 1, or send traffic) and rerun the demo from Step 11.1.2 — no recreate needed:
# Back to single-revision mode (auto-deactivates all but the latest revision)
az containerapp revision set-mode -n "$REV_APP" -g "$RG" --mode single
# Park at 0 replicas so the app reserves no compute while idle
az containerapp update -n "$REV_APP" -g "$RG" \
--min-replicas 0 --max-replicas 1
# Confirm: one active revision, mode=Single, and reserved cores heading to 0
az containerapp show -n "$REV_APP" -g "$RG" \
--query "{mode:properties.configuration.activeRevisionsMode, minReplicas:properties.template.scale.minReplicas}" -o table
az monitor metrics list --resource "$REV_APP_ID" \
--metric TotalCoresQuotaUsed --aggregation Maximum --interval PT1M -o table
Restart from the portal: open the app → Revisions and replicas (or Scale) and set min replicas back to 1, or just hit the ingress URL to wake it. The app keeps its single active revision, so you resume the demo instantly by re-running Step 11.1.2 onward.
12. Scaling & scale rules — right-size for cost
ACA scales horizontally (more/fewer replicas) using KEDA. Because every replica reserves its full cpu/memory (Section 1), the scale configuration is a direct cost lever: too-high min-replicas pays for idle capacity; the right rule keeps replicas close to real demand and lets idle apps fall to zero.
12.1 Default scaling (no rule provided)
If you create an app and don’t specify any scale rule, ACA applies a default HTTP rule:
|
Setting |
Default |
|
Trigger |
HTTP |
|
min-replicas |
0 |
|
max-replicas |
10 |
|
concurrentRequests |
10 (one extra replica per 10 concurrent requests) |
Default HTTP scaling from 0 to 10 replicas
⚠️ If you disable ingress and set neither a min-replicas nor a custom rule, the app scales to 0 with no trigger to start it back up. For non-HTTP apps always set a scale rule or min-replicas ≥ 1.
12.2 Right-size scaling to reduce cost
|
Lever |
Cheap setting |
Why it saves |
|
min-replicas |
0 (or 1 only if cold start is unacceptable) |
Idle replicas reserve capacity; zero replicas don't reserve any capacity |
|
max-replicas |
Cap to real peak |
Prevents runaway scale-out cost during traffic spikes/abuse |
|
Rule threshold |
Match real concurrency/utilization |
Fewer, busier replicas = less reserved compute |
|
Profile |
min-nodes 0 for workload-profile along with min-replicas 0 for ACA app deployed |
Dedicated bills the node even at 0 replicas (see note) |
Dedicated vs. scale-to-zero: On a Dedicated profile configured with min-nodes 1, scaling all app replicas to zero does not stop compute billing because one profile node remains running. If the profile is configured with min-nodes 0, it can scale to zero nodes when no assigned workloads require capacity; at zero nodes, workload-profile instance charges stop, although the Dedicated plan management fee and charges for related Azure services may remain. It's important to understand the cold start - with both min-nodes 0 and min-replicas 0, the first request may experience a longer cold start while ACA provisions a Dedicated node and starts an app replica.
12.3 Demo prerequisites — traffic & work generators (run in ACA, not your laptop)
Build two small images into the existing srinmantest registry so load is generated inside Azure, independent of your laptop.
export RG="aca-obs-rg"
export ENV="aca-obs-env"
export ACR="srinmantest"
export LOC="westus"
- a) CPU-work web app (/ is cheap; /burn?ms=200 busy-loops CPU — drives HTTP and CPU demos):
mkdir -p cpuwork && cd cpuwork
cat > app.py <<'EOF'
from flask import Flask, request
import time
app = Flask(__name__)
@app.get("/")
def home():
return "ok\n"
@app.get("/burn")
def burn():
ms = int(request.args.get("ms", "200"))
end = time.time() + ms / 1000.0
x = 0
while time.time() < end:
x += 1
return f"burned {ms}ms\n"
EOF
cat > Dockerfile <<'EOF'
FROM python:3.12-slim
RUN pip install --no-cache-dir flask gunicorn
WORKDIR /app
COPY app.py .
CMD ["gunicorn","-b","0.0.0.0:80","-w","4","--threads","8","app:app"]
EOF
az acr build --registry "$ACR" --image cpuwork:v1 .
cd ..
- b) Load generator (an always-on container that hammers a target URL from inside the environment):
mkdir -p loadgen && cd loadgen
cat > Dockerfile <<'EOF'
FROM alpine:3.20
RUN apk add --no-cache curl
ENV TARGET_URL="http://localhost/" CONCURRENCY=50
# POSIX sh, no bash/seq dependency; prints a heartbeat every 20 batches
CMD ["/bin/sh","-c","echo loadgen start -> $TARGET_URL x$CONCURRENCY; n=0; while true; do i=0; while [ $i -lt $CONCURRENCY ]; do curl -s -o /dev/null \"$TARGET_URL\" & i=$((i+1)); done; wait; n=$((n+1)); if [ $((n % 20)) -eq 0 ]; then echo \"completed $n batches\"; fi; done"]
EOF
az acr build --registry "$ACR" --image loadgen:v1 .
cd ..
The demo create commands below include --registry-server $ACR.azurecr.io — this is required so ACA wires up pull credentials for srinmantest; without it the create fails with an image-pull/authentication error.
12.4 Demo — CPU / memory scaling (70% threshold)
CPU/memory rules add replicas when average utilization across replicas crosses a target. They cannot scale to zero (the metric needs a live replica), so min-replicas ≥ 1.
CPU/memory scaling at a 70% utilization threshold
# CPU-bound app on the dedicated D4 profile, scale out at 70% CPU
az containerapp create \
--name cpu-scale-demo \
--resource-group "$RG" \
--environment "$ENV" \
--workload-profile-name d4prof \
--image "$ACR.azurecr.io/cpuwork:v1" \
--registry-server "$ACR.azurecr.io" \
--target-port 80 --ingress external \
--cpu 0.5 --memory 1.0Gi \
--min-replicas 1 --max-replicas 5 \
--scale-rule-name cpu70 \
--scale-rule-type cpu \
--scale-rule-metadata "type=Utilization" "value=70" \
--query properties.configuration.ingress.fqdn -o tsv
Drive CPU load from inside ACA and watch replicas climb:
CPU_FQDN=$(az containerapp show -n cpu-scale-demo -g "$RG" --query properties.configuration.ingress.fqdn -o tsv)
# Load generator app hitting the /burn endpoint (CPU-heavy)
az containerapp create \
--name loadgen-cpu \
--resource-group "$RG" \
--environment "$ENV" \
--workload-profile-name Consumption \
--image "$ACR.azurecr.io/loadgen:v1" \
--registry-server "$ACR.azurecr.io" \
--min-replicas 1 --max-replicas 1 \
--env-vars "TARGET_URL=https://$CPU_FQDN/burn?ms=250" "CONCURRENCY=80"
# Watch replicas and CPU% climb, then settle after each scale-out
watch -n 15 'az monitor metrics list --resource $(az containerapp show -n cpu-scale-demo -g aca-obs-rg --query id -o tsv) --metric Replicas CpuPercentage --aggregation Maximum Average --interval PT1M -o table | tail -n 8'
Deactivate the load generator revision and the target app scales back toward min-replicas 1:
LOADGEN_REVISION=$(az containerapp revision list -n loadgen-cpu -g "$RG" \
--query "[?properties.active].name | [0]" -o tsv)
az containerapp revision deactivate -n loadgen-cpu -g "$RG" \
--revision "$LOADGEN_REVISION"
Memory scaling is identical with --scale-rule-type memory --scale-rule-metadata "type=Utilization" "value=70".
This is a simple demo for explaining how ACA scales replicas based on cpu/mem. It's important to note that ACA also several additional scaling options out of the box.
12.7 Scaling cost takeaways
- ACA app min-replicas is a reserved capacity. Set it to 1+ only when cold-start latency is unacceptable; use a cron rule to keep replicas warm only during business hours instead of 24/7.
- Cap ACA app max-replicas to your real peak so a traffic spike (or abuse) can’t multiply reserved compute without bound.
- CPU/memory rules can’t scale to zero — pair with HTTP/event rules if you also need scale-to-0.
- Event-driven + Consumption is the cheapest pattern for bursty/async work: pay only while messages are in flight.
- Dedicated profiles min and max nodes - Set the minimum to the desired pre-provisioned baseline, or 0 to allow scale-to-zero, and set the maximum as the capacity and cost ceiling. ACA automatically adds or removes nodes within those bounds based on the aggregate CPU and memory required by app and job replicas assigned to the profile.
13. Reliability & cost — resiliency without overspending
Resiliency and cost are the same decision — you buy availability with standing capacity, so decide how much failure each part of the system must absorb, then pay exactly for that. Two references frame this:
- Reliability in Azure Container Apps
- Two zones or three? A design framework for zone-resilient Azure workloads
13.1 The key ACA insight: zone redundancy is free — replicas are the cost
Zone-redundant ACA across three availability zones
- Zone redundancy costs nothing extra. You pay the same vCPU/memory/request rates whether it’s on or off — ACA just spreads your replicas across the region’s availability zones for you (service-managed, active-active; ingress round-robins across all healthy replicas, any zone).
-
The cost consideration is the node capacity reserved by your minimum replicas. Keeping enough replicas warm for the *surviving* zones to carry full load can require additional dedicated profile nodes, increasing node-based billing.
- Enable it at environment creation — it’s immutable. You can’t turn zone redundancy on (or off) later; to change it you create a new environment and redeploy. Requires a VNet with a /27+ subnet (workload profiles) or /23+ (consumption-only).
- Stateless by design: ephemeral storage in a lost zone is gone. Keep state in ZRS Azure Files, Cosmos DB, or Azure SQL (each does its own cross-zone replication).
- Failover after a zone loss is transparent and recovery is driven by your health-probe settings.
Because zone redundancy is free and can only be set at create time, turn it on for every production environment by default.
13.2 Right-size minimum replicas for a zone loss (the real resiliency cost)
Your min-replicas defines the guaranteed capacity spread across zones. Set it so the replicas remaining after a zone loss still meet peak demand.
- Set min-replicas ≥ 2 so replicas are actually distributed across zones (one replica can’t span zones. Replica is an equivalent of a pod in Kubernetes).
- Every replica reserves its full cpu/memory (Section 1), so maintaining failover headroom may require additional workload-profile node capacity that remains provisioned 24x7. Trade it off against cold-start tolerance: fewer warm replicas is cheaper but you wait for the platform to start replacements in healthy zones during an outage. If you can’t tolerate any dip below your minimum, over-provision.
- Set cpu/memory requests deliberately — the scheduler uses them to place replicas evenly across zones; underspecified resources cause uneven distribution.
- The ACA SLA is based on the scale rules you set — a higher, well-distributed min-replicas is both your availability guarantee and your cost.
13.3 Decide resiliency per component, not per workload
The blog’s core point: don’t stamp “three zones everywhere.” Ask of each component — how many zones does it need to survive the loss of one?
|
Component |
Zone pattern |
ACA cost stance |
|
ACA app (stateless compute) |
Two or three zones both meet a single-zone objective |
Let Azure’s service-managed zone redundancy handle it (free). Spend on min-replicas, not on rolling your own zonal design |
|
Quorum / consensus / leader-election stores |
Three zones (or a witness) — a third failure domain, not just a third replica |
Runs outside ACA; budget for 3-zone here |
|
Other stateful deps (DB, cache, files) |
Two-zone, three-zone, or service-managed by RTO/RPO |
Prefer service-managed ZR where it meets requirements |
- Prefer service-managed zone redundancy wherever it fits — the ACA app is the textbook case.
- Replica count ≠ replica placement. Three replicas across two zones can still lose quorum when the majority-holding zone fails. ACA handles placement for you; your stateful dependencies may not.
- The cost conversation comes last. Decide the objective, then optimize — and use savings plans / reservations for the always-on baseline (min nodes with workload profiles), which is predictable spend.
13.4 Region failure is a separate, much larger cost tier
ACA is a single-region service — a region outage takes your environment down. Surviving it is a different (and far more expensive) exercise than zone resiliency:
Resiliency-cost ladder: transient faults, zone redundancy, multi-region
- Use zone redundancy with more replicas to improve the availability of your application within a region
- Duplicate the environment in a second region (each needs its own VNet/subnet).
- Geo-replicate the container registry (ACR geo-replication) so images pull in every region.
- Front the regions with Azure Front Door or Traffic Manager for failover or bring your own traffic manager such as F5 GTM for private network fail-over
- Replicate your data across regions (the app tier is stateless; the data tier is the hard part).
Climb the ladder only as far as the business needs: transient-fault handling (retries/health probes — free) → zone redundancy (free infra, node capacity for warm min-replicas) → multi-region (roughly double the footprint + global routing + data replication — reserve for mission-critical/DR).
13.5 Reliability-cost takeaways
- Turn zone redundancy on for prod by default — it’s free and can only be set at environment creation.
- min-replicas drive your resiliency footprint. Size them for surviving-zone demand, keep them ≥ 2 for zone redundancy, and remember the reserved CPU/memory requirements can translate into dedicated workload profile node capacity that remains provisioned 24/7.
- A 3-zone spread needs less standing headroom (1.5N) than 2-zone (2N) for the same post-failure capacity.
- Let Azure manage ACA zone redundancy; spend the resiliency budget on stateful dependencies and min-replicas, not a hand-rolled zonal design.
- Multi-region only when a region outage is unacceptable — it roughly doubles cost and adds global routing + data replication.
- Use savings plans / reservations for the always-on baseline, and set cpu/memory requests so the scheduler distributes replicas cleanly across zones.
14. Cost-optimization takeaways
- Reservation, not requests. In ACA cpu/memory is reserved. There is no low-request/high-limit oversubscription like AKS — so right-size the allocation to real usage.
- Watch the *Percentage metrics. Persistent low CpuPercentage/MemoryPercentage against a large allocation = pre-paid waste. Alert on it (Step 6) , investigate and shrink --cpu/--memory where it makes sense.
- Replicas multiply reserved compute. Tune min-replicas/scale rules; every replica consumes full capacity based on cpu and mem.
- On dedicated profiles you pay per node. Keep min-nodes tight; pack right-sized replicas with right sized nodes so you don’t provision extra nodes with idle cores.
- Kill stale revisions. Use single-revision mode; deactivate old active revisions to free reserved cores.
- Grafana needs no Azure Monitor workspace for built-in ACA metrics — only custom Prometheus metrics require the workspace + scrape pipeline.
15. Cleanup
az group delete --name "$RG" --yes --no-wait