serverless
227 TopicsRun an Ollaya decision model on Azure Container Apps
Many AI calls are not really conversations. A support message needs an intent. A workflow needs a route. A policy check needs a yes or no. An agent needs to select its next action from a known list. These are bounded decisions, but they are often sent to a general-purpose LLM. The LLM reads the request, generates an answer token by token, and then the application validates and parses the response. That is useful when the task needs reasoning or language generation. It is more machinery than necessary when the only valid answer is one of 60 labels. On September 15, 2026, TypeSafe released Jev, its first public System One Model, in early access. Jev has helped bring attention to models built specifically for decisions: structured state goes in, and typed choices with probabilities come out. Ollaya approaches the same problem from a self-hosted direction. It is a model runtime that can serve open decision models, including the winnow:e4b model used here. Jev and Ollaya are not connected products. Jev is a hosted model from TypeSafe; Ollaya provides a way to run decision models in infrastructure you control. They are related by the kind of work they target. What decision models are good at A decision model scores predefined options rather than generating open-ended text: Decision model Traditional LLM Selects from allowed choices Generates text Returns a probability for each decision Usually returns one generated answer Has a bounded, typed output Needs schema constraints and validation Supports confidence thresholds Often needs a separate confidence strategy Fits classification, routing, scoring, policy, and prioritization Fits generation, summarization, coding, and open-ended reasoning That narrower interface creates several practical benefits: Predictable outputs: the application receives an allowed value rather than text that must be repaired or parsed. Lower decision latency: there is no autoregressive output sequence to generate. Token savings: a decision can be returned without generating output tokens. Useful uncertainty: calibrated probabilities let the application act, reject, or escalate. Smaller infrastructure: a specialized model can fit on hardware that would be modest for a general LLM. Decision models are not replacements for every LLM call. They make sense when the possible outcomes are known before the request arrives. An LLM remains the better tool when the output itself is language, code, or open-ended reasoning. Why Azure Container Apps serverless GPU Azure Container Apps serverless GPUs make self-hosted inference feel much closer to consuming a managed API. You bring the container and model; Azure manages the underlying GPU infrastructure. The Consumption GPU profile provides: NVIDIA T4 or A100 GPUs without managing GPU nodes or a Kubernetes cluster. Automatic scaling with the option to scale to zero. Per-second GPU billing while replicas are running. Container Apps networking, identity, ingress, logging, and revision management. A private inference path where the model and request data stay inside your Azure environment. That last point matters for data-sensitive decisions. In the Ollaya-only configuration, text is sent to an internal Ollaya endpoint rather than an external model API. The model, API, persistent cache, and operational controls remain in the application's Azure environment. Serverless does not remove the need to think about cold starts. A model still needs to be loaded into VRAM. For low or sporadic traffic, scaling to zero can avoid idle GPU cost. For latency-sensitive traffic, a warm minimum replica avoids making a user wait for the model to load. Azure Files can retain the model layers across revisions so a new deployment does not download the full model again. The T4 is also an important part of this experiment. It is not the largest GPU available, but the goal is not to run the largest model. The goal is to match the hardware to a model designed for the task. A focused decision model on a T4 can compete with a hosted API when the workload is a bounded classification rather than text generation. The experiment I deployed winnow:e4b through Ollaya on an Azure Container Apps Consumption-GPU-NC8as-T4 workload profile. An authenticated API exposed the classifier while Ollaya remained on internal-only ingress. A second path sent the same requests to GPT-5.4 Nano through Azure OpenAI using managed identity. The evaluation used the complete 2,974-record test partition from Amazon MASSIVE 1.1. It contains 18 scenarios and 60 intents. Both providers received the same records, labels, and descriptions. The benchmark measured latency at concurrency 1 and throughput at concurrency 8. Providers and modes ran sequentially so one measurement did not load the service used by another. This was not a direct benchmark of Jev; it tested the same decision-model pattern with an open model that could run inside the Azure environment. What the results showed The benchmark ran on September 30, 2026. Provider Successful Intent accuracy Macro-F1 p50 p95 Winnow on T4 2,974 / 2,974 75.59% 76.40% 946 ms 960 ms GPT-5.4 Nano 2,974 / 2,974 79.12% 78.07% 1,369 ms 2,489 ms Nano led intent accuracy by 3.53 percentage points. Winnow was 30.9% faster at p50 and 61.4% faster at p95. This is the useful T4 result: a smaller GPU running the right specialized model matched and beat the hosted endpoint on response latency, though not on accuracy. At concurrency 8, Nano delivered 1.367 requests per second compared with Winnow's 1.120. Nano also had 25 requests fail after eight retries because the deployment exceeded its token-rate limit. Winnow completed all 2,974 requests, but its single loaded runner serialized work and increased queue time. The token comparison shows what the decision-only path removes: Provider, both benchmark passes Input tokens Cached input tokens Output tokens Winnow on T4 9,252,872 0 0 GPT-5.4 Nano 8,501,200 7,564,800 122,759 Winnow returned every decision without generating output tokens. Nano generated 122,759 output tokens across the latency and throughput passes. The rows do not represent equivalent billing models: Winnow consumes self-hosted GPU time, while Nano is metered by hosted token usage. The comparison isolates the generated tokens that a bounded decision did not need. Winnow also returned calibrated probabilities. At a 0.90 threshold, it accepted 53.73% of the records and was correct on 94.43% of those accepted decisions. An application could handle that high-confidence group locally and send only the uncertain remainder to an LLM or human reviewer. In brief A decision model is useful when software needs a bounded answer rather than generated language. In this experiment, Winnow on a serverless T4 traded some accuracy for lower latency, zero generated output tokens, private inference, and an explicit confidence signal. The practical design is often a combination: use the decision model for fast, high-confidence choices and reserve LLM calls for uncertain or open-ended work. Try it with the template The Azure Developer CLI template packages the Container Apps environment, T4 workload profile, authenticated API, internal Ollaya service, persistent model cache, GPU readiness checks, and benchmark. It supports two deployment modes: Mode What it deploys ollaya-only Private Winnow inference on a serverless T4 plus the authenticated API full The Ollaya deployment, GPT-5.4 Nano, and the comparison benchmark Read the deployment guide Inspect every benchmark prediction and retry Review the shared MASSIVE taxonomy120Views0likes0CommentsMeet the Hosted Skills Canvas: Build, Run, and Debug in GitHub Copilot
Build, run, and debug event-driven AI apps right inside GitHub Copilot. The new Azure Functions Hosted Skills canvas brings instructions, triggers, and live results into one workspace—so you can spend less time switching tools and more time building.
310Views1like0CommentsAzure Container Apps Sandboxes, Now Generally Available
Agents Act. You Set the Limits. Your platform needs to run untrusted code. Agents choose actions at runtime, from installing packages to calling APIs. In a multi-tenant service, you have to support both without letting one customer's code reach another customer's data. That requires an execution environment you control for each tenant, session, or task. Azure Container Apps Sandboxes is that execution environment, as a service. Each sandbox is a hardware-isolated microVM with its own Linux kernel. It starts in under a second, and you decide per sandbox what it can reach, how large it is, and how long it lives. For each user, controlling which credentials their agent can use and what it can reach is key. Untrusted code must stay isolated from other workloads and the host kernel. Preserving a session state without giving up scale to zero or instant startup is important, as is visibility into what each agent did. Let's unpack these one at a time. Control What Your Agents Can Reach The first question about an agent is not what it can do, but what it can reach. Per-sandbox egress policies answer that: an external proxy evaluates every outbound request against rules you set - by host, domain pattern, or CIDR. You can start with a default 'Deny' action and allow only the endpoints the task needs, so a prompt injection or a compromised dependency has nowhere to go. Network Audit shows you what was allowed and what was denied. Approved endpoints usually need credentials, and handing an API key to an agent means the key can be logged, echoed into a response, or carried off somewhere you did not intend. Transform rules can inject authentication headers outside the sandbox. The agent sends a request with no secret in it, the proxy adds the credential outside the sandbox, and the call goes through. The agent gets access to the service without ever getting the API key. When a static allowlist cannot express the rule you need, an egress webhook hands each request to your own service before it leaves. You can build a service that reviews all outbound calls, approves some and rejects others. Good results depend on the agent reaching the right data, and the data that matters usually sits on your private network. Outbound, VNet integration places the sandbox group on a dedicated subnet, so agents can reach internal APIs, databases, and services behind private endpoints. Egress rules are still enforced: routing is chosen per rule, so a call to an internal database is filtered, transformed, and recorded in Network Audit exactly like a call to the open internet. Inbound, a Private Endpoint brings the sandbox service into your VNet, so your own applications reach into your sandboxes without crossing the public internet. Run Untrusted Code Without Sharing a Kernel Each sandbox runs in a hardware-isolated microVM with its own Linux kernel and virtual hardware, with memory separation enforced through CPU virtualization. In contrast, typical container runtimes isolate processes while sharing the host kernel. In a sandbox, code calls into its own guest kernel, so the blast radius of a kernel exploit is one sandbox. You get that boundary with ACA Sandboxes. Inside that boundary, the filesystem is yours to choose. The quickest way to start is by creating a sandbox based on a platform-provided disk image: Public image Ready for ubuntu General-purpose Linux execution and development tools nginx Running a web server copilot, claude GitHub Copilot CLI or Claude Code workflows azure-dev Azure development with CLI tools and multiple language runtimes python-3.12-code-interpreter Isolated Python execution through REST and MCP python-3.11 to python-3.14 Python workloads node-22, node-24 Node.js workloads dotnet-8 to dotnet-10 .NET workloads php-8.3, php-8.4 PHP-FPM workloads Availability and versions change over time, please check the portal's catalog for the current list. Images are provided as is. For a custom environment with your own code and dependencies, you can bring a container image from a public or private registry. The platform converts it into an optimized, bootable disk image containing your application, agent harness, runtimes, and toolchain. For a private registry, you authenticate with registry credentials or a managed identity for Azure Container Registry. You can also start from a sandbox you already have running. A disk snapshot captures the filesystem as a new disk image, while a memory snapshot captures disk and memory together, so a sandbox created from it resumes where the original left off. Installed dependencies, or a cloned repo, are set up once, and every sandbox created from the snapshot starts with them. Give Every Task Its Own Sandbox A sandbox starts in under a second, and that is why it is practical to give every task its own machine. When a machine takes minutes to come up, everything ends up sharing that same space. When it starts instantly, each task, user, or tool call runs in isolation, and gets deleted when the work is done. Thousands at a time. Sandboxes are sized when created by selecting a resource tier. These range from 0.25 vCPU with 0.5 GB of memory and a 5 GB disk, enough for a short script or a quick evaluation, to 4 vCPU with 8 GB of memory and an 80 GB disk for compilation and heavier analysis, with multiple steps in between. Programmatically you can go up to 16 vCPU with 32 GB of memory and a 320 GB disk for the most demanding work. Sandboxes do not have to be discarded by hand. Lifecycle policies stop a sandbox once it goes idle, capturing the disk and optionally the memory. It can resume on a start command or automatically on arriving network traffic. A policy can also auto-delete a sandbox that has been stopped long enough, so abandoned work does not linger. Keep State and Bring Your Data Most agent tasks are short - run a script, or return a result. That is not the limit. A long-running agent needs state that outlives single runs, and volumes attach persistent storage to a sandbox at a path you choose. Code reads and writes to it through ordinary filesystem calls, and the data stays after the sandbox is deleted. There are three kinds: A Data Disk mounts to one sandbox at a time and gives it a fast, fully POSIX-compatible filesystem on local disk, which suits a working database, a build cache, or an agent's accumulated memory. An Azure Blob volume mounts to many sandboxes at once, for sharing a large read-heavy dataset rather than supporting concurrent writers. Azure Blob BYO volume (bring your own) does the same for blob storage you already own, referenced by resource ID. See What Your Sandboxes Are Doing Observability is how you know what a fleet of short-lived sandboxes did. Telemetry is opt-in and configured per sandbox when you create it, and it streams out while the sandbox runs, so you can go back to a run that finished hours ago. A sandbox emits four categories of data, and you choose which of them to collect: Category What it carries Console logs The stdout and stderr streams from each sandbox Sandbox metrics (platform) Platform-managed CPU usage, memory consumption, and network I/O per sandbox, sampled on an interval you set OpenTelemetry Signals emitted by your own application; the sandbox injects the OTLP endpoint, so an SDK you already use exports with no extra configuration Network egress decisions One record per outbound request, with the allow or deny decision the egress policy made Each category is then pointed at a destination: any OTLP-compatible collector, Log Analytics through the Logs Ingestion API, or Application Insights. They mix freely, so console logs can go to your collector while egress decisions go to Log Analytics. The credential that writes to the destination never enters the sandbox. OTLP resolves their from a sandbox-group secret, Log Analytics authenticates with a managed identity on the group. Separate from what a sandbox exports, the platform publishes cores and memory to Azure Monitor on the sandbox group, and that is what the portal shows for a Sandbox group. Sandbox group totals come by default, and you can opt-in for per sandbox details if you need that level of granularity. Use It From Code, a Shell, or a Browser Every one of these operations is available from code, a shell, a template, or a browser, so the choice comes down to what you are doing at the time. When the sandbox is part of your application, use an SDK. Today available for Python and TypeScript, with .NET on the way. An app or service you build creates sandboxes, writes and reads files, runs commands, mounts volumes, and captures snapshots as native objects in the language you use. When the work is scripted, the ACA CLI covers the same surface from Bash or PowerShell: aca sandbox create --disk ubuntu or aca sandbox snapshot. Commands accept label selectors, so automation and CI act on -l name=build-agent instead of tracking generated IDs. When the group itself is managed in source control, define it as infrastructure as code. Microsoft.App/sandboxGroups is a first-class ARM resource, so a Bicep template attaches a managed identity, links a delegated VNet subnet, and assigns data-plane roles. An ACA Terraform provider (Preview) covers the same group-level controls for teams standardized on Terraform. Sandboxes portal Designing a new resource type from scratch let us rethink the portal experience along with it. The creation flow asks the minimal set of questions - a disk image and resource tier - and you have a running sandbox. The Advanced section allows you to go deeper to ports, volumes, lifecycle policies, egress policies, logging and more. All there when you need it. What you see afterward is tailored the same way. A sandbox group gives you the overview of all sandboxes in that group: how many exist and how many are running, cores and memory in use over time. The most recent sandboxes with their state and size, and the disk images and snapshots the group can build from. A single sandbox gives you the machine: a terminal, live CPU, memory, storage, and network, a log stream, running processes, the files on disk, mounted volumes, and the egress traffic it generated with each request marked allowed or denied. You get that same experience wherever you start. Reach sandboxes from the Azure portal, alongside the rest of your resources and under the same subscriptions, RBAC, and policies, or go straight to the standalone ACA Sandboxes portal. It is the same experience either way, so there is nothing to relearn and nothing you can only do in one of them. How Much Does It Cost? Three types of charges: vCPU, per core-second while the sandbox runs. Memory, per GiB-second while the sandbox runs. Storage, per GB stored, for as long as you keep it. The resource tier of the sandbox determines the amount of vCPU and GiB of memory. These rates are on the Container Apps pricing page. Storage is charged (coming soon) at Premium Azure Blob ZRS rates and covers: Custom Disk Images, including Disk Snapshots. The platform converts your OCI container image into a bootable disk image. You pay to store one copy of that image for as long as you keep it, regardless of how many sandboxes boot from it. The OS disk each of those sandboxes runs on is not billed. Snapshots are the combined memory and disk snapshots of your sandboxes, including those taken automatically when a sandbox stops. Optional Data Disk Volumes and Azure Blob Volumes that can be attached to sandboxes. Thank You and What Comes Next General availability is not the destination, it is where many more of you get to start. During public preview that we announced in June 2026, the usage surpassed quickly a million sandboxes created every day. Our team worked closely with early customers that provided valuable feedback that shaped the product as it's today. From teams running real workloads on sandboxes ranging from cloud-native SaaS companies like Templafy to global organizations like KPMG and Cognite. KPMG built their Cowork AI Agent for their global workforce on ACA Sandboxes. Cognite uses ACA Sandboxes in their industrial Atlas AI system. Lastly, the Department for Education, South Australia - uses ACA sandboxes to power their EdChat - A safe place for every learner. We run the EdChat so students can learn by writing code and exploring data alongside AI, across 60,000 students and more than 40,000 staff. That model only works if every student gets an environment of their own, with clear guardrails, and can come back later to find their work exactly as they left it. Building that in house meant owning the machinery behind it. We estimate that moving to Azure Container Apps Sandboxes lets us retire close to 50,000 lines of code written to manage custom code interpreter and state ourselves. It is the per-user execution model that scales for our school system, and a lower maintenance burden for us. Cody Little, AI Technical Lead, Department for Education, South Australia In addition, many internal customers at Microsoft adopted ACA Sandboxes to build and enhance their products for their customers. Among those Microsoft Foundry built hosted agents, Copilot Studio hosts agents you create - both on ACA Sandboxes and – our very own Azure Container Apps Express built a modern and fast serverless container platform on sandboxes. Their feedback and your feedback set the next priorities, and we are grateful to you for it. Three focus areas of work that follow: Sandbox groups get more control and visibility, so a platform team can observe sandboxes, audit and enforce policies at the sandbox group level. Broad extensibility with more SDKs and tighter integration with VS Code. Lastly, extensive interoperability with Connectors and Triggers (now in preview) that will provide even larger customizability and on-behalf-of authentication (OBO), so agents securely reach the systems needed to do their job. Next Steps Open the ACA Sandboxes portal and create a group with a sandbox. Clone the samples repo and start with the working code. Read the documentation for the quick starts and the reference behind everything above. We appreciate your feedback, please submit it in the ACA Sandboxes portal, or open an issue in the Azure Container Apps repo Issues · microsoft/azure-container-apps.1.7KViews2likes0CommentsConnect Azure Functions to more services with managed connectors
Azure Functions can already connect to many Azure services through triggers and bindings. With managed connectors, your functions can access about 1,700 connectors across services such as Microsoft 365, Microsoft Teams, Dataverse, SharePoint, OneDrive, and third-party systems. Connector triggers deliver events from these services to your function, while typed connector clients let your code take actions against them. You get this broader integration surface without writing the webhook registration code or managing the OAuth tokens required to connect to each service. Focus on your function's business logic and let Azure Connector Namespace handles the connection. Azure Functions integration with Connector Namespace is currently in public preview. It supports .NET isolated, Python, and Node.js. Review the managed connectors overview for current language, hosting plan, and regional availability. To demonstrate how connector triggers and actions work together, this article follows a .NET sample that automates RFP intake across SharePoint, Azure Content Understanding, and Teams. From an uploaded RFP to Teams notification Consider an organization that receives requests for proposals (RFPs) in a shared SharePoint document library. Someone must read each document, identify the requested capabilities, determine which subject-matter experts should respond, and notify the right team. The automated RFP intake sample turns that process into an event-driven workflow: A customer uploads an RFP to a SharePoint document library. A SharePoint connector trigger invokes an Azure Function when the file is created. The function uses a typed SharePoint connector client to retrieve the file contents. Azure Content Understanding extracts the document’s text and layout. The function applies deterministic rules to identify the customer, required capabilities, and recommended subject-matter experts. The function uses a typed Teams connector client to post the results as an Adaptive Card in a channel. Connector Namespace manages the SharePoint and Teams connections. The function controls file processing, document analysis, routing rules, error handling, and notification content How the sample works The .NET sample demonstrates both parts of the connector programming model: a connector trigger receives an event from SharePoint, and typed connector clients provided by the Connector SDKs to perform actions against SharePoint and Teams. The function starts when the SharePoint When a file is created trigger detects a new RFP. It declares the trigger using the ConnectorTrigger attribute and receives a typed payload containing the file’s properties: [Function("OnNewFile")] public async Task OnNewFile( [ConnectorTrigger] SharePointOnlineOnNewFileItemsTriggerPayload payload, CancellationToken cancellationToken) { // Process the newly uploaded file. } Because the trigger provides file properties rather than its contents, the function uses a typed SharePoint client to retrieve the document: byte[] response = await _sharePoint.GetFileContentAsync( Uri.EscapeDataString(siteAddress), fileIdentifier, cancellationToken: cancellationToken); byte[] document = SharePointFileContent.Decode(response); The SharePoint and Teams clients are registered through dependency injection. Each client uses the runtime URL of its Connector Namespace connection and authenticates with DefaultAzureCredential: services.AddSingleton( new SharePointOnlineClient( new Uri(sharePointRuntimeUrl), credential)); services.AddSingleton( new TeamsClient( new Uri(teamsRuntimeUrl), credential)); The function sends the document to Content Understanding’s prebuilt-layout analyzer, which extracts its text and structure. It then applies deterministic C# rules to identify the customer and required capabilities and map those capabilities to predefined subject-matter expert roles. Finally, the function creates an Adaptive Card containing the results and posts it to the configured Teams channel with the typed Teams client: await _teams.PostCardToConversationAsync( postAs, postIn, request, cancellationToken); Connector Namespace handles the SharePoint and Teams connections, while the function controls the document analysis, routing logic, error handling, and notification content. Try the sample The RFP intake sample includes the function code, Bicep infrastructure, Azure Developer CLI configuration, and supporting scripts. Its README explains how to test the workflow locally and deploy it to Azure. Common connector patterns Managed connectors are useful when a function must react to events or perform operations in external systems. Common patterns include: Event to action: React to an event in one service and take an action in another. Event to enrich to action: Retrieve additional information related to an event before acting. Event to document analysis to action: Extract text and structure from a document, apply application rules, and send the result through another connector. Event to AI to action: Analyze event data with an AI service and write the result back through a connector. Extend an existing function app: Add connector-based integrations alongside HTTP, timer, queue, Service Bus, Event Grid, or Durable Functions workloads. The RFP sample combines several of these patterns. A SharePoint event starts the workflow, a SharePoint action retrieves the document, Content Understanding extracts its contents, application code enriches the result, and a Teams action sends the notification. Closing thoughts Managed connectors extend the external systems that can trigger your functions and the services your function code can act on. This brings services such as SharePoint, Teams, Microsoft 365, and many third-party systems into the Azure Functions programming model without requiring you to build the underlying webhook and OAuth infrastructure. Choose Azure Functions with managed connectors when you want this broader integration surface in a code-first application and need custom branching, application libraries and SDKs, other Functions bindings, document or AI processing, or application-specific logic between the trigger and action. If the workload primarily orchestrates connector operations, involves little custom code, and would benefit from a visual designer, Azure Logic Apps is usually the simpler choice. Resources Documentations Overview of managed connectors in Azure Functions Azure Functions connector samples Azure Connector Namespace overview Content Understanding prebuilt-layout analyzer Connector SDK GitHub repos .NET SDK Python SDK Node.js SDK285Views0likes0CommentsVirtual nodes on Azure Container Instances: a new compute layer for AKS
Meet virtual nodes on ACI Azure Kubernetes Service (AKS) gives you managed Kubernetes: the full Kubernetes API without operating the control plane yourself. Virtual nodes on Azure Container Instances go a step further, letting your pods run directly on Azure's serverless container platform, with the elasticity and with no capacity planning and no waiting for machines. Whether you already run AKS or want a managed Kubernetes that bursts without node management, this is for you. In short: virtual nodes on ACI attach Azure's serverless container platform to your cluster as Kubernetes nodes. Pods run as Hyper-V isolated containers, sized per pod rather than packed onto a fixed VM, up to 200 pods per virtual node. Run multiple virtual nodes, scaled as replicas, for more. They behave like any other pod: same kubectl, Helm, and GitOps. Kubernetes has always assumed a fixed set of machines underneath it. That assumption shapes everything above it: you size a node pool for a specific VM type in a specific region, you plan for peak rather than for average, and every workload on a node shares the same kernel and the same security boundary. Virtual nodes on ACI relaxs that assumption, which is what makes both elastic capacity and per container isolation possible without a different Kubernetes. If you've used the original AKS virtual nodes add-on (Virtual Kubelet based), this is not a rebrand. It is a new implementation that integrates far more deeply with Kubernetes, lifts most prior limitations (init containers, persistent volumes, managed identity, richer networking), and adds confidential containers as a first-class capability. The migration guide can be found here. Two capabilities carry the rest of this post: effortless burst capacity, and confidential containers. How virtual nodes on ACI work ACI runs every container as a Hyper-V isolated container, which means each one gets its own lightweight virtual machine boundary rather than sharing a kernel with its neighbors. Azure operates that platform. A virtual node connects it to your cluster. The cluster's control plane, the component that decides where each container runs, sees two kinds of destination: a small pool of virtual machines carrying cluster services, and one or more virtual nodes. From the application manifest's perspective, nothing changes. The pod lands on a virtual node; the virtual node hands it off to ACI. See Microsoft Learn: virtual nodes on ACI for the official capability and current limits. Virtual nodes on ACI in practice The rest of this post is hands on. You do not need to be a Kubernetes expert to follow it. kubectl is the command line tool for talking to a cluster, Helm installs packaged software into one, and a manifest is a text file describing what you want to run. If you have a cluster, everything below runs against it as written. The manifests behind the examples live in a companion demo repo. Setup is documented officially, and you can reproduce this end to end from the ACI virtual nodes documentation and the microsoft/virtualnodesOnAzureContainerInstances Helm repo. One requirement before you start: deploy into a delegated ACI subnet, meaning a subnet in your virtual network set aside for the ACI platform to place containers in. Size it for peak pod count plus headroom, since every pod consumes an address from it for its lifetime. Demo manifest files can be found in this repo, a personal sample repo provided as is and not a supported Microsoft artifact. Enable virtual nodes on ACI The virtual node is deployed via Helm. The Microsoft GitHub repo is itself a Helm repository, so a single helm install is all you strictly need. Cloning first, shown here, just makes it easier to customize values. Running kubectl get nodes afterward confirms the node registered. git clone https://github.com/microsoft/virtualnodesOnAzureContainerInstances.git helm install <yourReleaseName> ./virtualnodesOnAzureContainerInstances/Helm/virtualnode kubectl get nodes The virtual node appears alongside any existing capacity, ready to accept work. A virtual node is a Kubernetes node You target it the same way you would target any node. These few lines in a manifest say "run this on the virtual node": nodeSelector: virtualization: virtualnode2 kubernetes.io/os: linux tolerations: - key: virtual-kubelet.io/provider operator: Exists effect: NoSchedule That is the entire integration surface. No new API to learn, no separate deployment pipeline, no application changes. kubectl describe, kubectl logs, and kubectl exec, the standard commands for inspecting and troubleshooting, all work as they would anywhere else, including opening a shell inside a container running in a Hyper-V isolated boundary. Scaling stays trivial. kubectl scale deployment demo-deployment --replicas=10 lands every replica on the same virtual node, with no VMSS scale event, no provisioning latency, no climbing node-count chart. The same flow scales just as cleanly to hundreds. Cost follows the same shape. Each pod is billed per second against the cores and memory it requests, at ACI rates, and billing stops when the pod stops. Logs and metrics flow through the same path you already use, so existing dashboards and alerts keep working. One annotation makes a pod confidential Turning a regular container into a confidential one takes a single addition to its manifest: a policy that pins exactly which images, commands, environment variables, mounts, and capabilities are permitted inside the Trusted Execution Environment. The format is a base64 encoded Rego document, called a CCE (Container Confidential Enforcement) policy. You do not write that policy by hand. A tool generates it from the manifest you already have: az extension add -n confcom az confcom acipolicygen --virtual-node-yaml ./hello-world-deployment.yaml The tool pulls each image, hashes its layers, builds the allow-list, and injects the annotation back into the manifest. kubectl apply, and you're done. (acipolicygen has prerequisites of its own, including a working Docker installation; see the confcom documentation.) Here is why this is a genuinely new isolation primitive rather than a stronger version of an existing one. Most container security policy is enforced by software in the cluster, which means an attacker who compromises the host can potentially bypass it. This policy is enforced by the guest operating system inside the TEE instead. The underlying hardware, AMD SEV-SNP, also produces an attestation report, retrievable from inside the container, which is a cryptographic proof that the workload running is the workload you specified and nothing tampered with it. That is the guarantee regulated industries have been asking for, and increasingly the one AI workloads running untrusted code need too. The same per pod boundary is also what makes multi-tenancy on a single cluster realistic, though multi-tenancy in production still depends on your network and identity boundaries, which sit outside what the isolation layer itself provides. Background: Microsoft Learn: confidential containers on ACI. Wrapping up Virtual nodes on ACI give containers on Azure two things that were previously hard to deliver cleanly on Kubernetes: Effortless burst capacity on Azure's serverless container platform, billed per second for the cores and memory used, with no capacity planning and no waiting for machines. Confidential containers with hardware attested, per container isolation inside a Trusted Execution Environment. Virtual nodes are additive, not a replacement. Traditional node pools remain the right home for steady state, DaemonSet, and persistent volume workloads, and AKS features such as Node Auto Provisioning and Virtual Machine Node Pools already make that baseline more flexible. Virtual nodes on ACI absorb the spikes, the short-lived jobs, and the specialized isolation work on top. Where to start New to containers on Azure? Start with a small AKS cluster and add a virtual node from day one. You get a managed Kubernetes environment without having to guess your peak capacity in advance, and the elastic layer is there the first time you need it. Already running AKS? Add a virtual node to an existing cluster and move one bursty or short lived workload to it. Nothing else changes, and the comparison is immediate. Evaluating platforms? The capability that is hard to find elsewhere is the confidential containers path: hardware attested isolation per container, reachable through a standard Kubernetes manifest. The result: virtual nodes on ACI expand what AKS can run, with more capacity and stronger isolation, without changing the Kubernetes operating model you already use. Same kubectl, same manifests, same GitOps. New ceiling. For the high-level overview, official documentation, and Helm details, the Microsoft Learn is the source of truth. The companion repo holds the demo manifests used in this post. Acknowledgements I'd like to thank Gurpreet Virdi, Partner Group Engineering Manager, whose guidance shaped this post from the first outline through to publication. Her product leadership ensured this post reflects both the technical depth and the customer value of virtual nodes on ACI. Thanks to Gabriel Fuhrman, Senior Software Engineer, for his detailed technical review. His feedback refined the technical content and significantly improved the accuracy and depth of this post. Christopher Little, Principal CSA, shaped the enterprise adoption perspective, and Adam Sharif, CSA, reviewed the post from the earliest draft. Thanks also to Kirthi Maguluri, Senior Product Manager, and Varun Shandilya, Principal Product Manager, for their review of the blog.588Views1like0CommentsEnable Dynamic Workflows in Azure Functions hosted skills
Azure Functions already gives you a familiar way to build event-driven apps. A queue message, HTTP request, timer, or event triggers the code that handles the work. Azure Functions hosted skills (formerly Serverless Agents) add AI reasoning to that model. A hosted skill can read a request, use the regular tools you give it to inspect context, and choose the next step, while your triggers, tools, and business logic stay in place. When the work needs to keep going Consider an insurance policy servicing request. A hosted skill can use its regular tools to understand the requested change, look up the policy, and inspect the submitted documents. If the information is ready and the request can finish now, the normal tool loop, where the model calls a tool, reads the result, and decides the next step, is a good fit. That changes when the work must continue after the initial request. An insurance policy servicing request may need to inspect several documents in parallel, wait for a configured delay before checking again for missing information, and build a review packet after the checks it depends on complete. In a normal tool loop, each result returns to the model before the skill can decide what happens next. The application must keep the job alive, save its progress, and deliver the final result. At that point, the work needs to keep running independently of the original interaction instead of relying on the model and application to coordinate every step. For a queue or other non-HTTP trigger, the final result also needs to be written or sent somewhere useful because there is no response channel. Introducing Dynamic Workflows Dynamic Workflows brings a programmatic tool-calling pattern to Azure Functions hosted skills. Instead of sending every tool result back to the model so it can decide the next call, the model creates a structured, validated workflow plan once. Durable Functions then executes the allowed workflow-safe tool calls, waits, and subagent tasks, passing intermediate results through the workflow instead of the model context. This can reduce model turns and token use for multi-step work while making the work durable. That separation addresses the limits of the normal tool loop: the workflow store keeps state and intermediate results out of the model's context, independent checks can run in parallel, and durable timers resume waits without holding a worker open. Because a Durable Functions orchestration handles execution, the work can continue after the original request or a Functions worker restart. To test the difference, we ran the same structured multi-step task with the regular tool loop and with Dynamic Workflows, using a Foundry gpt-5.4-mini deployment. We ran it with inputs for one service and then ten services. In the Dynamic Workflows version, the model made the plan once, while the runtime kept intermediate tool results in the workflow store instead of sending them back to the model after every tool call. Dynamic Workflows used 56% fewer total model tokens for the one-service run and 93% fewer for the ten-service run, while producing the same final reports. Results will vary by workload and model, and small jobs can have planning overhead. The savings are largest when intermediate tool results would otherwise return to the model after every tool call. How it works Enable workflows in the hosted skill's Markdown front matter. The runtime then adds the management tools: start_workflow, get_workflow_status, list_workflows, cancel_workflow, and terminate_workflow. You do not implement those tools. You choose the workflow-safe tools and subagents that a plan can use. At run time, the AI model uses the hosted skill's instructions to generate a structured plan, limited to the workflow-safe tools and subagents you explicitly allow. The hosted skill calls start_workflow with that plan; the runtime validates it, starts a Durable Functions orchestration, and returns a workflow ID right away. --- name: Add Driver Review description: Prepares an add-driver document review for an insurance representative. workflows: enabled: true trigger: type: queue_trigger args: queue_name: policy-service-requests connection: AzureWebJobsStorage --- Put workflow-safe handlers under tools/ and decorate them for use in a workflow. Each handler must run synchronously, accept one dict argument, return JSON-serializable data, and be idempotent. A worker failure can cause a handler to run more than once, which is why that last point matters. Ordinary tools retain their existing behavior unless you explicitly make them available to a workflow. workflow_tool( description=( "Inspect one document from an add-driver request. Args: " "{document: <document>, position: int}. Returns the document and evidence state." ) ) def inspect_driver_document(args: dict[str, Any]) -> dict[str, Any]: document = args["document"] evidence_state = { "received": "present", "missing": "missing", "expired": "needs_current_copy", }[document["status"]] return { "position": args["position"], "document_id": document["document_id"], "type": document["type"], "file_name": document["file_name"], "evidence_state": evidence_state, } Dynamic Workflows runs on Durable Functions. You can configure Durable Task Scheduler in host.json and use its dashboard to see per-instance task state, retry history, and controls for work that is still running. Durable timers let a workflow wait without keeping a worker busy, then resume the steps that are ready. The workflow stays visible and durable instead of depending on an open request or a best-effort background task. Get started Build your first Azure Functions hosted skill dynamic workflow with the quickstart, then use the overview and sample to go deeper: Follow the Dynamic Workflows quickstart. Read the Dynamic Workflows overview. Browse the insurance policy review sample for an end-to-end implementation.342Views0likes0CommentsBring Your Own Orchestrator to Azure Container Apps Jobs
Azure Container Apps Jobs are a good fit for batch processing, ETL, machine learning, reports, and other tasks that run to completion. But when those tasks have dependencies, retries, or fan-out, you still need an orchestrator. Many teams already have one. The community-maintained Bring Your Own Orchestrator collection provides 13 templates that connect existing workflow engines to Azure Container Apps Jobs. The collection is also listed in the Microsoft Azure Container Apps template index. The idea is simple: Your orchestrator manages schedules, dependencies, retries, and workflow history. Azure Container Apps Jobs runs each containerized task and reports the result. You keep the control plane your team knows while ACA Jobs provides the execution layer. How it works Each integration follows the same flow: The orchestrator authenticates to Azure. It starts an ACA Job execution. It waits for that execution to succeed or fail. It uses the result to continue, retry, or stop the workflow. Your orchestrator ---> Azure Container Apps Job ^ | +---- execution result ---+ The templates package this flow in the native model of each platform: an Airflow operator, a Temporal Activity, an Argo workflow template, a Camunda service task, or visual actions in Logic Apps and n8n. The workload container stays independent of the orchestrator that launched it. Choose the orchestrator that fits the workflow There is no single best orchestrator for every workload. The useful question is which control plane matches the way your team models work. When this describes your team Start with Why You already operate Airflow Airflow on ACA Jobs Adds an ACA Jobs operator without replacing your Airflow deployment You need a complete Airflow environment Airflow hosted on ACA Deploys the Airflow control plane and the ACA Jobs integration Your workflows are Kubernetes-native and run from AKS Argo Workflows Uses Argo workflow templates and AKS workload identity You model long-running business processes in BPMN Camunda 8 Connects Camunda service tasks to ACA Job executions You use JSON-defined microservice workflows Conductor Uses Conductor workers and native FORK_JOIN workflows You need durable replay, heartbeats, and resilient retries Temporal Keeps Temporal as the durable control plane while ACA Jobs runs the workload You build asset-centric Python data pipelines Dagster Uses Dagster resources, ops, and dynamic mapping You build general Python flows and task automation Prefect Uses Prefect tasks, flows, and mapped execution You prefer visual automation and SaaS integrations n8n Provides visual workflows for starting and observing ACA Jobs You use Azure-native data pipelines Azure Data Factory and Fabric Provides pipeline definitions for Azure data integration workflows You need connector-rich application integration Logic Apps Standard Uses stateful workflows, connectors, and native control flow You want Azure-native, code-first durable orchestration Durable Functions Uses durable orchestrations, activities, retries, and fan-out/fan-in You already operate a Dapr-enabled workflow host Dapr Workflow Demonstrates Dapr Workflow directing external ACA Job workloads The Bring Your Own Orchestrator catalog keeps this comparison current and links to deployment instructions for every option. Before production The existing-orchestrator templates are designed around managed identity, scoped Azure RBAC, failure handling, and native fan-out/fan-in examples. Their fan-out samples default to five shards and accept configurations from 1 to 50. Treat higher shard counts as configuration support, not a throughput guarantee. Test them against your Azure quotas, orchestrator limits, and downstream systems. Two template-specific boundaries are worth calling out: Dapr Workflow is a preview architecture Azure Container Apps Jobs do not host Dapr sidecars. The Dapr workflow runtime must run in a separate Dapr-enabled host and start ACA Jobs through Azure Resource Manager. The template is therefore labeled preview architecture. Fabric still needs a native workspace run The Azure Data Factory path has live validation. The included Fabric pipeline is structurally validated but still needs a native run in a Fabric workspace. Each repository README documents its validation scope and limitations. Get started Open the template catalog. Choose the orchestrator your team already uses. Review that template's prerequisites and validation notes. Deploy the sample ACA Job with azd up . Run the single-job example, then test fan-out and failure behavior. For example, if Airflow is already your standard: git clone https://github.com/hetvip2/airflow-on-aca-jobs cd airflow-on-aca-jobs azd up The exact setup differs by orchestrator, but the target remains ACA Jobs. Try the templates Compare all 13 orchestrator templates and choose the control plane that matches your team. Review the Azure Container Apps Jobs documentation for triggers, permissions, and platform limits. Browse the ACA community template collections to find the collection in the Microsoft Azure Container Apps repository. Closing thoughts Using Azure Container Apps Jobs should not require an orchestrator migration. Keep the workflow engine your team already trusts and use ACA Jobs for containerized task execution. Explore all 13 options in the Bring Your Own Orchestrator to Azure Container Apps Jobs collection. References Azure Container Apps Jobs overview Azure Container Apps Jobs management API Managed identities in Azure Container Apps Azure Developer CLI documentation Bring Your Own Orchestrator template catalog Azure Container Apps community template collections712Views1like0CommentsOrchestrate Azure Container Apps Jobs with Apache Airflow
Azure Container Apps (ACA) Jobs are a great way to run work that starts, does something, and finishes: nightly batch, data processing, ETL, ML scoring, report generation. They scale to zero, bill per execution, and run any container you give them. But the moment your "one job" becomes "a set of jobs that depend on each other," a gap appears: How do I run twenty jobs in parallel, wait for all of them, then run one more job only if they all succeeded — and retry just the one that failed? A single ACA Job can't express that on its own. What you're describing is an orchestrator, and the most widely adopted one in the data world is Apache Airflow. This post introduces two open-source templates that connect the two, so Airflow becomes the brain and ACA Jobs become the muscle. Pick the one that matches what you already run: airflow-on-aca-jobs: you already have Airflow. Drop in an operator and point it at ACA Jobs. Host nothing new. airflow-hosted-on-aca: you don't have Airflow. Get a full one running on Azure Container Apps with one command. Both use the same operator and the same DAGs, so you can start with one and move to the other later without rewriting your workflows. See Airflow orchestrate real ACA Job executions with parallel fan-out, dependency ordering, and automatic retries. Why ACA Jobs need an orchestrator A plain ACA Job is great at one thing: run this container to completion, then stop. That covers a scheduled job or a one-off task perfectly. Real pipelines need more than that: Dependency ordering: step B runs only after step A succeeds. Parallel fan-out: launch one execution per file, per store, or per partition, all at once, then wait for the whole batch. Per-task retries: if one execution in a batch of fifty fails, retry just that one, not the other forty-nine. Backfills and scheduling: re-run yesterday's pipeline, or run every night with a full history of what happened. These are the problems an orchestrator solves. Instead of building that logic yourself, you let Airflow handle the graph, the scheduling, and the retries, while ACA Jobs run the compute. You get serverless, scale-to-zero workers, and you didn't have to stand up a scheduler to get them. The operator that ties them together Both templates ship the same small plugin: an Airflow operator called AzureContainerAppsJobOperator . In a DAG it looks like any other task: report_sales = AzureContainerAppsJobOperator( task_id="report_store_sales", subscription_id="{{ var.value.azure_subscription_id }}", resource_group="{{ var.value.aca_resource_group }}", job_name="{{ var.value.aca_job_name }}", image="python:3.12-slim", command=["python", "-c", MY_PROGRAM], env_vars={"STORE_NAME": "Seattle"}, deferrable=True, ) A few things make this operator easy to work with: Per-execution overrides. It takes the ACA Job you point it at and overrides the image , command , args , and env_vars for that run. You can drive many different workloads from a single ACA Job definition, and you don't need to build or push a custom image just to try something. The example above runs the stock python:3.12-slim image with an inline program. Deferrable by default. With deferrable=True , Airflow frees its worker slot while the ACA Job runs and resumes when it finishes. That means your fan-out width is bounded by ACA, not by how many Airflow workers you have. You can launch dozens of parallel executions cheaply. No secrets required. Authentication resolves in a sensible order: an Airflow Connection if you set one, otherwise an AZURE_ACCESS_TOKEN environment variable, otherwise DefaultAzureCredential (managed identity). In Azure, the hosted template uses a managed identity so nothing sensitive is stored in Airflow at all. Because both templates share this operator, a DAG written for one runs unchanged on the other. Option 1: Bring your own Airflow (host nothing) Choose airflow-on-aca-jobs if you already run Airflow: Azure Managed Airflow, MWAA, Astronomer, or your own deployment. You keep that Airflow exactly as it is and simply teach it to talk to ACA Jobs. +------------------------------------------+ | Your Airflow (you host it, unchanged) | | runs AzureContainerAppsJobOperator | +------------------------------------------+ | | ACA Jobs REST API v +------------------------------------------+ | ACA Job (Azure Container Apps) | | | | store 1 | store 2 | ... | store N | | parallel executions -> scale to zero | +------------------------------------------+ Your existing Airflow runs the operator; ACA Jobs run the work. You host nothing new. Adoption is three small steps: Copy the operator into your Airflow's plugins/ folder. Add a DAG that uses AzureContainerAppsJobOperator . Set three Airflow Variables so the operator knows which job to drive: Airflow Variable Value azure_subscription_id your subscription id aca_resource_group the resource group holding the ACA Job aca_job_name the ACA Job name That's the whole integration. Nothing new to host, no extra scheduler or database, no custom image. ACA Jobs just become another task type Airflow can call. If you want a job to point at first, the template includes an Azure Developer CLI ( azd ) deployment that stands up a sample ACA Job for you: git clone https://github.com/hetvip2/airflow-on-aca-jobs cd airflow-on-aca-jobs azd up # deploys a sample ACA Job, prints its resource group + name Then copy airflow/plugins/ and airflow/dags/ into your Airflow, set the three Variables, and trigger the DAG. Option 2: Airflow hosted on ACA (turnkey) Choose airflow-hosted-on-aca if you don't already have an orchestrator and want one running next to your jobs. One command provisions the whole thing on Azure Container Apps: azd up | v +------------------------------------------+ | Airflow control plane on ACA | | web | scheduler | triggerer | | Postgres (metadata) + Azure Files (dags)| | Managed Identity - no secrets stored | +------------------------------------------+ | | ACA Jobs REST API v +------------------------------------------+ | ACA Job (Azure Container Apps) | | | | store 1 | store 2 | ... | store N | | parallel executions -> scale to zero | +------------------------------------------+ One command deploys the whole Airflow control plane on ACA, right next to the jobs it drives. git clone https://github.com/hetvip2/airflow-hosted-on-aca cd airflow-hosted-on-aca azd env new my-airflow azd up # prints your Airflow URL when it finishes azd up deploys a complete, working Airflow control plane on ACA: airflow-web, airflow-scheduler, and airflow-triggerer running as Container Apps on LocalExecutor, so there's no Celery or Redis to operate. A Postgres metadata database. A user-assigned managed identity with permission to call the ACA Jobs API, so the operator authenticates with no secrets stored in Airflow. A sample ACA Job for Airflow to drive out of the box. Your DAGs and plugins live on a mounted Azure Files share, so you ship new workflows by re-uploading files rather than rebuilding an image: cp my_dag.py airflow/dags/ azd hooks run postprovision # uploads dags + plugins to the share Airflow picks up the change within a minute. You now own a real orchestrator, hosted serverlessly on the same platform as your jobs. Which one should you pick? Option 1: airflow-on-aca-jobs Option 2: airflow-hosted-on-aca Best when You already run Airflow You don't have Airflow yet Setup Copy the operator + a DAG + 3 Variables azd up (one command) Who hosts Airflow You do (unchanged) Azure Container Apps Authentication Connection or short-lived token Managed identity, nothing stored Ownership Lowest: nothing new to run Turnkey: a full orchestrator you own The important part: the workload never changes. The same DAG and the same operator drive the same ACA Job executions in both. Start wherever you are today, and switch later with zero changes to your pipelines. See it end to end Picture a retailer that wants one number every night: total sales across all stores. Each store reports its own sales as a separate ACA Job execution, all running in parallel. When every store is in, a final job adds them into the company total. That one workflow exercises exactly what a plain Job can't do alone: parallel fan-out: one ACA Job execution per store, all at once dependency ordering: the roll-up runs only after every store reports per-task retries: if a store's execution fails, Airflow retries just that store, and the nightly total still lands In Airflow's Graph view you watch the store tasks light up together, then the roll-up run last. In the Azure portal you watch real executions appear under your ACA Job and scale back to zero when they finish. Same job, same DAG, whichever template you chose. Call to action If you run batch, ETL, or any multi-step work on Azure Container Apps Jobs, give one of these templates a try: Already have Airflow? Start with airflow-on-aca-jobs. Need an orchestrator? Start with airflow-hosted-on-aca. Both are open source, deploy with azd up , and share the same operator so you can move between them freely. Try them out and let us know what you orchestrate.640Views1like3CommentsHow to build long-running MCP tools on Azure Functions
Recently, a customer building servers with the Azure Functions MCP extension reached out and asked: How do I handle tools that take longer than the client is willing to wait? This becomes especially relevant when tool calls move beyond simple request/response into multi-step workflows and long-running operations. At the same time, MCP is evolving to address exactly this. The Tasks extension is introduced in the 2026-07-28 release candidate, defining a standard way to model long-running work. In this post, we’ll walk through how to build long-running MCP tools on Azure Functions using Durable Functions , a framework for authoring stateful, long-running workflows as ordinary code, with checkpointing, scaling, and recovery handled automatically. MCP tools today Today, MCP tools are fundamentally request/response: the client issues a tools/call the server returns a result This works well for fast operations, but breaks down when: workflows take minutes execution depends on multiple steps latency is unpredictable In practice, clients enforce their own tool-call timeouts. These aren't standardized by the MCP spec and vary per client, but they're often in the ~30–60 second range. If a tool exceeds that window: In practice, clients often enforce short timeouts. If a tool exceeds that window: the client times out the agent observes a failed call the underlying work may still be running So the core issue is that you have synchronous tool calls don’t naturally model long-running work. The MCP Tasks extension The Tasks extension to address this. With the extension, a server can respond to a tools/call with an asynchronous task handle instead of a final result, and the client drives the lifecycle from there: tasks/get: poll the task's status tasks/update: submit input back to the server if the task reaches input_required tasks/cancel: cancel an in-flight task A task carries a status ("working", "input_required", "completed", "failed", or "cancelled") and on completion, the final result. Task creation is server-directed: the client advertises support by including the extension in its per-request capabilities, and the server decides per request whether to return a task. A server won't return a task to a client that hasn't advertised support. It's important to note that Tasks rely on ecosystem support. Clients must advertise the extension, and MCP SDKs must implement the task lifecycle, before servers can use it. So while Tasks is now a defined extension, broad client and SDK support is still in progress. Implement long-runng tasks with Durable Functions today Until the Tasks extension is broadly supported across clients, we need a pattern that works with existing request/response clients and supports long-running execution. The following samples show how, using Durable Functions: Python NET The long-running work in this sample mines a short chain of blocks. Each block requires solving a computational puzzle where the system keeps trying different inputs until it finds one that produces a result matching a specific pattern (for example, starting with a certain number of zeros). Because this involves lots of trial and error, it naturally takes time, making it a good example of a long-running workflow. The server in the sample exposes two tools: start_mining Starts a Durable Functions orchestration to mine the blocks Waits briefly (within a configurable budget) Returns result inline if completed within budget OR returns workflow_id if still running get_mining_result Takes the workflow_id Returns the current state, e.g. "completed", "running", "failed", or "not_found" To ensure that the agent calls the tools in the right order, workflow_id is a required parameter of get_mining_result, so the agent can't poll without starting a mining run first. Also, the "running" response carries a poll_after_seconds and a next instruction, ensuring the agent to poll again if work is not done rather than give up or assume completion. Even so, the poll path still relies on the agent correctly remembering, and not hallucinating, the workflow_id it was handed. If it garbles or invents an id, the poll lands on the wrong instance or none at all (which is why get_mining_result returns "not_found" rather than guessing). What changes with the Tasks extension Once the Tasks extension is fully implemented across clients and SDKs, the model becomes simpler and more reliable: the server returns a Task handle, the client manages the polling and lifecyle calls, and the SDK tracks execution state. This removes a key limitation of today’s solution, which requires the agent to remember and correctly pass identifiers like workflow_id. Call to action Try out the sample and let us know whether it addresses your MCP needs around long-running or workflow type tools!707Views0likes0Comments