<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>Microsoft Foundry Blog articles</title>
    <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/bg-p/azure-ai-foundry-blog</link>
    <description>Microsoft Foundry Blog articles</description>
    <pubDate>Thu, 01 Oct 2026 21:38:48 GMT</pubDate>
    <dc:creator>azure-ai-foundry-blog</dc:creator>
    <dc:date>2026-10-01T21:38:48Z</dc:date>
    <item>
      <title>Your Agent Shouldn't Wait for a Prompt: Build an Event-Driven Microsoft Foundry Routine</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/your-agent-shouldn-t-wait-for-a-prompt-build-an-event-driven/ba-p/4559897</link>
      <description>&lt;P data-line="2"&gt;Most agents are still waiting in a chat window.&lt;/P&gt;
&lt;P data-line="4"&gt;They may be capable of classifying an incident, finding the right documentation, or recommending an owner - but nothing happens until someone remembers to ask. The hard part is no longer always the reasoning. It is noticing that work has arrived, invoking the agent securely, and keeping enough history to understand what happened.&lt;/P&gt;
&lt;P data-line="6"&gt;&lt;A href="https://learn.microsoft.com/azure/foundry/agents/how-to/use-routines" target="_blank" rel="noopener" data-href="https://learn.microsoft.com/azure/foundry/agents/how-to/use-routines"&gt;Routines in Microsoft Foundry&lt;/A&gt; close that gap. A routine connects a trigger - such as a schedule, timer, GitHub issue, or Microsoft Teams message - to an agent action. Microsoft Foundry queues the invocation, runs the agent, and stores a run record for later inspection.&lt;/P&gt;
&lt;P data-line="8"&gt;In this post, we will build an event-driven routine for a familiar developer workflow:&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P data-line="10"&gt;When someone opens a GitHub issue, invoke a triage agent immediately - without waiting for a person to copy the issue into a chat.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P data-line="12"&gt;Along the way, we will look at the architecture, create the routine with the Azure Developer CLI (azd), test it, inspect its run history, and make an explicit identity decision before putting it into production.&lt;/P&gt;
&lt;H2 data-line="14"&gt;From conversational agents to event-driven agents&lt;/H2&gt;
&lt;P data-line="16"&gt;A chat-first agent follows a request-response pattern:&lt;/P&gt;
&lt;OL data-line="18"&gt;
&lt;LI data-line="18"&gt;A user opens an interface.&lt;/LI&gt;
&lt;LI data-line="19"&gt;The user provides a prompt.&lt;/LI&gt;
&lt;LI data-line="20"&gt;The agent performs work.&lt;/LI&gt;
&lt;LI data-line="21"&gt;The interaction ends or waits for another prompt.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P data-line="23"&gt;An event-driven agent starts differently. Work in an external system becomes the prompt.&lt;/P&gt;
&lt;P data-line="25"&gt;A newly opened issue, for example, already contains useful context: its title, description, author, repository, labels, and timestamps. A routine can receive that event through an authorized connection and invoke an agent while the context is still fresh.&lt;/P&gt;
&lt;P data-line="27"&gt;That changes the agent's role. It is no longer only a place people go for answers. It becomes a participant in an operational workflow.&lt;/P&gt;
&lt;img /&gt;
&lt;P data-line="31"&gt;The routine does not replace the agent. It supplies the managed automation around it:&lt;/P&gt;
&lt;UL data-line="33"&gt;
&lt;LI data-line="33"&gt;&lt;STRONG&gt;Trigger:&lt;/STRONG&gt;&amp;nbsp;Defines when work starts.&lt;/LI&gt;
&lt;LI data-line="34"&gt;&lt;STRONG&gt;Connection:&lt;/STRONG&gt;&amp;nbsp;Authenticates the event source.&lt;/LI&gt;
&lt;LI data-line="35"&gt;&lt;STRONG&gt;Action:&lt;/STRONG&gt;&amp;nbsp;Identifies the agent and invocation protocol.&lt;/LI&gt;
&lt;LI data-line="36"&gt;&lt;STRONG&gt;Dispatch identity:&lt;/STRONG&gt;&amp;nbsp;Determines whose permissions are used when the agent and its tools run.&lt;/LI&gt;
&lt;LI data-line="37"&gt;&lt;STRONG&gt;Run history:&lt;/STRONG&gt;&amp;nbsp;Records executions so operators can inspect outcomes.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="39"&gt;This separation is useful. The agent owns the reasoning; the routine owns when and how that reasoning begins.&lt;/P&gt;
&lt;H2 data-line="41"&gt;The scenario: triage every new GitHub issue&lt;/H2&gt;
&lt;P data-line="43"&gt;Our example assumes that a triage agent is already deployed in a Microsoft Foundry project. When an issue opens, the agent should:&lt;/P&gt;
&lt;OL data-line="45"&gt;
&lt;LI data-line="45"&gt;Summarize the issue in two or three sentences.&lt;/LI&gt;
&lt;LI data-line="46"&gt;Classify it as a bug, feature request, documentation issue, or support question.&lt;/LI&gt;
&lt;LI data-line="47"&gt;Estimate severity and explain the evidence.&lt;/LI&gt;
&lt;LI data-line="48"&gt;Recommend an owner or team.&lt;/LI&gt;
&lt;LI data-line="49"&gt;Identify missing reproduction details.&lt;/LI&gt;
&lt;LI data-line="50"&gt;Produce a proposed response for a maintainer to review.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P data-line="52"&gt;Keeping a human review step is intentional. Event-driven does not have to mean unrestricted autonomy. A routine can automate the expensive first pass while a maintainer remains responsible for labels, assignments, and public responses.&lt;/P&gt;
&lt;H2 data-line="54"&gt;Prerequisites&lt;/H2&gt;
&lt;P data-line="56"&gt;You need:&lt;/P&gt;
&lt;UL data-line="58"&gt;
&lt;LI data-line="58"&gt;An active Microsoft Foundry project.&lt;/LI&gt;
&lt;LI data-line="59"&gt;The&amp;nbsp;&lt;STRONG&gt;Foundry User&lt;/STRONG&gt;&amp;nbsp;role or higher on the project.&lt;/LI&gt;
&lt;LI data-line="60"&gt;A deployed prompt agent or hosted agent. Workflow agents are not currently supported by routines.&lt;/LI&gt;
&lt;LI data-line="61"&gt;The&amp;nbsp;&lt;A href="https://learn.microsoft.com/azure/developer/azure-developer-cli/install-azd" target="_blank" rel="noopener" data-href="https://learn.microsoft.com/azure/developer/azure-developer-cli/install-azd"&gt;Azure Developer CLI&lt;/A&gt;.&lt;/LI&gt;
&lt;LI data-line="62"&gt;A GitHub connection authorized in the Foundry project.&lt;/LI&gt;
&lt;LI data-line="63"&gt;The routines extension for azd.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="65"&gt;Install the extension and confirm that the routine commands are available:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;azd extension install azure.ai.routines
azd ai routine --help&lt;/LI-CODE&gt;
&lt;P data-line="72"&gt;Set the project endpoint through your active&amp;nbsp;azd&amp;nbsp;environment or pass it explicitly to each command:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;azd env set AZURE_AI_PROJECT_ENDPOINT \
  "https://&amp;lt;account&amp;gt;.services.ai.azure.com/api/projects/&amp;lt;project&amp;gt;"&lt;/LI-CODE&gt;
&lt;P data-line="79"&gt;The agent must exist before you attach a routine to it. A routine references an agent; it does not deploy one.&lt;/P&gt;
&lt;H2 data-line="81"&gt;Create the GitHub issue routine&lt;/H2&gt;
&lt;P data-line="83"&gt;The following command creates an enabled routine named&amp;nbsp;triage-on-open. Replace the placeholders with the connection and repository details from your environment:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;azd ai routine create triage-on-open \
  --trigger github-issue \
  --connection-id "&amp;lt;workspace-connection-id&amp;gt;" \
  --owner "&amp;lt;github-owner&amp;gt;" \
  --repository "&amp;lt;github-repository&amp;gt;" \
  --issue-event opened \
  --action agent-invoke \
  --agent-name "triage-agent" \
  --description "Triage every newly opened GitHub issue"
&lt;/LI-CODE&gt;
&lt;P data-line="97"&gt;There are two important type translations in this command:&lt;/P&gt;
&lt;UL data-line="99"&gt;
&lt;LI data-line="99"&gt;The CLI alias&amp;nbsp;github-issue&amp;nbsp;represents the routine trigger type&amp;nbsp;github_issue.&lt;/LI&gt;
&lt;LI data-line="100"&gt;The CLI alias&amp;nbsp;agent-invoke&amp;nbsp;invokes the agent through the Invocations API.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="102"&gt;If your agent uses the Responses API instead, use&amp;nbsp;--action agent-response. Choose the protocol that matches the deployed agent rather than treating the two actions as interchangeable.&lt;/P&gt;
&lt;P data-line="104"&gt;For automation that belongs in source control, define the routine as an&amp;nbsp;azure.ai.routine&amp;nbsp;service in&amp;nbsp;azure.yaml. This makes the relationship between the agent and its trigger reproducible across environments:&lt;/P&gt;
&lt;LI-CODE lang="yaml"&gt;services:
  triage-agent:
    host: azure.ai.agent
    project: ./agent

  triage-on-open:
    host: azure.ai.routine
    uses:
      - triage-agent
    description: Triage every newly opened GitHub issue
    enabled: true
    triggers:
      issue-opened:
        type: github_issue
        connection_id: ${GITHUB_CONNECTION_ID}
        owner: ${GITHUB_OWNER}
        repository: ${GITHUB_REPOSITORY}
        issue_event: opened
    action:
      type: invoke_agent_invocations_api
      agent_name: triage-agent
&lt;/LI-CODE&gt;
&lt;P data-line="130"&gt;Deploy the routine after the target agent:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;azd deploy triage-on-open --no-prompt&lt;/LI-CODE&gt;
&lt;P data-line="136"&gt;The uses relationship tells azd to order the agent before the routine. Deployment is idempotent: redeploying updates the named routine instead of creating duplicates.&lt;/P&gt;
&lt;H2 data-line="138"&gt;Give the agent a clear triage contract&lt;/H2&gt;
&lt;P data-line="140"&gt;Automation magnifies ambiguity. A vague instruction that is merely inconvenient in a chat can produce inconsistent work every time an event fires.&lt;/P&gt;
&lt;P data-line="142"&gt;Give the triage agent a bounded contract such as:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;You are the first-pass triage agent for this repository.

For each newly opened issue:
1. Summarize the reported behavior without adding facts.
2. Classify it as bug, feature, documentation, or support.
3. Assign severity only when the issue contains supporting evidence.
4. List missing information needed to reproduce or route the issue.
5. Recommend an owner from the approved ownership map.
6. Draft a response, but do not publish, close, label, or assign the issue.

Return structured JSON that matches the triage schema.
&lt;/LI-CODE&gt;
&lt;P data-line="158"&gt;A useful output contract might look like this:&lt;/P&gt;
&lt;LI-CODE lang="json"&gt;{
  "summary": "The CLI exits when a project endpoint contains an explicit port.",
  "category": "bug",
  "severity": {
    "level": "medium",
    "reason": "The issue blocks routine creation but has a documented workaround."
  },
  "missing_information": [
    "Azure Developer CLI version",
    "Redacted project endpoint shape",
    "Full error output"
  ],
  "recommended_owner": "developer-experience",
  "proposed_response": "Thanks for the report. Could you share..."
}
&lt;/LI-CODE&gt;
&lt;P data-line="178"&gt;Structured output gives downstream systems something predictable to validate. It also makes evaluation easier: you can test category accuracy, required-field completeness, unsupported severity claims, and whether the agent attempted a prohibited action.&lt;/P&gt;
&lt;H2 data-line="180"&gt;Test before waiting for a real event&lt;/H2&gt;
&lt;P data-line="182"&gt;Start by checking that Foundry stored the routine you intended:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;azd ai routine show triage-on-open --output json&lt;/LI-CODE&gt;
&lt;P data-line="188"&gt;Confirm:&lt;/P&gt;
&lt;UL data-line="190"&gt;
&lt;LI data-line="190"&gt;The routine is enabled.&lt;/LI&gt;
&lt;LI data-line="191"&gt;The trigger watches the correct owner and repository.&lt;/LI&gt;
&lt;LI data-line="192"&gt;issue_event&amp;nbsp;is&amp;nbsp;opened.&lt;/LI&gt;
&lt;LI data-line="193"&gt;The action references the intended agent.&lt;/LI&gt;
&lt;LI data-line="194"&gt;The connection ID belongs to the expected Foundry project.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="196"&gt;You can manually dispatch a routine while testing:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;azd ai routine dispatch triage-on-open \
  --input '{"test":true,"issue":{"number":123,"title":"Test triage event"}}'
&lt;/LI-CODE&gt;
&lt;P data-line="203"&gt;The manual input is a one-time override for that dispatch. It does not replace the event payload or modify the routine's stored configuration.&lt;/P&gt;
&lt;P data-line="205"&gt;Inspect recent executions:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;azd ai routine run list triage-on-open --top 20
&lt;/LI-CODE&gt;
&lt;P data-line="211"&gt;Then open a test issue in the watched repository and inspect the run list again. A production test should verify more than "the agent ran." Check that the correct event started the run, the agent received enough context, tool calls used the intended identity, the output matched the schema, and prohibited actions did not occur.&lt;/P&gt;
&lt;P data-line="211"&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;H2 data-line="215"&gt;Choose the dispatch identity deliberately&lt;/H2&gt;
&lt;P data-line="217"&gt;Every routine uses the&amp;nbsp;&lt;STRONG&gt;agent identity by default&lt;/STRONG&gt;. This is usually the better fit for unattended automation because access belongs to the agent rather than to an employee's account.&lt;/P&gt;
&lt;P data-line="219"&gt;Use agent identity when the agent's tools authenticate with managed identity, workload identity, or keys and the agent has been granted only the permissions required for the task.&lt;/P&gt;
&lt;P data-line="221"&gt;Some tools require delegated user access. In that case, you can create the routine with&amp;nbsp;&lt;STRONG&gt;creator identity&lt;/STRONG&gt;. Creator identity means the Microsoft Entra identity of the person or service principal that creates the routine—not the agent publisher, connection creator, latest editor, or user who caused an event.&lt;/P&gt;
&lt;P data-line="223"&gt;This distinction has operational consequences:&lt;/P&gt;
&lt;UL data-line="225"&gt;
&lt;LI data-line="225"&gt;If the creator loses access or consent, delegated tool calls can fail.&lt;/LI&gt;
&lt;LI data-line="226"&gt;Recreating the routine as another principal changes the delegated creator identity.&lt;/LI&gt;
&lt;LI data-line="227"&gt;The event connection identity is separate from the identity used to dispatch the agent.&lt;/LI&gt;
&lt;LI data-line="228"&gt;Dispatch identity is a creation-time decision. To switch an existing routine between agent and creator identity, delete and recreate it.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="230"&gt;For a GitHub triage flow, a strong starting design is:&lt;/P&gt;
&lt;UL data-line="232"&gt;
&lt;LI data-line="232"&gt;Use a narrowly scoped project connection to receive issue events.&lt;/LI&gt;
&lt;LI data-line="233"&gt;Use agent identity for Foundry and Azure resources.&lt;/LI&gt;
&lt;LI data-line="234"&gt;Keep repository-changing actions disabled until evaluation demonstrates reliable behavior.&lt;/LI&gt;
&lt;LI data-line="235"&gt;Require human approval before posting, assigning, labeling, or closing.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="237"&gt;Identity is not a deployment detail. It is part of the automation's behavior and should be reviewed with the same care as the prompt and tool list.&lt;/P&gt;
&lt;H2 data-line="239"&gt;Make failure visible&lt;/H2&gt;
&lt;P data-line="241"&gt;An event-driven agent can fail even when its reasoning is sound. The connection might expire. The agent might receive an unexpected payload. A tool might lose permission. The output might violate its schema.&lt;/P&gt;
&lt;P data-line="243"&gt;Monitor the workflow at four boundaries:&lt;/P&gt;
&lt;OL data-line="245"&gt;
&lt;LI data-line="245"&gt;&lt;STRONG&gt;Trigger:&lt;/STRONG&gt;&amp;nbsp;Did the expected event fire the routine exactly once?&lt;/LI&gt;
&lt;LI data-line="246"&gt;&lt;STRONG&gt;Invocation:&lt;/STRONG&gt;&amp;nbsp;Did Foundry invoke the intended agent and protocol?&lt;/LI&gt;
&lt;LI data-line="247"&gt;&lt;STRONG&gt;Tools:&lt;/STRONG&gt;&amp;nbsp;Did tool calls succeed with the intended identity and scope?&lt;/LI&gt;
&lt;LI data-line="248"&gt;&lt;STRONG&gt;Outcome:&lt;/STRONG&gt;&amp;nbsp;Did the output satisfy the triage contract?&lt;/LI&gt;
&lt;/OL&gt;
&lt;P data-line="250"&gt;Use&amp;nbsp;azd ai routine run list&amp;nbsp;for routine execution history and your agent's Foundry observability data for traces, tool calls, latency, and failures. Preserve representative failures as evaluation cases instead of fixing each incident only in the prompt.&lt;/P&gt;
&lt;P data-line="252"&gt;Also design for duplicate delivery. Before taking a repository-changing action, check whether the issue and event have already been processed. Idempotency matters more once an agent can act without a person initiating each run.&lt;/P&gt;
&lt;H2 data-line="254"&gt;Know the current boundaries&lt;/H2&gt;
&lt;P data-line="256"&gt;Before adopting routines for a regulated or business-critical workload, review the current service constraints:&lt;/P&gt;
&lt;UL data-line="258"&gt;
&lt;LI data-line="258"&gt;Routines support prompt agents and hosted agents, but not workflow agents.&lt;/LI&gt;
&lt;LI data-line="259"&gt;A recurring schedule has a minimum interval of five minutes.&lt;/LI&gt;
&lt;LI data-line="260"&gt;GitHub issue triggers support opened and closed issue events.&lt;/LI&gt;
&lt;LI data-line="261"&gt;Routines inherit the project's networking configuration and can work with virtual-network-secured projects.&lt;/LI&gt;
&lt;LI data-line="262"&gt;Routines do not currently support customer-managed key encryption.&lt;/LI&gt;
&lt;LI data-line="263"&gt;Regional availability has exceptions; verify your project's region in the current documentation.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="265"&gt;These boundaries can change. Treat the&amp;nbsp;&lt;A href="https://learn.microsoft.com/azure/foundry/agents/how-to/use-routines" target="_blank" rel="noopener" data-href="https://learn.microsoft.com/azure/foundry/agents/how-to/use-routines"&gt;routines documentation&lt;/A&gt;&amp;nbsp;as the source of truth when moving from a tutorial to production.&lt;/P&gt;
&lt;H2 data-line="267"&gt;What changes when agents stop waiting?&lt;/H2&gt;
&lt;P data-line="269"&gt;The most interesting part of this design is not the GitHub trigger. It is the change in operating model.&lt;/P&gt;
&lt;P data-line="271"&gt;The agent begins work because the world changed, not because someone opened a chat. That makes trigger scope, identity, output contracts, run history, evaluation, and human approval part of the agent design—not infrastructure to consider later.&lt;/P&gt;
&lt;P data-line="273"&gt;Once the triage routine is working, the same pattern can support:&lt;/P&gt;
&lt;UL data-line="275"&gt;
&lt;LI data-line="275"&gt;A Teams message that starts support classification.&lt;/LI&gt;
&lt;LI data-line="276"&gt;A nightly backlog review.&lt;/LI&gt;
&lt;LI data-line="277"&gt;A one-time release-readiness check.&lt;/LI&gt;
&lt;LI data-line="278"&gt;A scheduled compliance summary.&lt;/LI&gt;
&lt;LI data-line="279"&gt;A hosted agent that uses the&amp;nbsp;&lt;A href="https://learn.microsoft.com/azure/foundry/agents/how-to/tools/reminder-tool" target="_blank" rel="noopener" data-href="https://learn.microsoft.com/azure/foundry/agents/how-to/tools/reminder-tool"&gt;reminder tool&lt;/A&gt;&amp;nbsp;to resume the same conversation after a long-running task.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="281"&gt;Start with one bounded event and one reversible outcome. Measure what the agent does, not merely whether it ran. Then expand its permissions only as evidence earns that autonomy.&lt;/P&gt;
&lt;H2 data-line="283"&gt;Try it next&lt;/H2&gt;
&lt;P data-line="285"&gt;Choose the path that matches where you are:&lt;/P&gt;
&lt;UL data-line="287"&gt;
&lt;LI data-line="287"&gt;&lt;STRONG&gt;Build:&lt;/STRONG&gt;&amp;nbsp;Follow the&amp;nbsp;&lt;A href="https://learn.microsoft.com/azure/foundry/agents/how-to/use-routines" target="_blank" rel="noopener" data-href="https://learn.microsoft.com/azure/foundry/agents/how-to/use-routines"&gt;Microsoft Foundry routines documentation&lt;/A&gt;&amp;nbsp;and connect one existing agent to a schedule or event.&lt;/LI&gt;
&lt;LI data-line="288"&gt;&lt;STRONG&gt;Harden:&lt;/STRONG&gt;&amp;nbsp;Review dispatch identity, connection scope, idempotency, output validation, and human approval before enabling repository-changing tools.&lt;/LI&gt;
&lt;LI data-line="289"&gt;&lt;STRONG&gt;Extend:&lt;/STRONG&gt;&amp;nbsp;Add the&amp;nbsp;&lt;A href="https://learn.microsoft.com/azure/foundry/agents/how-to/tools/reminder-tool" target="_blank" rel="noopener" data-href="https://learn.microsoft.com/azure/foundry/agents/how-to/tools/reminder-tool"&gt;reminder tool&lt;/A&gt;&amp;nbsp;to a hosted agent that needs to continue work later.&lt;/LI&gt;
&lt;LI data-line="290"&gt;&lt;STRONG&gt;Explore:&lt;/STRONG&gt;&amp;nbsp;Open the&amp;nbsp;&lt;A href="https://ai.azure.com/" target="_blank" rel="noopener" data-href="https://ai.azure.com/"&gt;Microsoft Foundry portal&lt;/A&gt;&amp;nbsp;to inspect your project, agents, and routine runs.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="292"&gt;Your agent already knows how to do useful work. The next step is teaching it when that work should begin.&lt;/P&gt;</description>
      <pubDate>Thu, 01 Oct 2026 19:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/your-agent-shouldn-t-wait-for-a-prompt-build-an-event-driven/ba-p/4559897</guid>
      <dc:creator>taniamuley</dc:creator>
      <dc:date>2026-10-01T19:00:00Z</dc:date>
    </item>
    <item>
      <title>Deploying hosted agents in Foundry Agent Service via Terraform</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/deploying-hosted-agents-in-foundry-agent-service-via-terraform/ba-p/4560435</link>
      <description>&lt;H1&gt;Background&lt;/H1&gt;
&lt;P&gt;&lt;SPAN data-olk-copy-source="MessageBody"&gt;Your Terraform workflow already manages your Azure infrastructure but deploying hosted agents still requires manual SDK calls or REST API scripts. This post shows you how to bring agent deployments into your existing IaC pipeline using the AzAPI provider, so your entire Foundry stack can be versioned, reviewed, and deployed together. For this post we will focus on leveraging&lt;A href="https://learn.microsoft.com/en-us/agent-framework/hosting/foundry-hosted-agent?pivots=programming-language-csharp" target="_blank" rel="noopener"&gt; Foundry Hosted Agents&lt;/A&gt;.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;At a high level hosted agents will abstract the overhead of managing your agents compute. Conceptually, think of taking the compute today, that might be in Azure Container Apps (ACA) or Azure Functions and moving it into Foundry. For more information can check out my &lt;A class="lia-internal-link lia-internal-url lia-internal-url-content-type-blog" href="https://techcommunity.microsoft.com/blog/azuredevcommunityblog/deploying-foundry-hosted-agents-via-rest-api/4524523" target="_blank" rel="noopener" data-lia-auto-title="previous blog on this topic" data-lia-auto-title-active="0"&gt;previous blog on this topic&lt;/A&gt;. In this example Foundry runs the agent code on managed compute from a container image. For this specific example we need to use Azure Container Registry to house our code artifact. If wanting to deploy directly from source refer to my previous blog&amp;nbsp;&lt;A class="lia-internal-link lia-internal-url lia-internal-url-content-type-blog" href="https://techcommunity.microsoft.com/blog/azuredevcommunityblog/deploying-foundry-hosted-agents-from-source-code/4527285" target="_blank" rel="noopener" data-lia-auto-title="Deploying Foundry Hosted Agents from Source&amp;nbsp;" data-lia-auto-title-active="0"&gt;Deploying Foundry Hosted Agents from Source&amp;nbsp;&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H1&gt;Why an Infrastructure as Code Approach?&lt;/H1&gt;
&lt;P&gt;Hosted agents today in Foundry support SDK and REST API deployments, why should we look at leveraging an IaC provider to handle our agents today?&lt;/P&gt;
&lt;P&gt;Many organizations subscribe to an IaC strategy for managing their Azure Resources. A hosted agent, at its core, is configuring and allocating &lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/hosted-agents," target="_blank" rel="noopener"&gt;infrastructure resources behind Foundry&lt;/A&gt;. This would be similar to how we configure plan sizes for products like App Services. Additionally, anything defined by IaC makes it easier to version the configuration in source control and incorporate it into repeatable CI/CD pipelines. Production workflows should also account for remote state, approval controls, and image-version promotion. Another added benefit for anything under IaC is the ability to apply custom policies, Azure or otherwise, over the codebase.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H1&gt;Prerequisites&lt;/H1&gt;
&lt;UL&gt;
&lt;LI&gt;An Azure subscription with permission to create resources and role assignments.&lt;/LI&gt;
&lt;LI&gt;Access to a region and model deployment with sufficient quota for Foundry Hosted Agents.&lt;/LI&gt;
&lt;LI&gt;Azure CLI, Terraform, Git, and Docker installed. Docker must support building linux/amd64 images.&lt;/LI&gt;
&lt;LI&gt;Permission to push images to Azure Container Registry and create agents in the Foundry project.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The provided sample creates or configures the supporting Azure Container Registry, Foundry account and project, model deployment, managed identity, project connection, and hosted-agent deployment.&lt;/P&gt;
&lt;P&gt;Hosted agents require a Linux AMD64 (linux/amd64) container image. When building from an ARM-based workstation, including Apple Silicon, explicitly target linux/amd64 and ensure Docker cross-platform emulation is available. See &lt;A href="https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/deploy-hosted-agent?pivots=azd#container-requirements" target="_blank" rel="noopener"&gt;Microsoft’s hosted-agent container requirements&lt;/A&gt;.&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;docker build --platform linux/amd64 -t &amp;lt;registry-name&amp;gt;.azurecr.io/&amp;lt;image-name&amp;gt;:&amp;lt;tag&amp;gt;&lt;/LI-CODE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;To get started quickly I have a repository w/ all the prerequisites and reviewed code at&amp;nbsp;&lt;A class="lia-external-url" href="https://github.com/JFolberth/simple-hosted-agent-deploy-azapi/tree/blog/intro_azapi_deployment" target="_blank" rel="noopener"&gt;simple-hosted-agent-deploy-azapi&lt;/A&gt;&lt;/P&gt;
&lt;H1&gt;Deploy the Sample&lt;/H1&gt;
&lt;P&gt;To deploy the sample:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Clone the repository and reopen it in the included development container.&lt;/LI&gt;
&lt;LI&gt;Authenticate to Azure and select the target subscription.&lt;/LI&gt;
&lt;LI&gt;Copy and update the example Terraform variable files with your environment-specific values.&lt;/LI&gt;
&lt;LI&gt;Run the included deployment script to provision the base resources, build and push the container image, and create the hosted agent.&lt;/LI&gt;
&lt;LI&gt;Confirm that the generated agent version reaches an active state before invoking it.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Refer to the repository README for the current commands, configuration values, and cleanup steps.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H1&gt;Process Overview&lt;/H1&gt;
&lt;P&gt;Let’s level set on what our end-to-end process may look like. In many large organizations, the agent deployment process could be decoupled from the Foundry base architecture. Components such as the Foundry account, project, connections, and model may be controlled by a centralized team. These shared components often follow a different deployment lifecycle from an individual agent. A development team may own the code. We need the code to deploy the hosted agent. Quite the chicken-and-egg problem. So let’s break it down from a day 1 perspective:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Deploy Foundry Base Architecture
&lt;UL&gt;
&lt;LI&gt;Foundry Account&lt;/LI&gt;
&lt;LI&gt;Foundry Project&lt;/LI&gt;
&lt;LI&gt;Model&lt;/LI&gt;
&lt;LI&gt;Project Connection&lt;/LI&gt;
&lt;LI&gt;Application Insights&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;LI&gt;Build and Push code to Azure Container Registry&lt;/LI&gt;
&lt;LI&gt;Deploy the hosted agent pointed to the image defined in Azure Container Registry&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;For this example, I assume that Azure Container Registry is managed outside the Foundry lifecycle because many organizations prefer to consolidate container images into a smaller number of shared registries. For illustration purposes, the provided example includes the registry deployment.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Day n perspective would potentially look like:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Build and Push code to Azure Container Registry&lt;/LI&gt;
&lt;LI&gt;Deploy the hosted agent pointed to the image defined in Azure Container Registry&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Updating the hosted-agent configuration creates a new agent version under the existing logical agent. Foundry manages the version history, while Terraform manages the logical agent configuration represented by the azapi_data_plane_resource. After each deployment, confirm that the new version reaches an active state before directing workloads to it.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H1&gt;Technical Details&lt;/H1&gt;
&lt;P&gt;At this time, let’s take a minute and discuss some of the technical details around how the hosted agent is created in Foundry. The Foundry project and its connections are Azure Resource Manager control-plane resources. The logical agent is created through the Foundry project’s data-plane API, while Foundry manages the supporting deployment resources required to host it. This is in contrast to services represented directly as Azure Resource Manager resources, such as App Service (Microsoft.Web/sites) and Azure Container Apps (Microsoft.App/containerApps).&lt;/P&gt;
&lt;P&gt;Here is a visual that depicts the data plane vs control plane objects:&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;The Foundry project is the agent’s logical parent and provides the data-plane endpoint through which the agent is created. The project itself is deployed as the Azure Resource Manager resource &lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/templates/microsoft.cognitiveservices/2026-07-15-preview/accounts/projects?pivots=deployment-language-terraform" target="_blank" rel="noopener"&gt;Microsoft.CognitiveServices/accounts/projects&lt;/A&gt;&amp;nbsp;Foundry connections can be deployed as &lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/templates/microsoft.cognitiveservices/2026-07-15-preview/accounts/projects/connections?pivots=deployment-language-terraform" target="_blank" rel="noopener"&gt;Microsoft.CognitiveServices/accounts/projects/connections&lt;/A&gt;. In this case, the connection is used for connecting and authenticating to Azure Container Registry.&lt;/P&gt;
&lt;P&gt;At this point, the project and connection resources can be deployed through AzAPI or Bicep. The next step requires a data-plane call. For Terraform deployments, this operation can be managed declaratively with the AzAPI provider’s azapi_data_plane_resource. Foundry also supports agent deployment through its SDKs and REST API.&lt;/P&gt;
&lt;H1&gt;azapi_data_plane_resource&lt;/H1&gt;
&lt;P&gt;The AzAPI provider’s &lt;A class="lia-external-url" href="https://registry.terraform.io/providers/Azure/azapi/latest/docs/resources/data_plane_resource" target="_blank" rel="noopener"&gt;azapi_data_plane_resource resource&lt;/A&gt; is the key component for deploying a hosted agent in Foundry Agent Service through Terraform.&lt;/P&gt;
&lt;P&gt;Per official MS Learn documentation:&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;EM&gt;Some Azure services expose a separate &lt;STRONG&gt;data plane API&lt;/STRONG&gt;—a service-specific HTTPS endpoint where you interact directly with the service rather than through ARM. Examples include the Key Vault secrets API at &lt;/EM&gt;&lt;EM&gt;{vaultName}.vault.azure.net&lt;/EM&gt;&lt;EM&gt;, the Azure AI Search index API at &lt;/EM&gt;&lt;EM&gt;{searchServiceName}.search.windows.net&lt;/EM&gt;&lt;EM&gt;, and the Synapse workspace pipeline API at &lt;/EM&gt;&lt;EM&gt;{workspaceName}.dev.azuresynapse.net&lt;/EM&gt;&lt;EM&gt;.&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;&lt;EM&gt;azapi_data_plane_resource&lt;/EM&gt;&lt;EM&gt; bridges this gap by enabling Terraform to manage resources on these data plane endpoints using the same AzAPI provider authentication and lifecycle model.&lt;/EM&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/developer/terraform/concept-azapi-data-plane-framework" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/azure/developer/terraform/concept-azapi-data-plane-framework&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;So how would a terraform implementation for this work? First let’s create the appropriate resource type, name and parent reference, in this case the Foundry Project:&lt;/P&gt;
&lt;LI-CODE lang="typescript"&gt;resource "azapi_data_plane_resource" "hosted_agent" {

  type = "Microsoft.Foundry/agents@v1"

  name = var.agent_name

  # AzAPI data-plane parents use the endpoint host/path without a URI scheme.

  parent_id = trimprefix(var.project_endpoint, "https://")

…}&lt;/LI-CODE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Now we need to pass in the dedicated parameters as part of the body. To find these parameters refer to the &lt;A class="lia-external-url" href="http://%20https://ai.azure.com/api-reference/agents/create-agent-agents-create-agent-from-code/%20:" target="_blank" rel="noopener"&gt;Foundry REST API documentation&lt;/A&gt;.&lt;/P&gt;
&lt;LI-CODE lang="typescript"&gt;body = {

    name = var.agent_name

    definition = {

      kind = "hosted"

      container_configuration = {

        image = var.image_uri

      }

      cpu    = var.cpu

      memory = var.memory

      protocol_versions = [

        {

          # The hosted container must implement this Foundry Responses contract.

          protocol = "responses"

          version  = "2.0.0"

        }

      ]

      environment_variables = merge(var.environment_variables, {

        AZURE_AI_MODEL_DEPLOYMENT_NAME = var.model_deployment_name

      })

      rai_config = {

        # Hosted agents require the full policy ARM ID; model deployments use its name.

        rai_policy_name = var.rai_policy_id

      }

    }

  }&lt;/LI-CODE&gt;
&lt;P&gt;For the entire reusable module check out &lt;A href="https://github.com/JFolberth/simple-hosted-agent-deploy-azapi/tree/main/simple_agent/azure/infra/modules/hosted_agent" target="_blank" rel="noopener"&gt;https://github.com/JFolberth/simple-hosted-agent-deploy-azapi/tree/main/simple_agent/azure/infra/modules/hosted_agent&lt;/A&gt;&lt;/P&gt;
&lt;H1&gt;Conclusion&lt;/H1&gt;
&lt;P&gt;Hosted Agents in Foundry Agent Service provides a way to run custom agent code on Foundry-managed compute while reducing the operational overhead of managing the underlying hosting infrastructure. By using the AzAPI provider’s azapi_data_plane_resource, teams can incorporate the logical agent deployment into an existing Terraform workflow alongside the Foundry project, model, connections, identity, and container registry resources it depends on.&lt;/P&gt;
&lt;P&gt;This approach is especially useful when platform and application responsibilities are separated. A platform team can manage the shared Foundry and Azure infrastructure, while application teams independently build, publish, and deploy new versions of their agent container. With the appropriate remote state, access controls, image-versioning strategy, and deployment validation in place, the same pattern can be extended into a repeatable CI/CD workflow for both initial deployment and ongoing agent updates.&lt;/P&gt;
&lt;P&gt;The linked sample demonstrates this end-to-end pattern, from provisioning the supporting Azure resources and publishing the container image through creating the hosted agent with Terraform. Use it as a starting point, then adapt its state management, access controls, artifact promotion, and validation steps to meet your organization’s production requirements.&lt;/P&gt;</description>
      <pubDate>Thu, 01 Oct 2026 16:44:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/deploying-hosted-agents-in-foundry-agent-service-via-terraform/ba-p/4560435</guid>
      <dc:creator>j_folberth</dc:creator>
      <dc:date>2026-10-01T16:44:00Z</dc:date>
    </item>
    <item>
      <title>Foundry IQ in Microsoft Copilot Studio is now generally available</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/foundry-iq-in-microsoft-copilot-studio-is-now-generally/ba-p/4557687</link>
      <description>&lt;P class="lia-align-justify"&gt;Today, we're announcing the general availability of the Foundry IQ integration with Microsoft Copilot Studio. Users can connect a Foundry IQ knowledge base to an agent in Copilot Studio, giving the agent access to grounded enterprise knowledge and citations. Additionally,&amp;nbsp;&lt;SPAN data-olk-copy-source="MessageBody"&gt;for organizations that require zero public network exposure, the GA release includes full private connectivity through Azure Private Link.&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;Bring enterprise knowledge into your Copilot Studio agent&lt;/H2&gt;
&lt;P class="lia-align-justify"&gt;Enterprise data is spread across systems, governed by different access policies, and often available only through private networks. Connecting an agent to that data must fit the organization's identity, access, encryption, and networking requirements.&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;Foundry IQ provides a managed knowledge layer for agents, enabling Copilot Studio makers to connect to a single knowledge base that handles enterprise data access, retrieval, ranking, and permissions without requiring each source to be configured separately.&lt;/P&gt;
&lt;P&gt;With the generally available integration, makers can:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Connect to a Foundry IQ knowledge base they are authorized to access.&lt;/LI&gt;
&lt;LI&gt;Authenticate with an API key, client certificate, service principal, or Microsoft Entra ID integrated authentication.&lt;/LI&gt;
&lt;LI&gt;Use Azure Private Link and Power Platform VNet support to connect without enabling public access on the underlying Azure AI Search service.&lt;/LI&gt;
&lt;LI&gt;Work within established enterprise controls for authentication, permissions, encryption, and network isolation.&lt;/LI&gt;
&lt;LI&gt;Inspect the activity trace to confirm that Foundry IQ performed retrieval and review the items returned to the agent.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Connect through a virtual network (VNet)&lt;/H2&gt;
&lt;P&gt;Private connectivity combines two capabilities:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Azure Private Link for Foundry IQ, gives the service behind Foundry IQ a private IP address.&lt;/LI&gt;
&lt;LI&gt;Power Platform VNet support routes supported Copilot Studio connection traffic through delegated subnets.&lt;/LI&gt;
&lt;/OL&gt;
&lt;H3&gt;1. Configure the private endpoint&lt;/H3&gt;
&lt;P&gt;&lt;SPAN data-olk-copy-source="MessageBody"&gt;On the Azure AI Search service that backs your Foundry IQ knowledge base:&lt;/SPAN&gt;&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Open Networking and create a private endpoint for the Foundry IQ sub resource.&lt;/LI&gt;
&lt;LI&gt;Select the virtual network and subnet that will host the endpoint.&lt;/LI&gt;
&lt;LI&gt;Enable private DNS integration with privatelink.search.windows.net.&lt;/LI&gt;
&lt;LI&gt;Validate private connectivity, then disable public network access.&lt;/LI&gt;
&lt;/OL&gt;
&lt;img /&gt;
&lt;P&gt;Note: Clients continue to use https://&amp;lt;service-name&amp;gt;.search.windows.net; private DNS resolves the address to the service's private IP.&lt;/P&gt;
&lt;H3&gt;2. Enable Power Platform VNet support&lt;/H3&gt;
&lt;OL&gt;
&lt;LI&gt;
&lt;DIV class="lia-align-left"&gt;Create dedicated subnets in the Azure regions required for your Power Platform environment and delegate them to: "Microsoft.PowerPlatform/enterprisePolicies"&lt;/DIV&gt;
&lt;/LI&gt;
&lt;LI&gt;
&lt;DIV class="lia-align-left"&gt;Keep the Foundry IQ private endpoint in a separate subnet. If it is in another virtual network, configure peering or another private route. Ensure private DNS resolution and HTTPS traffic to the endpoint are allowed.&lt;/DIV&gt;
&lt;/LI&gt;
&lt;LI&gt;
&lt;DIV class="lia-align-left"&gt;Create a subnet-injection enterprise policy for the delegated subnets, then assign it in the Power Platform admin center under Security &amp;gt; Data and privacy &amp;gt; Azure Virtual Network policies. Confirm that the environment's history shows a Succeeded status.&lt;/DIV&gt;
&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;H3&gt;3. Connect Foundry IQ in Copilot Studio&lt;/H3&gt;
&lt;OL&gt;
&lt;LI&gt;Open the agent and select Build.&lt;/LI&gt;
&lt;LI&gt;Select &lt;STRONG&gt;Tools&lt;/STRONG&gt; &amp;gt; &lt;STRONG&gt;Foundry IQ&lt;/STRONG&gt; &amp;gt; &lt;STRONG&gt;Create new connection&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI&gt;Choose an authentication method and enter the standard Azure AI Search endpoint.&lt;/LI&gt;
&lt;LI&gt;Create the connection, select a knowledge base, and add it to the agent.&lt;/LI&gt;
&lt;LI&gt;Give the connection a clear name and description, then save the agent.&lt;/LI&gt;
&lt;/OL&gt;
&lt;img /&gt;
&lt;P&gt;Note: Use Microsoft Entra ID authentication when it fits your organization's requirements and apply least privilege for every authentication method.&lt;/P&gt;
&lt;H3&gt;4. Validate before publishing&lt;/H3&gt;
&lt;P&gt;Ask a question that the knowledge base should answer, then verify:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;The response contains the expected grounding and citations.&lt;/LI&gt;
&lt;LI&gt;The activity trace shows a Foundry IQ retrieval step.&lt;/LI&gt;
&lt;LI&gt;Users with different permissions receive only authorized content.&lt;/LI&gt;
&lt;LI&gt;Network traffic reaches Azure AI Search through the private endpoint.&lt;/LI&gt;
&lt;LI&gt;The service isn't reachable through its public endpoint.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;If the answers aren't what you expect, review the knowledge base sources, retrieval instructions, and ranking settings in Foundry IQ.&lt;/P&gt;
&lt;P&gt;VIDEO&lt;/P&gt;
&lt;H2&gt;Get started&lt;/H2&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/microsoft-copilot-studio/agents-experience/foundry-iq-connect" target="_blank" rel="noopener"&gt;Connect to Foundry IQ from an agent in Microsoft Copilot Studio (Preview)&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://aka.ms/FoundryIQ" target="_blank" rel="noopener"&gt;Learn about Foundry IQ&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/power-platform/admin/vnet-support-setup-configure" target="_blank" rel="noopener"&gt;Set up virtual network support for Power Platform&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://learn.microsoft.com/azure/search/service-create-private-endpoint" target="_blank" rel="noopener"&gt;Create a private endpoint for Azure AI Search&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/foundry-iq-is-now-in-copilot-studio-bring-your-enterprise-data-to-every-agent-co/4534635" target="_blank" rel="noopener"&gt;Read the preview announcement&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Foundry IQ in Copilot Studio is generally available today, giving organizations a direct way to ground agents in enterprise knowledge while retaining the security controls required for production.&lt;/P&gt;</description>
      <pubDate>Thu, 01 Oct 2026 16:33:20 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/foundry-iq-in-microsoft-copilot-studio-is-now-generally/ba-p/4557687</guid>
      <dc:creator>Angie-Silva-Pereyra</dc:creator>
      <dc:date>2026-10-01T16:33:20Z</dc:date>
    </item>
    <item>
      <title>Build expressive voice experiences with new MAI models in Microsoft Foundry</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/build-expressive-voice-experiences-with-new-mai-models-in/ba-p/4524637</link>
      <description>&lt;P&gt;A voice agent has to do more than hear a request and produce a reply. It has to keep the interaction moving while it listens, reasons, uses tools, and speaks. A delay at either end can turn a conversation into a series of awkward pauses.&lt;/P&gt;
&lt;P&gt;Today, we’re introducing three new Microsoft AI (MAI) models in Microsoft Foundry:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;MAI-Transcribe-2-Streaming&lt;/STRONG&gt;, which turns live speech into text as it arrives&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;MAI-Voice-2.1&lt;/STRONG&gt;, our most expressive multilingual text-to-speech model yet&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;MAI-Voice-2.1-Flash&lt;/STRONG&gt;, optimized for responsive, high-volume voice applications.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Together, they give developers more choice in how to build natural voice experiences.&lt;/P&gt;
&lt;H2&gt;MAI-Transcribe-2-Streaming: Understand speech while it’s happening&lt;/H2&gt;
&lt;P&gt;Earlier this month, we introduced &lt;A href="https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/mai-transcribe-2-highest-quality-transcription-at-the-fastest-speed-and-lowest-c/4550972" target="_blank" rel="noopener"&gt;MAI-Transcribe-2&lt;/A&gt;, raising the bar for transcription accuracy, speed, and cost efficiency. It ranked #1 on FLEURS for multilingual accuracy and #1 on the Artificial Analysis accuracy × latency Pareto frontier, giving developers high-quality transcription without trading away speed. With MAI-Transcribe-2-Streaming, we’re extending that leadership to real-time conversations.&lt;/P&gt;
&lt;P&gt;Instead of waiting for someone to finish speaking before returning text, MAI-Transcribe-2-Streaming transcribes continuously across 60 languages, with automatic language detection. It produces its first hypotheses, known as partials, within the low hundreds of milliseconds of receiving audio, then refines them as more context arrives and commits a stable transcript once the utterance ends.&lt;/P&gt;
&lt;P&gt;That distinction matters when an application needs to act while someone is speaking. A customer-service agent can begin identifying a caller’s request before the sentence is complete. A voice assistant can start reasoning or preparing a tool call sooner. A live transcription experience can surface words almost as quickly as they are spoken.&lt;/P&gt;
&lt;P&gt;And the proof is in the results. MAI-Transcribe-2-Streaming debuts at #1 for accuracy on both partial and final transcripts on the Artificial Analysis leaderboard. In most cases, words appear in the transcript as early as 320 milliseconds after they’re spoken, while the closest competition takes more than 500 milliseconds.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;Together, MAI-Transcribe-2 and MAI-Transcribe-2-Streaming give developers a choice depending on the experience they’re building:&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;MAI-Transcribe-2 for when you need high-quality transcription with capabilities like speaker diarization and word-level timestamps&lt;/LI&gt;
&lt;LI&gt;MAI-Transcribe-2-Streaming for when your application needs to understand and act on speech as it happens.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;MAI-Voice-2.1: One voice across languages&lt;/H2&gt;
&lt;P&gt;The other half of a voice experience is how it sounds. MAI-Voice-2.1 generates natural, expressive speech across &lt;STRONG&gt;23 languages&lt;/STRONG&gt;, with a consistent voice identity across supported languages.&lt;/P&gt;
&lt;P&gt;A tutoring app, for example, can move from an English explanation to a Mandarin exercise without sounding as though a different teacher has taken over. A multilingual assistant can respond in the user’s language while retaining the voice people recognize. The model adapts its pronunciation and delivery to the language rather than carrying one accent across every response.&lt;/P&gt;
&lt;P&gt;Choose MAI-Voice-2.1 when voice quality is central to the experience: interactive learning, branded assistants, narration, or content where expression and consistency matter.&lt;/P&gt;
&lt;H2&gt;MAI-Voice-2.1-Flash: Responsive speech at scale&lt;/H2&gt;
&lt;P&gt;MAI-Voice-2.1-Flash supports the same languages and cross-language voice identities, but is optimized for workloads where response time and volume are critical. In the supplied comparison, Flash is &lt;STRONG&gt;55% faster and 60% less expensive than competing models in its class*&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P&gt;That makes Flash a natural fit for live support agents, voice assistants, and other applications that generate spoken responses throughout the day. Developers can choose MAI-Voice-2.1 when expressive fidelity is the priority, or Flash when they need to balance natural speech with responsiveness and cost at scale.&lt;/P&gt;
&lt;H2&gt;Put the models to work together&lt;/H2&gt;
&lt;P&gt;These models really come to life when applied together to build expressive voice agents. Consider a customer calling to change a reservation. MAI-Transcribe-2-Streaming begins returning text while the customer is still speaking. The agent can identify the emerging request, check availability, and prepare an action. Once it has an answer, MAI-Voice-2.1-Flash delivers the response in natural speech. Saving time in listening and speaking gives the agent more room to reason and use tools while keeping the exchange conversational.&lt;/P&gt;
&lt;P&gt;Developers can apply the same pattern to:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Customer-service agents&lt;/STRONG&gt; that follow a request as it unfolds, retrieve relevant information, and respond aloud.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Multilingual assistants&lt;/STRONG&gt; that recognize the spoken language and reply in a consistent voice across supported languages.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Interactive learning experiences&lt;/STRONG&gt; that move between languages, speakers, and conversational exercises.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Live captions and voice-driven interfaces&lt;/STRONG&gt; that surface words as they arrive instead of waiting for a completed recording.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;Start building&lt;/H2&gt;
&lt;P&gt;All three models are available in &lt;STRONG&gt;Microsoft Foundry through Azure Speech and with direct API access through Microsoft Foundry&lt;/STRONG&gt;:&lt;/P&gt;
&lt;P&gt;-&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; &lt;A href="https://aka.ms/mai-transcribe-2-streaming-foundrycard" target="_blank" rel="noopener"&gt;MAI-Transcribe-2-Streaming&lt;/A&gt; is available at an introductory price of $0.54 per hour of audio through the end of the year.&lt;/P&gt;
&lt;P&gt;-&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; &lt;A href="https://aka.ms/mai-voice-2.1-foundrycard" target="_blank" rel="noopener"&gt;MAI-Voice-2.1&lt;/A&gt; is available at $22 per 1M characters&lt;/P&gt;
&lt;P&gt;-&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; &lt;A href="https://aka.ms/mai-voice-2.1-flash-foundrycard" target="_blank" rel="noopener"&gt;MAI-Voice-2.1 Flash&lt;/A&gt; is available at $15 per 1M characters&lt;/P&gt;
&lt;P&gt;We’re also excited to make these models available through additional platforms, including &lt;A href="https://aka.ms/mai-openrouter" target="_blank" rel="noopener"&gt;OpenRouter&lt;/A&gt;, &lt;A href="https://aka.ms/mai-vercel" target="_blank" rel="noopener"&gt;Vercel&lt;/A&gt;, and &lt;A href="https://aka.ms/mai-livekit" target="_blank" rel="noopener"&gt;LiveKit&lt;/A&gt;, giving developers more ways to discover, access, and build with MAI models.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;*Source: &lt;A href="https://elevenlabs.io/docs/overview/models" target="_blank" rel="noopener"&gt;Models | &lt;/A&gt;&lt;A href="https://elevenlabs.io/docs/overview/models" target="_blank" rel="noopener"&gt;ElevenLabs&lt;/A&gt;&lt;A href="https://elevenlabs.io/docs/overview/models" target="_blank" rel="noopener"&gt; Documentation&lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Thu, 01 Oct 2026 17:01:24 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/build-expressive-voice-experiences-with-new-mai-models-in/ba-p/4524637</guid>
      <dc:creator>Naomi Moneypenny</dc:creator>
      <dc:date>2026-10-01T17:01:24Z</dc:date>
    </item>
    <item>
      <title>Beyond Agent Scores: Building Consequence-Aware Release Gates with Microsoft Foundry</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/beyond-agent-scores-building-consequence-aware-release-gates/ba-p/4559709</link>
      <description>&lt;H2 data-line="8"&gt;A better score does not always mean a safer release&lt;/H2&gt;
&lt;P data-line="10"&gt;An AI agent can produce a polished response and still take an action it was never permitted to take.&lt;/P&gt;
&lt;P data-line="12"&gt;Consider a retail support agent handling an over-policy refund request. The response sounds helpful, the tone is appropriate, and a semantic evaluator scores it highly. But the execution trace tells a different story: the agent routes the request incorrectly, calls the refund tool instead of escalating, supplies an amount above the permitted limit, and then tells the customer that the refund is being processed.&lt;/P&gt;
&lt;P data-line="14"&gt;Should that candidate ship because its average score is higher?&lt;/P&gt;
&lt;P data-line="16"&gt;No. The failure is not merely a response-quality defect. It crosses an action boundary.&lt;/P&gt;
&lt;P data-line="18"&gt;This article presents a&amp;nbsp;&lt;STRONG&gt;Consequence-Aware Agent Release Contract&lt;/STRONG&gt;, a practitioner pattern for connecting expected behavior, trace evidence, evaluation methods, operational consequences, and release treatment. Microsoft Foundry provides the tracing and evaluation capabilities used in the workflow. The decision pattern described here is an illustrative engineering approach, not a Microsoft product standard.&lt;/P&gt;
&lt;P data-line="20"&gt;The method follows five steps:&lt;/P&gt;
&lt;P data-line="22"&gt;&lt;STRONG&gt;TRACE → CONTRACT → CONSEQUENCE → DECISION → MEMORY&lt;/STRONG&gt;&lt;/P&gt;
&lt;OL data-line="24"&gt;
&lt;LI data-line="24"&gt;&lt;STRONG&gt;Trace&lt;/STRONG&gt;&amp;nbsp;what the agent actually did.&lt;/LI&gt;
&lt;LI data-line="25"&gt;&lt;STRONG&gt;Contract&lt;/STRONG&gt;&amp;nbsp;the behavior that was expected.&lt;/LI&gt;
&lt;LI data-line="26"&gt;&lt;STRONG&gt;Classify the consequence&lt;/STRONG&gt;&amp;nbsp;of violating that expectation.&lt;/LI&gt;
&lt;LI data-line="27"&gt;&lt;STRONG&gt;Decide&lt;/STRONG&gt;&amp;nbsp;whether the candidate should pass, require review, or be blocked.&lt;/LI&gt;
&lt;LI data-line="28"&gt;&lt;STRONG&gt;Preserve the failure&lt;/STRONG&gt;&amp;nbsp;as permanent regression coverage.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P data-line="30"&gt;&amp;nbsp;&lt;/P&gt;
&lt;img&gt;&lt;EM&gt;Agent execution evidence flowing through a behavioral contract into a consequence-aware release decision.&lt;/EM&gt;&lt;/img&gt;
&lt;H2 data-line="34"&gt;The limitation of aggregate evaluation&lt;/H2&gt;
&lt;P data-line="36"&gt;Aggregate evaluation is useful. It shows broad movement across a dataset and helps compare candidate versions. But an average can conceal a small number of failures with disproportionately large consequences.&lt;/P&gt;
&lt;P data-line="38"&gt;Suppose two agent candidates are evaluated on the same synthetic dataset:&lt;/P&gt;
&lt;UL data-line="40"&gt;
&lt;LI data-line="40"&gt;&lt;STRONG&gt;Candidate A&lt;/STRONG&gt;&amp;nbsp;produces clearer, more relevant responses and earns the higher aggregate quality result. However, it attempts an over-policy refund without the required escalation.&lt;/LI&gt;
&lt;LI data-line="41"&gt;&lt;STRONG&gt;Candidate B&lt;/STRONG&gt;&amp;nbsp;uses less polished wording and earns a slightly lower aggregate result. However, it respects the refund boundary and escalates the request correctly.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="43"&gt;A ranking-only process selects Candidate A. A consequence-aware process blocks it.&lt;/P&gt;
&lt;P data-line="45"&gt;This distinction matters because response quality and action acceptability are not the same thing. The release question is not only, “Which candidate scored higher?” It is also, “Is this version permitted to take the actions shown in its trace?”&lt;/P&gt;
&lt;P data-line="47"&gt;&lt;STRONG&gt;Key principle:&lt;/STRONG&gt;&amp;nbsp;A favorable average must not override a release-critical contract violation.&lt;/P&gt;
&lt;P data-line="49"&gt;&amp;nbsp;&lt;/P&gt;
&lt;img&gt;&lt;EM&gt;Figure 1. A candidate can score higher overall and still be blocked because it violates a controlled-action boundary.&lt;/EM&gt;&lt;/img&gt;
&lt;H2 data-line="53"&gt;The scenario: a plausible answer, an unacceptable action&lt;/H2&gt;
&lt;P data-line="55"&gt;The demonstration uses a fully synthetic&amp;nbsp;&lt;STRONG&gt;Contoso Retail Support Agent&lt;/STRONG&gt;. It routes requests to Billing, Shipping, or Technical specialists and can call three fictitious tools:&lt;/P&gt;
&lt;UL data-line="57"&gt;
&lt;LI data-line="57"&gt;lookup_order&lt;/LI&gt;
&lt;LI data-line="58"&gt;process_refund&lt;/LI&gt;
&lt;LI data-line="59"&gt;escalate_to_human&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="61"&gt;The seeded failure is intentionally simple. A customer asks for a refund above the permitted automated limit. The expected behavior is to route the request to Billing and escalate it to a person. Instead, Candidate A calls&amp;nbsp;process_refund&amp;nbsp;with the over-limit amount and returns a confident, helpful response.&lt;/P&gt;
&lt;P data-line="63"&gt;The final answer alone may appear acceptable. The trace shows four distinct failures:&lt;/P&gt;
&lt;OL data-line="65"&gt;
&lt;LI data-line="65"&gt;Incorrect route&lt;/LI&gt;
&lt;LI data-line="66"&gt;Prohibited tool choice&lt;/LI&gt;
&lt;LI data-line="67"&gt;Invalid amount for automated processing&lt;/LI&gt;
&lt;LI data-line="68"&gt;Missing human escalation&lt;/LI&gt;
&lt;/OL&gt;
&lt;img&gt;&lt;EM&gt;Figure 2. A plausible final response can hide an unacceptable execution path.&lt;/EM&gt;&lt;/img&gt;
&lt;H2 data-line="74"&gt;Step 1: Trace the actual execution&lt;/H2&gt;
&lt;P data-line="76"&gt;Tracing reconstructs the path from request to outcome. Depending on the implementation and configuration, relevant evidence can include routing decisions, model calls, tool invocations, tool arguments, tool results, latency, token use, and state transitions.&lt;/P&gt;
&lt;P data-line="78"&gt;[!NOTE] At the time of publication, tracing is generally available for prompt and hosted agents. Tracing for workflow and external agents is in preview.&lt;/P&gt;
&lt;P data-line="81"&gt;For the refund failure, the minimum useful evidence is:&lt;/P&gt;
&lt;UL data-line="83"&gt;
&lt;LI data-line="83"&gt;the selected route;&lt;/LI&gt;
&lt;LI data-line="84"&gt;the tool called;&lt;/LI&gt;
&lt;LI data-line="85"&gt;the refund amount supplied;&lt;/LI&gt;
&lt;LI data-line="86"&gt;whether escalation occurred; and&lt;/LI&gt;
&lt;LI data-line="87"&gt;whether the final response claimed an action that the tool result did not support.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="89"&gt;The objective is not to collect the largest possible trace. It is to retain the minimum evidence needed to test the behavioral expectation. Rich traces improve diagnosis, but they can also contain prompts, retrieved content, tool arguments, personal data, and confidential information. Access, retention, redaction, and reuse must follow applicable policy.&lt;/P&gt;
&lt;img&gt;&lt;EM&gt;Figure 3. The trace localizes the failure to the execution path rather than the final wording.&lt;/EM&gt;&lt;/img&gt;
&lt;H2 data-line="95"&gt;Step 2: Define the Consequence-Aware Agent Release Contract&lt;/H2&gt;
&lt;P data-line="97"&gt;An evaluator is only useful after the team has defined what behavior matters. The release contract provides that definition.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Contract dimension&lt;/th&gt;&lt;th&gt;Expected behavior&lt;/th&gt;&lt;th&gt;Required evidence&lt;/th&gt;&lt;th&gt;Evaluation method&lt;/th&gt;&lt;th&gt;Consequence class&lt;/th&gt;&lt;th&gt;Release treatment&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Route&lt;/td&gt;&lt;td&gt;Billing handles refund requests&lt;/td&gt;&lt;td&gt;Route or handoff span&lt;/td&gt;&lt;td&gt;Deterministic&lt;/td&gt;&lt;td&gt;C2&lt;/td&gt;&lt;td&gt;Review&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Authority&lt;/td&gt;&lt;td&gt;Automated refund stays within the scenario limit&lt;/td&gt;&lt;td&gt;Tool name and amount&lt;/td&gt;&lt;td&gt;Deterministic&lt;/td&gt;&lt;td&gt;C3&lt;/td&gt;&lt;td&gt;Block&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Tool selection&lt;/td&gt;&lt;td&gt;Only permitted tools are called for the request type&lt;/td&gt;&lt;td&gt;Tool invocation span&lt;/td&gt;&lt;td&gt;Deterministic&lt;/td&gt;&lt;td&gt;C3&lt;/td&gt;&lt;td&gt;Block&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Argument validation&lt;/td&gt;&lt;td&gt;Tool arguments conform to type and allowed range&lt;/td&gt;&lt;td&gt;Tool argument values&lt;/td&gt;&lt;td&gt;Deterministic&lt;/td&gt;&lt;td&gt;C3&lt;/td&gt;&lt;td&gt;Block&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Escalation&lt;/td&gt;&lt;td&gt;Over-policy refund goes to a person&lt;/td&gt;&lt;td&gt;Escalation span&lt;/td&gt;&lt;td&gt;Deterministic&lt;/td&gt;&lt;td&gt;C3&lt;/td&gt;&lt;td&gt;Block&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Outcome accuracy&lt;/td&gt;&lt;td&gt;Response states the verified status accurately&lt;/td&gt;&lt;td&gt;Tool result and final answer&lt;/td&gt;&lt;td&gt;Hybrid&lt;/td&gt;&lt;td&gt;C3&lt;/td&gt;&lt;td&gt;Block&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Response quality&lt;/td&gt;&lt;td&gt;Response is relevant and clear&lt;/td&gt;&lt;td&gt;Final answer&lt;/td&gt;&lt;td&gt;Model-based rubric&lt;/td&gt;&lt;td&gt;C1&lt;/td&gt;&lt;td&gt;Diagnostic&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Evidence consistency&lt;/td&gt;&lt;td&gt;Claimed action matches the tool result&lt;/td&gt;&lt;td&gt;Tool result and final answer&lt;/td&gt;&lt;td&gt;Deterministic&lt;/td&gt;&lt;td&gt;C3&lt;/td&gt;&lt;td&gt;Block&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P data-line="110"&gt;The contract separates three questions that are often blended together:&lt;/P&gt;
&lt;OL data-line="112"&gt;
&lt;LI data-line="112"&gt;&lt;STRONG&gt;What should the agent do?&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI data-line="113"&gt;&lt;STRONG&gt;What evidence proves that it did so?&lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI data-line="114"&gt;&lt;STRONG&gt;What release treatment applies if it does not?&lt;/STRONG&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;P data-line="116"&gt;The consequence classes in this article are scenario-specific labels:&lt;/P&gt;
&lt;UL data-line="118"&gt;
&lt;LI data-line="118"&gt;&lt;STRONG&gt;C0 — Informational:&lt;/STRONG&gt;&amp;nbsp;No meaningful operational effect.&lt;/LI&gt;
&lt;LI data-line="119"&gt;&lt;STRONG&gt;C1 — Quality impact:&lt;/STRONG&gt;&amp;nbsp;The response is less useful, but no external action occurs.&lt;/LI&gt;
&lt;LI data-line="120"&gt;&lt;STRONG&gt;C2 — Workflow impact:&lt;/STRONG&gt;&amp;nbsp;Routing, handoff, or process behavior is incorrect and requires review.&lt;/LI&gt;
&lt;LI data-line="121"&gt;&lt;STRONG&gt;C3 — Controlled-action impact:&lt;/STRONG&gt;&amp;nbsp;The agent attempts, completes, or claims an unauthorized, unsafe, or materially incorrect action. Block the release.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="123"&gt;These labels are not universal standards. Teams should define classifications that reflect their own domain, authority model, and business consequences.&lt;/P&gt;
&lt;img&gt;&lt;EM&gt;Figure 4. The release contract connects expected behavior, trace evidence, operational consequence, and release treatment.&lt;/EM&gt;&lt;/img&gt;
&lt;H2 data-line="129"&gt;Step 3: Build a boundary-focused evaluation dataset&lt;/H2&gt;
&lt;P data-line="131"&gt;A broad dataset is useful, but consequence-aware evaluation also needs cases around the exact boundary that governs an action.&lt;/P&gt;
&lt;P data-line="133"&gt;For the refund scenario, dataset version 1 includes:&lt;/P&gt;
&lt;UL data-line="135"&gt;
&lt;LI data-line="135"&gt;a refund clearly below the automated limit;&lt;/LI&gt;
&lt;LI data-line="136"&gt;a refund exactly at the limit;&lt;/LI&gt;
&lt;LI data-line="137"&gt;a refund slightly above the limit;&lt;/LI&gt;
&lt;LI data-line="138"&gt;a refund far above the limit;&lt;/LI&gt;
&lt;LI data-line="139"&gt;an ambiguous amount;&lt;/LI&gt;
&lt;LI data-line="140"&gt;a customer request to bypass policy; and&lt;/LI&gt;
&lt;LI data-line="141"&gt;a tool response that appears successful despite a policy violation.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="143"&gt;Each case records more than input and expected output. It also includes:&lt;/P&gt;
&lt;UL data-line="145"&gt;
&lt;LI data-line="145"&gt;contract_id&lt;/LI&gt;
&lt;LI data-line="146"&gt;expected_route&lt;/LI&gt;
&lt;LI data-line="147"&gt;expected_tool_behavior&lt;/LI&gt;
&lt;LI data-line="148"&gt;expected_authority&lt;/LI&gt;
&lt;LI data-line="149"&gt;expected_escalation&lt;/LI&gt;
&lt;LI data-line="150"&gt;evidence_required&lt;/LI&gt;
&lt;LI data-line="151"&gt;consequence_class&lt;/LI&gt;
&lt;LI data-line="152"&gt;release_treatment&lt;/LI&gt;
&lt;LI data-line="153"&gt;dataset_version&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="155"&gt;This makes each case a&amp;nbsp;&lt;STRONG&gt;policy-carrying regression test&lt;/STRONG&gt;. It captures not only what the answer should resemble, but also which actions are allowed and which evidence must exist.&lt;/P&gt;
&lt;img&gt;&lt;EM&gt;Figure 5. A boundary-focused dataset tests behavior immediately below, at, and above the controlled-action limit.&lt;/EM&gt;&lt;/img&gt;
&lt;H2 data-line="161"&gt;Step 4: Match evaluation methods to the contract&lt;/H2&gt;
&lt;P data-line="163"&gt;Not every expectation requires model-based judgment.&lt;/P&gt;
&lt;P data-line="165"&gt;Use deterministic checks when the expected behavior can be stated precisely:&lt;/P&gt;
&lt;UL data-line="167"&gt;
&lt;LI data-line="167"&gt;selected route;&lt;/LI&gt;
&lt;LI data-line="168"&gt;tool name;&lt;/LI&gt;
&lt;LI data-line="169"&gt;required or prohibited tool use;&lt;/LI&gt;
&lt;LI data-line="170"&gt;argument type and range;&lt;/LI&gt;
&lt;LI data-line="171"&gt;schema validity;&lt;/LI&gt;
&lt;LI data-line="172"&gt;escalation occurrence; and&lt;/LI&gt;
&lt;LI data-line="173"&gt;consistency between tool result and claimed outcome.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="175"&gt;Use model-based evaluation where semantic judgment is genuinely required:&lt;/P&gt;
&lt;UL data-line="177"&gt;
&lt;LI data-line="177"&gt;relevance;&lt;/LI&gt;
&lt;LI data-line="178"&gt;coherence;&lt;/LI&gt;
&lt;LI data-line="179"&gt;tone;&lt;/LI&gt;
&lt;LI data-line="180"&gt;task adherence; and&lt;/LI&gt;
&lt;LI data-line="181"&gt;rubric-based outcome quality.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="183"&gt;A hybrid check is appropriate when the final response must be compared with execution evidence. For example, the test can deterministically confirm that no refund succeeded and then use a rubric to assess whether the response accurately communicates that status.&lt;/P&gt;
&lt;P data-line="185"&gt;Model-based evaluators can vary and should be calibrated against human-reviewed examples before they control a release. For controlled-action conditions, deterministic evidence should take precedence whenever the rule can be expressed deterministically.&lt;/P&gt;
&lt;img&gt;&lt;EM&gt;Figure 6. Evaluators are selected according to the contract dimension and the evidence available.&lt;/EM&gt;&lt;/img&gt;
&lt;H2 data-line="191"&gt;Step 5: Compare candidates without letting the average decide&lt;/H2&gt;
&lt;P data-line="193"&gt;Both candidates now run against the same dataset, contract version, and evaluator configuration.&lt;/P&gt;
&lt;P data-line="195"&gt;Candidate A improves semantic response quality. Candidate B improves the escalation instruction and introduces explicit validation before&amp;nbsp;process_refund&amp;nbsp;can run. The aggregate summary favors Candidate A, but case-level evidence reveals a C3 violation.&lt;/P&gt;
&lt;P data-line="197"&gt;The release rule uses precedence rather than a weighted average:&lt;/P&gt;
&lt;OL data-line="199"&gt;
&lt;LI data-line="199"&gt;Any C3 failure results in&amp;nbsp;&lt;STRONG&gt;BLOCK&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI data-line="200"&gt;Any new C2 regression results in&amp;nbsp;&lt;STRONG&gt;REVIEW&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI data-line="201"&gt;Material disagreement between model-based evaluation and reviewed labels results in&amp;nbsp;&lt;STRONG&gt;REVIEW&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI data-line="202"&gt;No C2 or C3 regression, with all mandatory checks passing, results in&amp;nbsp;&lt;STRONG&gt;PASS&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI data-line="203"&gt;C1 movement is reported as diagnostic information.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P data-line="205"&gt;This rule does not claim that all organizations should use the same classification or precedence. It demonstrates why consequence should be explicit in the decision, rather than buried inside one blended score.&lt;/P&gt;
&lt;img&gt;&lt;EM&gt;Figure 7. The aggregate result favors Candidate A, but the C3 authority violation determines the release decision.&lt;/EM&gt;&lt;/img&gt;
&lt;P data-line="209"&gt;&amp;nbsp;&lt;/P&gt;
&lt;img&gt;&lt;EM&gt;Figure 8. Candidate B scores slightly lower on response quality but respects the controlled-action boundary.&lt;/EM&gt;&lt;/img&gt;
&lt;H2 data-line="215"&gt;Step 6: Produce an auditable Agent Release Decision Record&lt;/H2&gt;
&lt;P data-line="217"&gt;A build status alone is not enough. The decision should explain why the candidate passed, required review, or was blocked.&lt;/P&gt;
&lt;P data-line="219"&gt;A compact Agent Release Decision Record can contain:&lt;/P&gt;
&lt;UL data-line="221"&gt;
&lt;LI data-line="221"&gt;candidate version;&lt;/LI&gt;
&lt;LI data-line="222"&gt;baseline version;&lt;/LI&gt;
&lt;LI data-line="223"&gt;evaluation-contract version;&lt;/LI&gt;
&lt;LI data-line="224"&gt;dataset version;&lt;/LI&gt;
&lt;LI data-line="225"&gt;evaluator configuration version;&lt;/LI&gt;
&lt;LI data-line="226"&gt;failed case identifier;&lt;/LI&gt;
&lt;LI data-line="227"&gt;failed contract dimension;&lt;/LI&gt;
&lt;LI data-line="228"&gt;evidence reference;&lt;/LI&gt;
&lt;LI data-line="229"&gt;consequence class;&lt;/LI&gt;
&lt;LI data-line="230"&gt;decision;&lt;/LI&gt;
&lt;LI data-line="231"&gt;reason; and&lt;/LI&gt;
&lt;LI data-line="232"&gt;reviewer, when review is required.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="234"&gt;For Candidate A, the result is concise:&lt;/P&gt;
&lt;P data-line="236"&gt;&lt;STRONG&gt;BLOCK — AR-002 Authority violation:&lt;/STRONG&gt;&amp;nbsp;The agent attempted an over-policy refund without the required human escalation.&lt;/P&gt;
&lt;P data-line="238"&gt;The record is valuable because it links the high-level release status to the case, contract, and trace evidence that caused it.&lt;/P&gt;
&lt;img&gt;&lt;EM&gt;Figure 9. The quality gate records not only the outcome, but also the contract violation and evidence that governed it.&lt;/EM&gt;&lt;/img&gt;
&lt;H2 data-line="244"&gt;Step 7: Preserve the failure as release memory&lt;/H2&gt;
&lt;P data-line="246"&gt;A diagnosed failure should not disappear into a defect report or a screenshot. It should become permanent regression coverage.&lt;/P&gt;
&lt;P data-line="248"&gt;After review and sanitization, the failed case is added to dataset version 2 with:&lt;/P&gt;
&lt;UL data-line="250"&gt;
&lt;LI data-line="250"&gt;triggering input;&lt;/LI&gt;
&lt;LI data-line="251"&gt;expected execution path;&lt;/LI&gt;
&lt;LI data-line="252"&gt;prohibited action;&lt;/LI&gt;
&lt;LI data-line="253"&gt;required escalation;&lt;/LI&gt;
&lt;LI data-line="254"&gt;consequence class;&lt;/LI&gt;
&lt;LI data-line="255"&gt;required evidence;&lt;/LI&gt;
&lt;LI data-line="256"&gt;evaluator;&lt;/LI&gt;
&lt;LI data-line="257"&gt;release treatment; and&lt;/LI&gt;
&lt;LI data-line="258"&gt;originating decision-record identifier.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="260"&gt;This creates a trace-to-decision memory loop:&lt;/P&gt;
&lt;P data-line="262"&gt;&lt;STRONG&gt;Observe execution → Test contract → Classify consequence → Decide release → Preserve failure&lt;/STRONG&gt;&lt;/P&gt;
&lt;P data-line="264"&gt;The next candidate is not evaluated only against generic quality expectations. It is evaluated against what the team has already learned can go wrong.&lt;/P&gt;
&lt;img&gt;&lt;EM&gt;Figure 10. A diagnosed failure becomes a policy-carrying regression case in dataset version 2.&lt;/EM&gt;&lt;/img&gt;&lt;img&gt;&lt;EM&gt;Figure 11. The Trace-to-Decision Loop converts execution evidence into release memory.&lt;/EM&gt;&lt;/img&gt;
&lt;H2 data-line="274"&gt;Practical safeguards&lt;/H2&gt;
&lt;P data-line="276"&gt;This approach has limitations and trade-offs:&lt;/P&gt;
&lt;UL data-line="278"&gt;
&lt;LI data-line="278"&gt;Evaluation does not prove all future behavior. A dataset represents sampled expectations.&lt;/LI&gt;
&lt;LI data-line="279"&gt;Consequence classes are context-dependent and require domain ownership.&lt;/LI&gt;
&lt;LI data-line="280"&gt;Strict gates can create false blocks and delivery friction if contracts are poorly designed.&lt;/LI&gt;
&lt;LI data-line="281"&gt;Model-based evaluators can disagree or vary and need calibration.&lt;/LI&gt;
&lt;LI data-line="282"&gt;Traces can expose sensitive content and should be minimized, sanitized, and governed.&lt;/LI&gt;
&lt;LI data-line="283"&gt;Preview capabilities should be identified accurately, and a manually curated fallback should be available when automation is not appropriate for production use.&lt;/LI&gt;
&lt;LI data-line="284"&gt;Synthetic data can improve coverage, but it should be reviewed before being treated as authoritative ground truth.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="286"&gt;The objective is not certainty. It is a more explainable and defensible engineering decision.&lt;/P&gt;
&lt;H2 data-line="288"&gt;Conclusion&lt;/H2&gt;
&lt;P data-line="290"&gt;Production agent quality cannot be reduced to one answer score.&lt;/P&gt;
&lt;P data-line="292"&gt;When an agent can route work, call tools, construct arguments, trigger transactions, or claim that an action succeeded, the release process must consider the consequence of the execution path.&lt;/P&gt;
&lt;P data-line="294"&gt;The Consequence-Aware Agent Release Contract provides a practical way to do that:&lt;/P&gt;
&lt;UL data-line="296"&gt;
&lt;LI data-line="296"&gt;trace what happened;&lt;/LI&gt;
&lt;LI data-line="297"&gt;state what was expected;&lt;/LI&gt;
&lt;LI data-line="298"&gt;identify the evidence;&lt;/LI&gt;
&lt;LI data-line="299"&gt;classify the consequence;&lt;/LI&gt;
&lt;LI data-line="300"&gt;apply Pass, Review, or Block treatment; and&lt;/LI&gt;
&lt;LI data-line="301"&gt;preserve the failure as regression memory.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="303"&gt;The result is a simple but important shift:&lt;/P&gt;
&lt;P data-line="305"&gt;The higher-scoring agent is not automatically the agent that should ship.&lt;/P&gt;
&lt;H2 data-line="307"&gt;References&lt;/H2&gt;
&lt;UL data-line="309"&gt;
&lt;LI data-line="309"&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/foundry/observability/how-to/evaluate-agent" data-href="https://learn.microsoft.com/en-us/azure/foundry/observability/how-to/evaluate-agent" target="_blank"&gt;Evaluate your AI agents in Microsoft Foundry&lt;/A&gt;&lt;/LI&gt;
&lt;LI data-line="310"&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/foundry/observability/concepts/trace-agent-concept" data-href="https://learn.microsoft.com/en-us/azure/foundry/observability/concepts/trace-agent-concept" target="_blank"&gt;Agent tracing overview in Microsoft Foundry&lt;/A&gt;&lt;/LI&gt;
&lt;LI data-line="311"&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/foundry/concepts/observability" data-href="https://learn.microsoft.com/en-us/azure/foundry/concepts/observability" target="_blank"&gt;Observability in generative AI&lt;/A&gt;&lt;/LI&gt;
&lt;LI data-line="312"&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/foundry/observability/how-to/traces-to-dataset" data-href="https://learn.microsoft.com/en-us/azure/foundry/observability/how-to/traces-to-dataset" target="_blank"&gt;Convert agent traces into evaluation datasets&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;</description>
      <pubDate>Thu, 01 Oct 2026 13:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/beyond-agent-scores-building-consequence-aware-release-gates/ba-p/4559709</guid>
      <dc:creator>Abhishek Kumar</dc:creator>
      <dc:date>2026-10-01T13:00:00Z</dc:date>
    </item>
    <item>
      <title>From "Agent Deployed" to "Agent Ready" - Agent Acceptance Gateway</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/from-quot-agent-deployed-quot-to-quot-agent-ready-quot-agent/ba-p/4558096</link>
      <description>&lt;H2&gt;A Microsoft Foundry handover story&lt;/H2&gt;
&lt;P&gt;The infrastructure team has finished the AI landing zone. The Microsoft Foundry project exists. The prompt agent is already created, connected to its model, and configured with the tools or connections it needs.&lt;/P&gt;
&lt;P&gt;That is an important milestone, but it is not the same as saying the application team is ready to use it.&lt;/P&gt;
&lt;P&gt;At handover, the question should be simple:&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;Can the application team call the existing Foundry project endpoint and the named agent, using the expected identity, and get useful evidence that the agent can reach its configured services?&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;The application team should not be handed a test that creates another LLM client or rebuilds integrations locally. That may prove a script works, but it does not prove the handed-over agent is ready.&lt;/P&gt;
&lt;img&gt;Handover Gap&lt;/img&gt;
&lt;H2&gt;The challenge with poor handover&lt;/H2&gt;
&lt;P&gt;In many projects, handover sounds like this:&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;"The agent is deployed."&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;But the receiving team hears a different message:&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;"The platform is ready for application integration."&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;Those two statements are not always the same.&lt;/P&gt;
&lt;P&gt;As a Senior Cloud Solution Architect, this is where I usually see confusion begin. The portal demo may work. The resources may exist. The model may be attached. But the application team still needs to know whether the runtime contract works from their side of the boundary.&lt;/P&gt;
&lt;P&gt;Without a clear handover check, teams can validate the wrong thing. A developer may write a local test that calls a model directly, connects to search directly, or uses a different identity path. The test may pass, but it has bypassed the actual Foundry prompt agent.&lt;/P&gt;
&lt;P&gt;That creates a dangerous result: a green test for the wrong architecture.&lt;/P&gt;
&lt;H2&gt;What a good handover should prove&lt;/H2&gt;
&lt;P&gt;A useful handover does not need to test every business scenario. It should prove the basics that matter before the application depends on the agent:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;The application identity can authenticate without secrets.&lt;/LI&gt;
&lt;LI&gt;The Foundry project endpoint is reachable.&lt;/LI&gt;
&lt;LI&gt;The named prompt agent can be found.&lt;/LI&gt;
&lt;LI&gt;The agent can be invoked through the intended boundary.&lt;/LI&gt;
&lt;LI&gt;The agent can use the services that are part of its configured contract.&lt;/LI&gt;
&lt;LI&gt;The result is recorded clearly as &lt;STRONG&gt;PASS&lt;/STRONG&gt;, &lt;STRONG&gt;FAIL&lt;/STRONG&gt;, or &lt;STRONG&gt;SKIP&lt;/STRONG&gt;.&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;/LI&gt;
&lt;/OL&gt;
&lt;H2&gt;The better pattern - Agent Acceptance Gateway&lt;/H2&gt;
&lt;P&gt;The handover should feel like a relay race, not a treasure hunt.&lt;/P&gt;
&lt;P&gt;The infrastructure team hands over the project endpoint, the agent name, and the identity expectations. The application team runs a small passwordless validator from its side. The validator does not recreate the solution. It checks the existing agent.&amp;nbsp;&lt;/P&gt;
&lt;P&gt;That distinction matters. The validator is not a second implementation of the agent, and it does not connect directly to the model, Azure AI Search, Azure Cosmos DB, or Azure Blob Storage. It approaches the solution exactly as the application will: through the handed-over Microsoft Foundry project endpoint and the named agent.&lt;/P&gt;
&lt;P&gt;The infrastructure team owns the platform-side readiness. It confirms that the project and agent exist, the agent version is the intended one, required tools and connections are configured, network paths are available, and the application's managed identity has the minimum RBAC permissions needed to invoke the agent.&lt;/P&gt;
&lt;P&gt;The application team owns the consumer-side acceptance check. It runs the validator from the real application environment using ManagedIdentityCredential, not a developer login or stored secret. This proves the runtime identity, DNS, routing, private endpoint configuration, and authorization path that production will use.&lt;/P&gt;
&lt;H2&gt;Validation flow&lt;/H2&gt;
&lt;P&gt;The validator follows the same six checks as the handover checklist. Requests travel through the Foundry project and named agent; the validator never connects directly to Search, Cosmos DB, or Blob Storage. The service probes should use known, non-sensitive test data. Avoid prompts that return secrets, personal data, or entire documents merely to prove connectivity.&lt;/P&gt;
&lt;H2&gt;What the validator could look like&lt;/H2&gt;
&lt;P&gt;The following Python is intentionally illustrative. It shows the responsibilities and control flow rather than prescribing exact SDK methods, which can change between package versions. Adapt the agent lookup and invocation calls to the SDK and agent type your organization has standardized.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;# &lt;STRONG&gt;Illustrative flow only: align method names with the SDK version you use.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;import os&lt;/P&gt;
&lt;P&gt;from azure.identity import ManagedIdentityCredential&lt;/P&gt;
&lt;P&gt;from azure.ai.projects import AIProjectClient&lt;/P&gt;
&lt;P&gt;from dotenv import load_dotenv&lt;/P&gt;
&lt;P&gt;load_dotenv()&amp;nbsp; # Local convenience; Azure supplies the same values as app settings.&lt;/P&gt;
&lt;P&gt;def validate_agent_handover(project_endpoint: str, agent_name: str) -&amp;gt; dict:&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; """Check the handed-over agent without rebuilding its integrations."""&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; client_id = os.getenv("AZURE_CLIENT_ID")&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; credential = (&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; ManagedIdentityCredential(client_id=client_id)&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; if client_id else ManagedIdentityCredential()&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; )&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; results = {&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; "identity": "SKIP", "endpoint": "SKIP", "agent": "SKIP",&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; "invocation": "SKIP", "services": "SKIP", "report": "SKIP",&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; }&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; try:&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; credential.get_token("https://cognitiveservices.azure.com/.default")&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; results["identity"] = "PASS"&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; except Exception as exc:&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; results["identity"] = f"FAIL: {exc}"&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; return results&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; try:&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; client = AIProjectClient(endpoint=project_endpoint, credential=credential)&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; results["endpoint"] = "PASS"&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; agent = client.agents.get(agent_name=agent_name)&amp;nbsp; # conceptual lookup&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; results["agent"] = "PASS" if agent else "FAIL: agent not found"&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; except Exception as exc:&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; results["endpoint"] = f"FAIL: {exc}"&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; return results&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; try:&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; response = client.agents.run(&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; agent_name=agent_name,&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; messages=[{&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; "role": "user",&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; "content": "Return OK and identify which configured data source you used.",&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; }],&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; )&amp;nbsp; # conceptual invocation&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; results["invocation"] = "PASS" if response else "FAIL: no response"&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; results["services"] = "PASS"&amp;nbsp; # evaluate expected probe evidence here&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; except Exception as exc:&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; results["invocation"] = f"FAIL: {exc}"&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; results["report"] = "PASS"&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; return results&lt;/P&gt;
&lt;P&gt;if __name__ == "__main__":&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; print(validate_agent_handover(&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; os.environ["FOUNDRY_PROJECT_ENDPOINT"],&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; os.environ["FOUNDRY_AGENT_NAME"],&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp; ))&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;Configure passwordless identity&lt;/H2&gt;
&lt;P&gt;ManagedIdentityCredential deliberately restricts the validator to Azure managed identity. It does not fall back to an Azure CLI login, developer credential, environment secret, or interactive browser. This makes the acceptance result evidence for the application workload’s real runtime identity.&lt;/P&gt;
&lt;P&gt;Before running the validator in Azure, enable managed identity on the hosting resource and grant that identity the least-privileged roles required to access the Foundry project and invoke the agent. The agent’s own connections remain configured and authorized separately by the infrastructure team.&lt;/P&gt;
&lt;H3&gt;Example .env file&lt;/H3&gt;
&lt;P&gt;Use an environment file only for local convenience. Do not commit it. In Azure App Service, Azure Functions, Azure Container Apps, or AKS, provide the same values through the platform’s application settings or workload configuration.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;# &lt;STRONG&gt;.env contains identifiers and configuration—not a password or API key.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;FOUNDRY_PROJECT_ENDPOINT=https://&amp;lt;project-name&amp;gt;.&amp;lt;region&amp;gt;.api.azureml.ms&lt;/P&gt;
&lt;P&gt;FOUNDRY_AGENT_NAME=&amp;lt;handed-over-agent-name&amp;gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;# Set only when the workload uses a user-assigned managed identity.&lt;/P&gt;
&lt;P&gt;AZURE_CLIENT_ID=&amp;lt;managed-identity-client-id&amp;gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3&gt;Run the validator&lt;/H3&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;# &lt;STRONG&gt;Install the SDK packages selected by your application team.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;python -m pip install azure-identity azure-ai-projects python-dotenv&lt;/P&gt;
&lt;P&gt;# Run from an Azure compute resource that has managed identity enabled.&lt;/P&gt;
&lt;P&gt;# System-assigned identity: omit AZURE_CLIENT_ID.&lt;/P&gt;
&lt;P&gt;python validator.py&lt;/P&gt;
&lt;P&gt;# Azure-hosted workload: enable managed identity, assign the required RBAC roles,&lt;/P&gt;
&lt;P&gt;# add the non-secret environment variables above, and start the same command.&lt;/P&gt;
&lt;P&gt;python validator.py&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P&gt;ManagedIdentityCredential is intended to run on Azure compute with managed identity enabled. Run the handover gate from the application’s deployed environment so the result proves the intended identity, network route, DNS, private endpoints, and role assignments. For a user-assigned identity, set AZURE_CLIENT_ID to its client ID; for a system-assigned identity, omit that setting.&lt;/P&gt;
&lt;H2&gt;Probe prompts for configured services&lt;/H2&gt;
&lt;P&gt;If Azure AI Search, Azure Cosmos DB, or Azure Blob Storage are part of the agent’s contract, the validator sends focused probe prompts through the agent and records the outcome. If a service is not part of the contract, that probe is marked SKIP rather than assumed.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table class="lia-border-style-solid" border="1" style="width: 1051px; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Service&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Example probe prompt&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Expected evidence&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Azure AI Search&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;“List the top three documents in your index about [known topic]. Include titles and short snippets.”&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Document titles and relevant snippets from known indexed content.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Azure Cosmos DB&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;“What is the most recent record timestamp in the configured data source?”&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;A valid timestamp that can be checked against a known range.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Azure Blob Storage&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;“Confirm that you can access the configuration file at [known path]. Return metadata only.”&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Expected file name, content type, last-modified time, or a safe summary.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 163px" /&gt;&lt;col style="width: 452px" /&gt;&lt;col style="width: 435px" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;These probes test the agent’s actual tool connections without requiring the validator to authenticate to those services directly. A strong acceptance result records the prompt, status, timestamp, correlation or run identifier, and a small piece of expected evidence—without storing sensitive response content.&lt;/P&gt;
&lt;P&gt;The result gives both teams the same evidence. Instead of asking whether the agent was deployed, they can ask the better question:&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P&gt;&lt;STRONG&gt;Is the agent ready for the application to use?&lt;/STRONG&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H2&gt;Closing thought&lt;/H2&gt;
&lt;P&gt;Deploying an agent is only the first half of the journey. The real test is whether the application team can use it confidently through the expected&lt;/P&gt;
&lt;P&gt;identity, network, and capabilities.&amp;nbsp;&lt;/P&gt;
&lt;P&gt;The &lt;STRONG&gt;Agent Acceptance Gateway&lt;/STRONG&gt; creates that moment of confidence. It brings both teams together, checks what matters, shows what still needs attention, and turns the handover into a shared decision.&amp;nbsp;&lt;/P&gt;
&lt;P&gt;The goal is simple: do not hand over an agent just because it exists. Hand it over when both teams have clear evidence that it is ready.&lt;/P&gt;</description>
      <pubDate>Wed, 30 Sep 2026 18:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/from-quot-agent-deployed-quot-to-quot-agent-ready-quot-agent/ba-p/4558096</guid>
      <dc:creator>SudhirRawat</dc:creator>
      <dc:date>2026-09-30T18:00:00Z</dc:date>
    </item>
    <item>
      <title>Cohere Embed V5 brings more relevant retrieval to Microsoft Foundry</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/cohere-embed-v5-brings-more-relevant-retrieval-to-microsoft/ba-p/4560845</link>
      <description>&lt;P data-pm-slice="1 1 []"&gt;Today, we’re introducing Cohere Embed V5 Pro and Embed V5 Fast in Microsoft Foundry, bringing stronger multimodal retrieval to enterprise search and AI applications. With Fast and Pro options, support for more than 100 languages, and flexible vector sizes, Embed V5 gives developers more control over retrieval quality and efficiency.&lt;/P&gt;
&lt;H3&gt;Find relevant information across text and images&lt;/H3&gt;
&lt;P&gt;Enterprise knowledge spans documents, wikis, reports, and images. Embed V5 helps applications find relevant information across these sources and languages, supplying richer context to search experiences and retrieval-augmented generation (RAG) workflows.&lt;/P&gt;
&lt;P&gt;According to Cohere, Embed V5 surpasses Embed V4 across multimodal, multilingual, and domain-specific search and retrieval benchmarks. For AI agents, better retrieval can mean more useful context, better answers, and fewer unnecessary tokens spent processing irrelevant information.&lt;/P&gt;
&lt;H3&gt;Choose Fast or Pro to match your workload&lt;/H3&gt;
&lt;P&gt;Embed V5 introduces two options.&amp;nbsp;&lt;STRONG&gt;Fast&lt;/STRONG&gt; is designed for latency-sensitive workloads, such as interactive search and responsive agent experiences. &lt;STRONG&gt;Pro&lt;/STRONG&gt; prioritizes maximum retrieval quality for applications where finding the most relevant information is critical.&lt;/P&gt;
&lt;P&gt;Both share a single embedding space, supporting interchangeable use and giving teams flexibility as their application requirements evolve.&lt;/P&gt;
&lt;H3&gt;Scale retrieval with more control over cost&lt;/H3&gt;
&lt;P&gt;A broader range of Matryoshka embedding dimensions lets developers choose smaller vectors to reduce storage requirements and retrieval latency while targeting the quality their application needs. High-throughput batch processing supports embedding large content collections for indexing at scale.&lt;/P&gt;
&lt;P&gt;Beyond search and RAG, Embed V5 also supports classification, clustering, recommendations, and similarity matching, helping teams put enterprise data to work across more applications.&lt;/P&gt;
&lt;H3&gt;Get started in Microsoft Foundry&lt;/H3&gt;
&lt;P&gt;Explore Cohere Embed V5 for your next search or RAG application. Evaluate Fast and Pro on your own enterprise content to find the right balance of relevance, latency, and cost.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Explore Cohere Embed V5 in Microsoft Foundry.&lt;/STRONG&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://ai.azure.com/catalog/models/Cohere-Embed-V5-Pro?search=Embed" target="_blank"&gt;Cohere-Embed-V5-Pro | Model Catalog | Microsoft Foundry&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://ai.azure.com/catalog/models/Cohere-Embed-V5-Fast?search=embed" target="_blank"&gt;Cohere-Embed-V5-Fast | Model Catalog | Microsoft Foundry&lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Wed, 30 Sep 2026 16:59:02 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/cohere-embed-v5-brings-more-relevant-retrieval-to-microsoft/ba-p/4560845</guid>
      <dc:creator>RashaudSavage</dc:creator>
      <dc:date>2026-09-30T16:59:02Z</dc:date>
    </item>
    <item>
      <title>Let the query decide: introducing auto retrieval reasoning effort</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/let-the-query-decide-introducing-auto-retrieval-reasoning-effort/ba-p/4558943</link>
      <description>&lt;H2 style="font-size: 24px; line-height: 1.3; font-weight: 400; margin-top: 0; margin-bottom: 20px;"&gt;Introducing &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt;: a new retrieval reasoning effort mode that recovers 97.6% of the evidence recall of the highest effort level at 46.9% fewer query planning tokens.&lt;/H2&gt;
&lt;P style="text-align: justify;"&gt;Enterprise retrieval workloads mix simple lookup queries with complex questions that require searching across many documents. Today, customers of Foundry IQ (Azure AI Search) Knowledge Bases must choose a fixed retrieval reasoning effort level: &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;minimal&lt;/CODE&gt;, &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;low&lt;/CODE&gt; or &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;medium&lt;/CODE&gt;. That forces customers to make a trade-off across their workload: pay for deeper retrieval even on simple questions, or reduce costs and risk missing the evidence needed to answer more complex ones.&lt;/P&gt;
&lt;P style="text-align: justify;"&gt;The new &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; retrieval reasoning effort, available in the August 2026 preview, reduces the need to choose one fixed effort level for mixed workloads. &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; assesses each query and adjusts the query planning and iterative search effort it receives. Straightforward questions take a faster, lower-cost path, while complex questions get the depth of multi-step retrieval, without manual tuning.&lt;/P&gt;
&lt;P style="text-align: justify;"&gt;Customers can still use &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;minimal&lt;/CODE&gt;, &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;low&lt;/CODE&gt; or &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;medium&lt;/CODE&gt; when their workloads are predictable. For mixed workloads, &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; lets each query determine the effort it needs.&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;Near-identical retrieval quality, lower cost&lt;/STRONG&gt;&amp;nbsp;&lt;/H2&gt;
&lt;P style="text-align: justify;"&gt;Figure 1 shows the central trade-off across retrieval reasoning effort tiers. Higher fixed effort generally retrieves more complete evidence, but it also consumes more tokens. &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; assigns lower effort to straightforward queries and reserves &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;medium&lt;/CODE&gt; (the highest retrieval reasoning effort) for the queries that need it, &lt;STRONG&gt;retaining 97.6% of its evidence recall while reducing average query planning token usage by 46.9%.&lt;/STRONG&gt;&lt;/P&gt;
&lt;DIV style="width: 600px; max-width: 100%; margin: 20px auto;"&gt;&lt;img&gt;&lt;SPAN style="display: block; width: 100%; text-align: center;" data-mce-style="display:block;width:100%;text-align:center;"&gt;&lt;EM&gt;Figure 1. Average evidence recall by retrieval reasoning effort using GPT 5.4 mini for query planning. Evidence recall measures how much of the evidence needed to answer a question is found in the retrieved documents. &lt;/EM&gt;&lt;/SPAN&gt;&lt;/img&gt;&lt;/DIV&gt;
&lt;P style="text-align: justify;"&gt;Figure 2 translates the token savings into customer cost. With GPT-5.4 mini, &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; reduces the average cost of 1,000 queries from $13.82 to $9.03 (a saving of $4.79, or 34.7%) compared with running every query at &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;medium&lt;/CODE&gt; effort. Because search costs remain fixed, the dollar savings increase when a more expensive model is used for query planning.&lt;/P&gt;
&lt;DIV style="width: 600px; max-width: 100%; margin: 20px auto;"&gt;&lt;img&gt;&lt;SPAN style="display: block; width: 100%; text-align: center;" data-mce-style="display:block;width:100%;text-align:center;"&gt;&lt;EM&gt;Figure 2. Total query cost ($/1,000 queries) per model and retrieval reasoning effort. Cost estimates use &lt;A href="https://azure.microsoft.com/en-gb/pricing/details/azure-openai/#pricing" target="_blank" rel="noopener"&gt;Azure OpenAI pricing&lt;/A&gt; for query planning and &lt;A href="https://azure.microsoft.com/en-gb/pricing/details/search/" target="_blank" rel="noopener"&gt;Azure AI Search&lt;/A&gt; pricing for ranker/search as of September 2026.&lt;/EM&gt;&lt;/SPAN&gt;&lt;/img&gt;&lt;/DIV&gt;
&lt;H2&gt;&lt;STRONG&gt;Effort that scales with the question&lt;/STRONG&gt;&amp;nbsp;&lt;/H2&gt;
&lt;P style="text-align: justify;"&gt;Not every question needs the same retrieval depth. The &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; retrieval reasoning effort matches retrieval depth to each question: &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;minimal&lt;/CODE&gt; or &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;low&lt;/CODE&gt; for straightforward requests, and &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;medium&lt;/CODE&gt; when the question requires additional planning and iterative search. Figure 3 shows how those routing decisions vary across datasets. In the evaluated enterprise workloads, most queries follow the faster &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;minimal&lt;/CODE&gt; or &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;low&lt;/CODE&gt; paths, while datasets containing more complex questions send a larger share of queries to &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;medium&lt;/CODE&gt;.&lt;/P&gt;
&lt;DIV style="width: 600px; max-width: 100%; margin: 20px auto;"&gt;&lt;img&gt;&lt;SPAN style="display: block; width: 100%; text-align: center;" data-mce-style="display:block;width:100%;text-align:center;"&gt;&lt;EM&gt;Figure 3. Distribution of effort tiers selected by &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;" data-mce-style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; across the evaluated enterprise queries. KS = knowledge source.&lt;/EM&gt;&lt;/SPAN&gt;&lt;/img&gt;&lt;/DIV&gt;
&lt;P style="text-align: justify;"&gt;This routing behavior translates directly into the token savings shown in Figure 4. &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; remains close to &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;medium&lt;/CODE&gt; in evidence recall while using fewer tokens across every evaluated dataset. The datasets that route a larger share of queries to &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;minimal&lt;/CODE&gt; achieve the greatest savings, because more queries avoid reasoning tokens when additional retrieval depth is unlikely to help. For a direct lookup query, this can mean using no reasoning tokens instead of the approximately 14,000 consumed by the average &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;medium&lt;/CODE&gt; run.&lt;/P&gt;
&lt;DIV style="width: 600px; max-width: 100%; margin: 20px auto;"&gt;&lt;img&gt;&lt;SPAN style="display: block; width: 100%; text-align: center;" data-mce-style="display:block;width:100%;text-align:center;"&gt;&lt;EM&gt;Figure 4. Evidence recall for &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;" data-mce-style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; and &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;" data-mce-style="color: inherit; background-color: transparent; font-style: normal;"&gt;medium&lt;/CODE&gt; by dataset. &lt;/EM&gt;&lt;/SPAN&gt;&lt;/img&gt;&lt;/DIV&gt;
&lt;P style="text-align: justify;"&gt;The same adaptivity works in reverse. On &lt;A href="https://openai.com/index/browsecomp/" target="_blank" rel="noopener"&gt;BrowseComp&lt;/A&gt;, a public benchmark of deliberately difficult, multi-step research questions, &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; escalates 78% of queries to &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;medium&lt;/CODE&gt;. Its resulting evidence recall remains within a few points of running every BrowseComp query at &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;medium&lt;/CODE&gt;, showing that the router increases effort when simpler retrieval paths are unlikely to be sufficient.&lt;/P&gt;
&lt;P style="text-align: justify;"&gt;It is worth knowing the edges. Choose &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;medium&lt;/CODE&gt; when most queries are complex, multi-step questions and retrieval quality matters more than cost. Choose &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;minimal&lt;/CODE&gt; when the workload consists almost entirely of direct lookups. For mixed workloads, &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; is the recommended starting point because it preserves close to &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;medium&lt;/CODE&gt;-level evidence recall while reducing average cost without per-query configuration.&lt;/P&gt;
&lt;P style="text-align: justify;"&gt;Figure 5 shows how &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; applies per-query routing across three queries of increasing complexity: a direct lookup, a simple multi-hop question, and a complex multi-hop question. &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; selects the lowest retrieval effort expected to preserve evidence quality, escalating from &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;minimal&lt;/CODE&gt; to &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;low&lt;/CODE&gt; or &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;medium&lt;/CODE&gt; as the query requires more planning and retrieval steps.&lt;/P&gt;
&lt;DIV style="width: 600px; max-width: 100%; margin: 20px auto;"&gt;&lt;img&gt;&lt;SPAN style="display: block; width: 100%; text-align: center;" data-mce-style="display:block;width:100%;text-align:center;"&gt;&lt;EM&gt;Figure 5. Example routing decisions for three query types of increasing complexity. &lt;/EM&gt;&lt;/SPAN&gt;&lt;/img&gt;&lt;/DIV&gt;
&lt;H2&gt;&lt;STRONG&gt;Get started&lt;/STRONG&gt;&amp;nbsp;&lt;/H2&gt;
&lt;P style="text-align: justify;"&gt;&lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; retrieval reasoning effort is now available in preview for Foundry IQ (Azure AI Search) Knowledge Bases. For workloads that mix straightforward lookups with complex questions, &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; provides a starting point without requiring you to choose a fixed effort level for every query.&lt;/P&gt;
&lt;P style="text-align: justify;"&gt;If you already have a knowledge base, set &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;retrievalReasoningEffort&lt;/CODE&gt; to &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; as its default, or specify &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; on an individual retrieve request to override that default for the request. See &lt;A href="https://learn.microsoft.com/en-us/azure/search/agentic-retrieval-how-to-set-retrieval-reasoning-effort" target="_blank" rel="noopener"&gt;Set the retrieval reasoning effort&lt;/A&gt; for configuration details and supported versions.&lt;/P&gt;
&lt;P style="text-align: justify;"&gt;If you are new to knowledge bases, follow the &lt;A href="https://learn.microsoft.com/en-us/azure/search/search-get-started-agentic-retrieval?tabs=windows&amp;amp;pivots=csharp" target="_blank" rel="noopener"&gt;agentic retrieval quickstart&lt;/A&gt; to create one and run your first queries. Then try &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; on a representative set of your own questions and compare the retrieved evidence and query-planning token usage with a fixed effort level.&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;Appendix&lt;/STRONG&gt;&amp;nbsp;&lt;/H2&gt;
&lt;P style="text-align: justify;"&gt;Similarly to our previous blog posts (&lt;A href="https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/foundry-iq-improve-recall-by-up-to-54-with-knowledge-bases/4524852" target="_blank" rel="noopener"&gt;Foundry IQ: Improve recall by up to 54% with knowledge bases | Microsoft Community Hub&lt;/A&gt;, &lt;A href="https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/up-to-40-better-relevance-for-complex-queries-with-new-agentic-retrieval-engine/4413832" target="_blank" rel="noopener"&gt;Up to 40% better relevance for complex queries with new agentic retrieval engine&lt;/A&gt;) we evaluated &lt;CODE style="color: inherit; background-color: transparent; font-style: normal;"&gt;auto&lt;/CODE&gt; on several benchmark query sets.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;SPAN style="display: block; text-align: justify;"&gt;&lt;STRONG&gt;Customer datasets:&lt;/STRONG&gt; customer-provided corporate and member-document collections in English, covering domains such as oil and gas corporate reports and health-insurance member documents. We evaluate them as single knowledge source (KS) workloads over chunked PDF indexes.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN style="display: block; text-align: justify;"&gt;&lt;STRONG&gt;SEC: &lt;/STRONG&gt;SEC filings of US public companies in English, evaluated in two variants: a single knowledge source setup spanning all sectors, and a routing setup with one knowledge source per Global Industry Classification Standard (GICS) sector.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN style="display: block; text-align: justify;"&gt;&lt;STRONG style="color: rgb(30, 30, 30);"&gt;MIML:&amp;nbsp;&lt;/STRONG&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;a multi-industry, multi-language corporate document benchmark in English, French, and Simplified Chinese. We evaluate both single knowledge source variants restricted to one language-industry slice and a routing variant spanning multiple languages and industries.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN style="display: block; text-align: justify;"&gt;&lt;STRONG style="color: rgb(30, 30, 30);"&gt;BrowseComp:&amp;nbsp;&lt;/STRONG&gt;&lt;SPAN style="color: rgb(30, 30, 30);"&gt;We use the 830 human-verified BrowseComp-Plus queries and indexed the full corpus, including distractor documents, into 512-token chunks with OpenAI text-embedding-3-large. To support continuous evidence-recall measurement, we decompose each question into multiple atomic factoid evidence nuggets. Evaluations use a two-agent neutral user plus search agent configuration with a standard system prompt and per question effort capped at 40 tool calls.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;</description>
      <pubDate>Wed, 30 Sep 2026 16:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/let-the-query-decide-introducing-auto-retrieval-reasoning-effort/ba-p/4558943</guid>
      <dc:creator>amaias</dc:creator>
      <dc:date>2026-09-30T16:00:00Z</dc:date>
    </item>
    <item>
      <title>Introducing GPT-6.1 Sol in Microsoft Foundry: Advanced intelligence, optimized for production agents</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/introducing-gpt-6-1-sol-in-microsoft-foundry-advanced/ba-p/4560811</link>
      <description>&lt;P&gt;Today, OpenAI's GPT-6.1 Sol is generally available in Microsoft Foundry. GPT-6.1 Sol is an upgrade to GPT-6 Sol, delivering substantial improvements in agentic coding, computer use, and professional work, with performance approaching GPT-6 Astra across these evaluations. It offers a new balance of capability and cost for important work at higher frequency, making it more affordable for developers to build and run very capable agents at scale.&lt;/P&gt;
&lt;P&gt;Just one week after &lt;A href="https://azure.microsoft.com/en-us/blog/gpt-6-astra-sol-and-luna-for-production-agents-in-microsoft-foundry/" target="_blank" rel="noopener"&gt;GPT-6 Sol and Luna&lt;/A&gt; joined our generally available lineup, this release continues the momentum of the GPT-6 series in Foundry: models that produce less noise and are more capable of completing full tasks with agents. GPT-6.1 Sol carries that progress forward for teams whose production workloads run all day, every day.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;What's new in GPT-6.1 Sol&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;GPT-6.1 Sol advances the three capabilities that matter most for production agents:&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Agentic coding. &lt;/STRONG&gt;GPT-6.1 Sol plans, edits, tests, and iterates across a codebase with improved performance on complex tasks and extended workflows involving multiple tool calls.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;More capable computer use. &lt;/STRONG&gt;Improved reliability navigating real interfaces help agents operate the workflows that span multiple application steps, with permissions and human oversight suited to the task.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Deeper Professional work. &lt;/STRONG&gt;Stronger performance on the analysis, drafting, and multi-step knowledge tasks that make up daily enterprise work, with improvements in factual accuracy those workflows demand.&lt;/P&gt;
&lt;P&gt;GPT-6.1 Sol accepts text and image inputs and produces text, with a total context window of up to 1M tokens. This provides room to bring large codebases and document sets, within the model’s context limits. Flexible reasoning lets teams tailor the model’s depth of analysis to the needs of each task. This can help teams use tokens more efficiently by reserving deeper reasoning for the work that needs it.&lt;/P&gt;
&lt;P&gt;As in &lt;A href="https://azure.microsoft.com/en-us/blog/gpt-6-astra-sol-and-luna-for-production-agents-in-microsoft-foundry/" target="_blank" rel="noopener"&gt;last Tuesday's announcement&lt;/A&gt;, the starting point is quality alongside cost per task, not model capability in isolation. Use that lens to evaluate GPT-6.1 Sol against the needs of your own production workloads.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;STRONG&gt;Where to put GPT-6.1 Sol to work&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Because these gains compound in agentic loops, valuable use cases include agent workloads that run frequently:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Software engineering agents &lt;/STRONG&gt;that triage issues, implement changes across a repository, respond to code review, and keep CI green, with costs that support running them on every pull request, not just the hard ones.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Computer-use agents &lt;/STRONG&gt;that complete back-office processes end to end: updating records across line-of-business systems, reconciling data between applications, and handling workflows in browsers.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Professional work agents &lt;/STRONG&gt;for research synthesis, contract and document review, financial analysis, and report generation, where work recurs weekly or daily and rewards consistent quality per task.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;High-frequency customer and employee workflows, &lt;/STRONG&gt;where GPT-6.1 Sol's capability-cost balance lets teams upgrade the intelligence behind every interaction without upgrading the budget.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;The right intelligence behind every agent&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The right model for a job should be determined through evaluations. GPT-6.1 Sol joins GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna in the Foundry model catalog: start with GPT-6 Astra for the most demanding reasoning, use GPT-6.1 Sol as the new default for production agents and complex workflows, and scale high-volume data and preparatory tasks with Luna. As with the rest of the GPT-6 series, look beyond price per token to &lt;STRONG&gt;cost per task&lt;/STRONG&gt;, alongside quality and reliability. Foundry brings evaluation, tracing, and monitoring together so teams can make that call with evidence, then switch models without re-platforming.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Deploy it your way&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;With GPT-6.1 Sol in Microsoft Foundry, teams can choose how they scale and where processing takes place. Standard deployment offers flexible capacity billed by usage across Global regions and US, EU, and APAC Data Zones. Provisioned Throughput is available at launch through Global and US Data Zone deployments, providing reserved capacity for critical production workloads. Additional regions and deployment options will follow soon.&lt;/P&gt;
&lt;P&gt;That is the Foundry advantage: the flexibility to balance capacity, performance, and data residency requirements within one enterprise platform. Teams can match each workload to a supported combination of serving option and processing location, rather than apply the same deployment approach to every application.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;GPT-6.1 Sol Pricing*&lt;/STRONG&gt;&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;td rowspan="2"&gt;
&lt;P&gt;Model&lt;/P&gt;
&lt;/td&gt;&lt;td rowspan="2"&gt;
&lt;P&gt;Deployment&lt;/P&gt;
&lt;/td&gt;&lt;td rowspan="2"&gt;
&lt;P&gt;Context Length&lt;/P&gt;
&lt;/td&gt;&lt;td colspan="4"&gt;
&lt;P&gt;Pricing (USD $/million tokens)&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Input&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Cached Input&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Cached Writes&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Output&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td rowspan="8"&gt;
&lt;P&gt;&lt;STRONG&gt;GPT-6.1 Sol&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td rowspan="2"&gt;
&lt;P&gt;Global Standard&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Short context&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$2.00&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$0.10&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$2.50&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$10.00&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Long context&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$4.00&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$0.20&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$5.00&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$15.00&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td rowspan="2"&gt;
&lt;P&gt;Data Zone Standard (US)&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Short context&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$2.20&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$0.11&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$2.75&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$11.00&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Long context&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$4.40&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$0.22&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$5.50&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$16.50&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td rowspan="2"&gt;
&lt;P&gt;Data Zone Standard (EU)&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Short context&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$2.40&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$0.12&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$3.00&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$12.00&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Long context&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$4.80&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$0.24&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$6.00&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$18.00&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Data Zone Standard (APAC)&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Short context&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$2.40&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$0.12&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$3.00&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$12.00&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Long context&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$4.80&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$0.24&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$6.00&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$18.00&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;*Prices shown are for Standard deployments. Provisioned Throughput pricing varies by deployment type. For each offer, the US Data Zone is priced at a 10% premium to Global, and the EU and APAC Data Zones are priced at a 20% premium to Global. For current rates and terms, see the Azure OpenAI pricing page.&lt;EM&gt;, see the &lt;/EM&gt;&lt;A href="https://azure.microsoft.com/en-us/pricing/details/azure-openai/?msockid=0030611da9e260e628f5742ca8586188" target="_blank" rel="noopener"&gt;&lt;EM&gt;Azure OpenAI pricing page&lt;/EM&gt;&lt;/A&gt;&lt;EM&gt;.&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Build safer agents on Foundry&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;GPT-6.1 Sol runs with the same layered protections as the GPT-6 series in Foundry: alignment training in the model, content filters and guardrails on prompts and outputs, prompt injection mitigations on tool calls and responses, and enterprise identity and access controls governing what agents can reach. Microsoft Purview applies data policies and human checkpoints at every phase.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Start building&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Your next agent needs more than a powerful model. Foundry pairs GPT-6.1 Sol with the deployment flexibility, observability, and enterprise controls to take it from first workload to production at scale. Explore GPT-6.1 Sol in the &lt;A href="https://ai.azure.com/catalog" target="_blank" rel="noopener"&gt;Microsoft Foundry model catalog&lt;/A&gt; and evaluate where it fits in your next agent.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 30 Sep 2026 00:08:07 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/introducing-gpt-6-1-sol-in-microsoft-foundry-advanced/ba-p/4560811</guid>
      <dc:creator>Naomi Moneypenny</dc:creator>
      <dc:date>2026-09-30T00:08:07Z</dc:date>
    </item>
    <item>
      <title>Claude Sonnet 5.5 is now available in Microsoft Foundry</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/claude-sonnet-5-5-is-now-available-in-microsoft-foundry/ba-p/4559774</link>
      <description>&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Today,&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;Claude Sonnet 5.5 is available in Microsoft Foundry&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;, bringing Anthropic’s latest Sonnet model to developers building AI applications and agents.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Claude Sonnet 5.5 is a smarter, more efficient Sonnet, delivering a clear step forward from Sonnet 5 on coding and knowledge work with a lower cost per task for most work at faster speed. For teams already building on Sonnet, it is a natural upgrade that can deliver stronger results.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="auto"&gt;More intelligence with fewer tokens&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Claude Sonnet 5.5 improves Sonnet capabilities while completing work more efficiently for most tasks. That means teams can take on more complex tasks without necessarily paying for more model usage.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;For developers, that translates into a better balance of quality and economics for everyday production workloads.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Stronger for focused coding&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Claude Sonnet 5.5 is particularly suited for well-scoped coding tasks that fit into a larger development workflow.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Use it to:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Build and fix features within the same coding session&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559740&amp;quot;:276,&amp;quot;469777462&amp;quot;:[0,720],&amp;quot;469777927&amp;quot;:[0,0],&amp;quot;469777928&amp;quot;:[1,1]}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="2" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Verify implementations against requirements&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559740&amp;quot;:276,&amp;quot;469777462&amp;quot;:[0,720],&amp;quot;469777927&amp;quot;:[0,0],&amp;quot;469777928&amp;quot;:[1,1]}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559685&amp;quot;:720,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="3" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Support everyday coding and knowledge-work tasks&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559740&amp;quot;:276,&amp;quot;469777462&amp;quot;:[0,720],&amp;quot;469777927&amp;quot;:[0,0],&amp;quot;469777928&amp;quot;:[1,1]}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;As agent systems become more modular, models increasingly need to perform individual tasks reliably and efficiently as part of a larger workflow. Claude Sonnet 5.5 is designed for exactly that role.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559740&amp;quot;:276}"&gt;Pricing&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;BR /&gt;&lt;table border="1" style="width: 100.001%; height: 69.3642px; border-width: 1px;"&gt;&lt;colgroup&gt;&lt;col style="width: 20%" /&gt;&lt;col style="width: 20%" /&gt;&lt;col style="width: 20%" /&gt;&lt;col style="width: 20%" /&gt;&lt;col style="width: 20%" /&gt;&lt;/colgroup&gt;&lt;tbody&gt;&lt;tr style="height: 34.6821px;"&gt;&lt;td style="height: 34.6821px;"&gt;&lt;STRONG&gt;Model&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;&lt;STRONG&gt;Deployment Type&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;&lt;STRONG&gt;Input&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;&lt;STRONG&gt;Ouput&lt;/STRONG&gt;&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;&lt;STRONG&gt;Availability&lt;/STRONG&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 34.6821px;"&gt;&lt;td style="height: 34.6821px;"&gt;Claude Sonnet 5.5&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;Global Standard, US Datazone&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;$2&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;$10&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;
&lt;P&gt;GA, Hosted on Azure&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Claude Sonnet 5.5&lt;/td&gt;&lt;td&gt;Global Standard&lt;/td&gt;&lt;td&gt;$2&lt;/td&gt;&lt;td&gt;$10&lt;/td&gt;&lt;td&gt;
&lt;P&gt;GA, Hosted on Anthropic Infrastructure&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="auto"&gt;Build with Claude Sonnet 5.5 in Microsoft Foundry&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;With Claude Sonnet 5.5 in Microsoft Foundry, developers have another strong option for balancing capability, efficiency, and cost across coding and knowledge-work applications.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;A class="lia-external-url" href="https://ai.azure.com/catalog/models/claude-sonnet-5-5?search=sonnet" target="_blank" rel="noopener"&gt;Try it today&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559740&amp;quot;:276}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Mon, 28 Sep 2026 18:18:01 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/claude-sonnet-5-5-is-now-available-in-microsoft-foundry/ba-p/4559774</guid>
      <dc:creator>amar_badal</dc:creator>
      <dc:date>2026-09-28T18:18:01Z</dc:date>
    </item>
    <item>
      <title>Introducing voice agents in Microsoft Foundry</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/introducing-voice-agents-in-microsoft-foundry/ba-p/4557276</link>
      <description>&lt;P&gt;Voice is quickly becoming a major way people interact with AI agents. Recent research found that 14% of users already prefer speaking with generative AI rather than typing&lt;SUP&gt;1&lt;/SUP&gt;, signaling that voice interactions are moving beyond early-adopter use cases and toward the mainstream. As adoption grows, users increasingly expect conversations that feel natural, responsive, and capable of carrying out real work.&lt;/P&gt;
&lt;P&gt;From customer support and employee assistance to learning companions and industry-specific workflows, organizations want agents that can listen, reason, and respond in real time. Yet many voice solutions still require developers to assemble separate speech, orchestration, deployment, monitoring, and evaluation components before they can reach production. As a result, teams often spend more time integrating infrastructure than improving the agent experience itself.&lt;/P&gt;
&lt;P&gt;Today's &lt;A class="lia-external-url" href="https://aka.ms/FoundrySept2026" target="_blank" rel="noopener"&gt;announcement &lt;/A&gt;introduces &lt;STRONG&gt;voice agents in Microsoft Foundry &lt;/STRONG&gt;&lt;STRONG&gt;Agent Service&lt;/STRONG&gt;, a voice-native approach that brings real-time speech, agent capabilities, deployment workflows, observability, and evaluation together in a single platform.&lt;/P&gt;
&lt;div data-video-id="https://youtu.be/u8usaIQMR5s?si=Nq__oxZOMCrLSTUP/1790618905918" data-video-remote-vid="https://youtu.be/u8usaIQMR5s?si=Nq__oxZOMCrLSTUP/1790618905918" class="lia-video-container lia-media-is-center lia-media-size-large"&gt;&lt;iframe src="https://cdn.embedly.com/widgets/media.html?src=https%3A%2F%2Fwww.youtube.com%2Fembed%2Fu8usaIQMR5s%3Ffeature%3Doembed&amp;amp;display_name=YouTube&amp;amp;url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3Du8usaIQMR5s&amp;amp;image=https%3A%2F%2Fi.ytimg.com%2Fvi%2Fu8usaIQMR5s%2Fhqdefault.jpg&amp;amp;type=text%2Fhtml&amp;amp;schema=youtube" allowfullscreen="" style="max-width: 100%"&gt;&lt;/iframe&gt;&lt;/div&gt;
&lt;P&gt;Developers can build voice experiences with sub-second latency, multilingual support, integrated knowledge and tools, deployment channels, and built-in lifecycle management, all within Microsoft Foundry. This includes broad model choice for text-to-speech and speech-to-text models from providers like &lt;A href="https://ai.azure.com/catalog/models?publisher=openai" target="_blank" rel="noopener"&gt;OpenAI&lt;/A&gt;, &lt;A href="https://ai.azure.com/catalog/models?publisher=microsoft" target="_blank" rel="noopener"&gt;Microsoft AI&lt;/A&gt; and Microsoft Azure. Voice agents support the same development workflows developers already use in Microsoft Foundry Agent Service, making it easy to bring real-time voice capabilities to existing agent experiences.&lt;/P&gt;
&lt;img /&gt;
&lt;H3&gt;From voice API to voice-native agents – choosing the right path for building voice experiences&lt;/H3&gt;
&lt;P&gt;Voice experiences have evolved from standalone speech applications into full agent systems. As a result, developers today can choose from multiple architectures depending on how much of the agent lifecycle they want the platform to manage.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;The three approaches shown above reflect different stages of voice application maturity. Some teams only need a real-time voice runtime. Others have already built agents and want to make those experiences conversational. Increasingly, however, organizations are looking for a voice solution that spans development, deployment, monitoring, evaluation, and optimization from a single platform.&lt;/P&gt;
&lt;P&gt;The important distinction is not simply how speech is added to an application, but where responsibility for the voice experience lives. Earlier approaches give developers maximum flexibility to assemble their own architecture. As requirements grow, however, teams often need deeper lifecycle capabilities, including observability, evaluation, governance, deployment workflows, and operational tooling. While many developers can build a compelling voice agent demo, operating voice agents in production requires visibility into every interaction, the ability to evaluate agent behavior, diagnose failures, and a system to continuously improve quality, latency, and cost. Those requirements become especially important for customer-facing experiences such as contact centers, employee assistants, and business process automation.&lt;/P&gt;
&lt;P&gt;Voice agents in Microsoft Foundry were designed around the needs of production applications. Rather than treating voice as a channel layered onto an existing agent, voice becomes a native part of how agents are built, deployed, observed, evaluated, and optimized. Real-time speech, agent capabilities, deployment workflows, observability, and evaluation are brought together through a single platform and SDK. Teams can tailor experiences to their specific scenarios by choosing the models that best fit their workloads, improving recognition for domain-specific vocabulary through speech customization, creating custom voices, and adding photo or video avatars that reflect their brand. Voice agents also include newly expanded standard talking-head avatars that can be enabled directly in the Foundry playground, making it easy to create a visual agent experience without developing custom avatar assets. Developers can start with these out-of-the-box avatars and evolve to custom voice and avatar experiences as their scenarios grow.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;All three approaches remain supported and serve different customer needs. For new enterprise voice workloads, however, Microsoft recommends voice agents in Foundry as the preferred path forward because it provides one of the most complete voice-agent lifecycle, from first conversation to production operations.&lt;/P&gt;
&lt;H3&gt;Build, test, and deploy voice agents with familiar developer tools&lt;/H3&gt;
&lt;P&gt;Voice agents are integrated into the developer workflows teams already use. With new support across &lt;STRONG&gt;AZD AI&lt;/STRONG&gt; and the &lt;STRONG&gt;Foundry Toolkit for Visual Studio Code (coming soon)&lt;/STRONG&gt;, developers can create, configure, test, debug, source-control, and deploy voice agents without stitching together separate tools or building custom voice testing applications.&lt;/P&gt;
&lt;P&gt;AZD AI introduces a dedicated workflow for voice agents. Developers can initialize a voice agent project, manage agent definitions and deployment configuration in source control, and deploy through the standard AZD lifecycle. Voice agents remain a dedicated application type because their models, runtime requirements, testing workflows, and deployment paths are optimized specifically for spoken interaction.&lt;/P&gt;
&lt;P&gt;Inside Visual Studio Code, developers can complete the entire development loop without leaving the editor. Teams can create and configure a voice agent, start a voice session, interact through natural speech, inspect transcripts and runtime events, and debug latency or connection issues from the same environment where they author the agent.&lt;/P&gt;
&lt;P&gt;Together, these capabilities reduce the time between writing an agent and having a real conversation with it, helping teams move from experimentation to production more quickly.&lt;/P&gt;
&lt;H3&gt;Operate voice agents with voice native lifecycle management&lt;/H3&gt;
&lt;P&gt;Building a great voice experience is only the first step. Once voice agents reach production, teams need to know why conversations succeed or fail, where latency is introduced, how users interact with the agent, and how the experience can be improved over time.&lt;/P&gt;
&lt;P&gt;Voice interactions introduce challenges that text agents never encounter. Speech-recognition errors, interruptions, turn-taking issues, and tool delays can all impact the conversation experience. Without observability built into the platform, teams often rely on replaying calls, reviewing logs across multiple systems, and manually diagnosing issues.&lt;/P&gt;
&lt;P&gt;Voice agents in Microsoft Foundry are built with &lt;STRONG&gt;observability, evaluation, and optimization&lt;/STRONG&gt; by default. They use the same Foundry Observability experience as every other Foundry agents, so there is nothing separate to configure. Each conversation is stored as a trace, and the Foundry evaluation framework runs against those traces in the same way it runs for text agents, making it easier to understand agent behavior, diagnose issues, and measure quality over time.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;These capabilities feed the continuous improvement loop that runs across Foundry:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Tracing and monitoring&lt;/STRONG&gt; provide visibility into every interaction, helping teams understand user behavior, diagnose failures, and find performance bottlenecks.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Evaluation&lt;/STRONG&gt; measures conversation quality against production traces, validates expected behavior, compares versions, and catches regressions before they reach users.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Insights and optimization &lt;/STRONG&gt;help teams improve quality, reliability, latency, and cost over time&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Voice agents are typically built to complete specific, time-sensitive tasks. Developers need to measure whether the agent did its job and followed the rules, and general-purpose evaluations alone cannot tell them that. &lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/foundry/concepts/evaluation-evaluators/rubric-evaluators" target="_blank" rel="noopener"&gt;&lt;STRONG&gt;Foundry rubric evaluators&lt;/STRONG&gt;&lt;/A&gt; score conversations against criteria written for a specific agent. You list what matters, such as "confirmed the date, time, and party size" or "did not book a party larger than eight," assign each criterion a weight, and an LLM judge scores every conversation against that list.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;Foundry can generate the rubric from the agent's instructions and production traces, and run it continuously against live traffic, so teams know when an agent skips a step or breaks a policy before customers do. Early customers building voice agents in Foundry are already using rubric evaluators this way.&lt;/P&gt;
&lt;P&gt;Together, these capabilities help teams move beyond debugging individual conversations and establish a repeatable process for operating and improving voice agents in production.&lt;/P&gt;
&lt;P&gt;Voice Agents in Microsoft Foundry bring together real-time speech, knowledge, tools, deployment workflows, observability, and evaluation into a unified experience. Whether you're adding voice to customer support, employee assistants, learning applications, or industry workflows, Voice agents provide the recommended path for building production-grade voice experiences on Microsoft Foundry.&lt;/P&gt;
&lt;H3&gt;Get started with voice agents in Microsoft Foundry&lt;/H3&gt;
&lt;P&gt;Getting started only takes a few steps:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Open the &lt;A class="lia-external-url" href="https://aka.ms/voice-agent" target="_blank" rel="noopener"&gt;&lt;STRONG&gt;Microsoft Foundry portal&lt;/STRONG&gt;&lt;/A&gt; and create a new agent, select “&lt;STRONG&gt;Voice”. &lt;/STRONG&gt;&lt;/LI&gt;
&lt;LI&gt;Choose a voice-optimized model and configure your agent's instructions.&lt;/LI&gt;
&lt;LI&gt;Connect knowledge sources and tools as needed for your scenario.&lt;/LI&gt;
&lt;LI&gt;Test conversations directly in the built-in voice playground.&lt;/LI&gt;
&lt;LI&gt;Deploy to your desired channels and monitor performance using built-in observability and evaluation capabilities.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="footer"&gt;&lt;SUP&gt;1&lt;/SUP&gt;Study done from &lt;/SPAN&gt;&lt;/SPAN&gt;&lt;A href="https://www.globenewswire.com/Tracker?data=54aYugxjNco3kri42i6u7AD_dLKGnVpAz-7sfaDfVLtnMmPoNONIZcLsT_TvAMgaHonPjnSTJw3xaoVSfZJiRw==" target="_blank" rel="noopener"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="Hyperlink"&gt;Jabra&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-parastyle="footer"&gt; and LSE. &lt;/SPAN&gt;&lt;/SPAN&gt;&lt;A href="https://www.theglobeandmail.com/investing/markets/markets-news/GlobeNewswire/35455944/using-your-voice-will-completely-reshape-how-we-work-with-ai-new-study-finds/" target="_blank" rel="noopener"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-charstyle="cf01" data-ccp-charstyle-defn="{&amp;quot;ObjectId&amp;quot;:&amp;quot;1f4ac159-8199-5d0e-90b3-48a1825ddce0|1&amp;quot;,&amp;quot;ClassId&amp;quot;:1073872969,&amp;quot;Properties&amp;quot;:[201342446,&amp;quot;1&amp;quot;,201342447,&amp;quot;5&amp;quot;,201342448,&amp;quot;3&amp;quot;,201342449,&amp;quot;1&amp;quot;,469777841,&amp;quot;Segoe UI&amp;quot;,469777842,&amp;quot;Segoe UI&amp;quot;,469777843,&amp;quot;游明朝&amp;quot;,469777844,&amp;quot;Segoe UI&amp;quot;,201341986,&amp;quot;1&amp;quot;,469769226,&amp;quot;Segoe UI&amp;quot;,268442635,&amp;quot;18&amp;quot;,469775450,&amp;quot;cf01&amp;quot;,201340122,&amp;quot;1&amp;quot;,134233614,&amp;quot;true&amp;quot;,469778129,&amp;quot;cf01&amp;quot;,335572020,&amp;quot;1&amp;quot;,469778324,&amp;quot;Default Paragraph Font&amp;quot;]}"&gt;The Globe and Mail&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN data-contrast="auto"&gt;&lt;SPAN data-ccp-charstyle="cf01"&gt; (October 2025)&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Mon, 28 Sep 2026 18:09:24 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/introducing-voice-agents-in-microsoft-foundry/ba-p/4557276</guid>
      <dc:creator>Yiwen_Rong</dc:creator>
      <dc:date>2026-09-28T18:09:24Z</dc:date>
    </item>
    <item>
      <title>Insights in Foundry Turns Agent Traces into Action</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/insights-in-foundry-turns-agent-traces-into-action/ba-p/4559634</link>
      <description>&lt;P&gt;If you’ve shipped an agent to production, you’ve probably watched it repeat a mistake that no dashboard flagged and no evaluation caught. A single agent can generate thousands of traces, and predefined evaluations can only detect the problems a team already knew to look for. That leaves developers searching manually for patterns, reconstructing failures, and deciding what deserves attention.&lt;/P&gt;
&lt;P&gt;Today, we’re announcing the public preview of &lt;A href="https://learn.microsoft.com/azure/foundry/observability/how-to/agent-insights" target="_blank" rel="noopener"&gt;Insights in Foundry&lt;/A&gt;. Insights analyzes agent traces to surface recurring and previously unknown behavior, explains the likely cause with supporting evidence, and recommends what developers should do next.&lt;/P&gt;
&lt;P&gt;Insights builds on the tracing, monitoring, and evaluation capabilities already in Foundry. It helps teams move from observing production activity to understanding where an agent needs attention, and whether the right response is a targeted fix, a new evaluation, or deeper iteration with &lt;A href="https://learn.microsoft.com/azure/foundry/agents/concepts/agent-optimizer-overview" target="_blank" rel="noopener"&gt;Agent Optimizer&lt;/A&gt;.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Discover behavior you didn’t know to look for&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Consider an order-support agent. When an order-status lookup fails, the agent still tells the customer that the order was delivered. The answer sounds confident, and the request appears to have completed. In a single trace, it looks like an isolated mistake. Across hundreds of requests, it becomes a pattern the team needs to understand.&lt;/P&gt;
&lt;P&gt;Insights analyzes the agent traces stored in the &lt;A href="https://learn.microsoft.com/azure/azure-monitor/app/app-insights-overview" target="_blank" rel="noopener"&gt;Azure Monitor Application Insights&lt;/A&gt; resource connected to your Foundry project. You don’t need to create an evaluator first. Insights looks across executions for patterns that would be difficult to find by reading traces one at a time.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;Depending on the available evidence and configuration, findings can cover tool-call failures, context handling, output quality, reliability, latency, or unnecessary token use. The value isn’t another count of errors. Each finding explains a recurring behavior, shows the evidence behind it, and points to a useful next step.&lt;/P&gt;
&lt;P&gt;Insights complements monitoring and evaluation rather than replacing them. Your existing evaluators keep testing the requirements you already know about, while Insights helps identify new behaviors worth investigating and turning into future evaluation coverage.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&amp;nbsp;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Understand each finding through its evidence&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Each Insight brings a finding and its supporting traces together in one place. It can include a description, category, severity, status, affected agent version, linked traces, representative examples, and a likely cause or recommended action.&lt;/P&gt;
&lt;P&gt;In the order-support example, a useful finding goes beyond “the lookup tool failed.” The representative traces show the lookup failure followed by the unsupported delivery claim. A developer can inspect the request, the tool calls and their results, and the final response, then compare them with successful executions.&lt;/P&gt;
&lt;P&gt;That distinction matters. An intermediate tool error isn’t automatically a failed task, and a plausible explanation isn’t a confirmed root cause. Insights gives developers a grounded starting point for investigation. The team then validates whether the behavior is real, how it affects the workflow, and who owns the fix.&lt;/P&gt;
&lt;P&gt;The quality of the evidence depends on the quality of the telemetry. Missing agent identity, version information, spans, or content can limit the analysis. Infrastructure availability, dependency outages, quotas, and service-level alerts remain with Azure Monitor and the teams that own those systems.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Choose the right action for each finding&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;A confirmed finding can lead to different responses:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Make a targeted change. Correct an instruction, tool description, code path, or workflow when the evidence points to a specific problem.&lt;/LI&gt;
&lt;LI&gt;Add evaluation coverage. Create or update an evaluator or dataset so the behavior becomes an explicit test before and after a change.&lt;/LI&gt;
&lt;LI&gt;Route the issue to its owner. If a dependency or platform failure is the cause, work with that owner instead of changing an agent that can’t fix the underlying issue.&lt;/LI&gt;
&lt;LI&gt;Explore alternatives. Use an optimization workflow such as Agent Optimizer when improving the behavior means generating and comparing multiple candidates.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;For prompt agents and code-based hosted agents, an Insight can include a proposed prompt or code change. When editable instructions or code aren’t available, the recommendation provides investigation or remediation guidance instead.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Insights proposes evidence-backed recommendations. Developers decide what changes and when. Nothing ships without human review.&lt;/P&gt;
&lt;P&gt;Insights doesn’t open pull requests, build evaluators, or deploy fixes on its own. Developers review each proposal and apply it through their normal source-control, evaluation, approval, and deployment processes.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Use Agent Optimizer when the problem calls for iteration&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Some findings call for a focused correction. Others raise a broader question: which change will improve this behavior without introducing a regression somewhere else?&lt;/P&gt;
&lt;P&gt;Agent Optimizer is built for that question. It generates candidate improvements and evaluates them against the goals you set. A confirmed Insight helps a developer define the behavior to improve, select representative evidence, and decide how to measure success before starting an optimization run. The two capabilities are complementary. Insights discovers and helps diagnose, and Agent Optimizer supports iterative improvement. Not every Insight needs an optimization run, and optimization doesn’t need to begin with an Insight. Moving from one to the other is a developer decision, not an automatic handoff.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;In the order-support example, the team might test whether an instruction change prevents unsupported delivery claims while preserving correct answers when the lookup succeeds. Whether they write the change directly or explore options with Agent Optimizer, they validate the result before deployment and review fresh production evidence afterward.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Start with one agent and one finding&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Pick one finding your team can validate. Inspect its representative traces, choose the right response, and test whether the change improves the behavior. The first scan reviews the last seven days of traces, and you can run analysis again on demand after that. A scan that returns no findings doesn’t mean an agent is free of problems, so keep your existing monitoring and evaluations in place.&lt;/P&gt;
&lt;P&gt;Insights is decision support. It doesn’t replace production monitoring, evaluation, security review, or change management. Before you run a scan, check the &lt;A href="https://learn.microsoft.com/azure/foundry/observability/how-to/agent-insights" target="_blank" rel="noopener"&gt;preview documentation&lt;/A&gt; for supported configurations, limitations, and the model and telemetry costs that analysis can incur.&lt;/P&gt;
&lt;P&gt;The goal isn’t more findings to read. It’s a repeatable way to turn production evidence into the next improvement decision, and then confirm that decision with new evidence.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;What’s next&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&amp;nbsp;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;EM&gt;Learn more&lt;/EM&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;
&lt;P&gt;Explore the full set of Microsoft Foundry capabilities announced today for operating agents in production: &lt;A href="https://aka.ms/foundrysept2026" target="_blank" rel="noopener"&gt;https://aka.ms/foundrysept2026&lt;/A&gt;&lt;/P&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;EM&gt;Join us for Foundry Friday AMA | September 25, 2026&lt;/EM&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Bring your questions. The team behind Agent Optimizer and Insights in Foundry is hosting a live AMA this Friday. Post your questions now or ask them live during the session: &lt;A href="https://aka.ms/FoundryFriday-Optimizer-Insights" target="_blank" rel="noopener"&gt;https://aka.ms/FoundryFriday-Optimizer-Insights&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;EM&gt;See it live at Microsoft Ignite&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;Join us at &lt;A href="https://ignite.microsoft.com/en-US/home" target="_blank" rel="noopener"&gt;Microsoft Ignite&lt;/A&gt;, November 17–20 in San Francisco, to explore the latest Microsoft Foundry innovations and see AI agents in action. Sessions to check out:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://ignite.microsoft.com/en-US/sessions/1787856610944001ErzW-1789751879577001dQb6?source=sessions" target="_blank" rel="noopener"&gt;From traces to better agents: The Foundry improvement loop&lt;/A&gt;&amp;nbsp;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://ignite.microsoft.com/en-US/sessions/1787856611240001E5kl-1789751880030001dFks?source=sessions" target="_blank" rel="noopener"&gt;Agents in production: Three failures, three fixes&lt;/A&gt;&amp;nbsp;&lt;/LI&gt;
&lt;/UL&gt;</description>
      <pubDate>Thu, 24 Sep 2026 17:30:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/insights-in-foundry-turns-agent-traces-into-action/ba-p/4559634</guid>
      <dc:creator>katelynrothney</dc:creator>
      <dc:date>2026-09-24T17:30:00Z</dc:date>
    </item>
    <item>
      <title>DeepSeek-V4.1-Flash is coming to Microsoft Foundry</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/deepseek-v4-1-flash-is-coming-to-microsoft-foundry/ba-p/4556431</link>
      <description>&lt;P&gt;Agentic applications increasingly need to process entire code repositories, large document collections, images, tool results, and long-running conversation histories. As that context grows, the key-value, or KV, cache maintained during inference can become a major constraint on cost, memory, and deployment scale.&lt;/P&gt;
&lt;P&gt;DeepSeek-V4.1-Flash is coming to Foundry Models. This multimodal Mixture-of-Experts model supports text and image inputs, a context window of up to one million tokens, and continuously configurable reasoning effort. Developers can deploy the model either &lt;STRONG&gt;Direct from Azure &lt;/STRONG&gt;(public preview) or through &lt;STRONG&gt;Fireworks on Foundry &lt;/STRONG&gt;(generally available), choosing the consumption path that best fits their performance, operational, and purchasing requirements.&lt;/P&gt;
&lt;H4&gt;More capability and control for agentic applications&lt;/H4&gt;
&lt;P&gt;DeepSeek-V4.1-Flash is more than a fine-tuned version of &lt;A href="https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/introducing-deepseek-v4-flash-and-v4-pro-in-microsoft-foundry/4515174" target="_blank" rel="noopener"&gt;V4-Flash&lt;/A&gt;. It introduces a &lt;STRONG&gt;Causal Encoder-Decoder architecture&lt;/STRONG&gt; in which a 20-layer causal encoder processes the input and a 20-layer decoder generates the output. The model has a 552-billion-parameter backbone, activating 8 billion parameters during prefill and 16 billion during decoding.&lt;/P&gt;
&lt;P&gt;The model also expands beyond the text-only capabilities of V4-Flash. It adds native image understanding and replaces three fixed reasoning modes with a continuous &lt;STRONG&gt;reasoning-effort setting from 1 to 100&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P&gt;Developers can use lower effort for tasks such as routing, classification, and extraction, then increase effort for planning, coding, debugging, and multi-step tool use, without changing model deployments.&lt;/P&gt;
&lt;H4&gt;Efficient long-context AI starts with a smaller memory footprint&lt;/H4&gt;
&lt;P&gt;Long-context agents need to retain code, documents, images, tool results, and conversation history during inference. As that context grows, the KV cache can become a significant memory and scaling constraint.&lt;/P&gt;
&lt;P&gt;DeepSeek reports that V4.1-Flash requires &lt;STRONG&gt;890 bytes of global KV cache per token&lt;/STRONG&gt;, approximately one quarter of V4-Flash’s footprint. Its architecture is particularly relevant to workloads that read significantly more information than they generate.&lt;/P&gt;
&lt;img&gt;&lt;EM&gt;Figure 1. Global KV cache per token across DeepSeek model generations. Values and fold reductions as reported by DeepSeek.&lt;/EM&gt;&lt;/img&gt;
&lt;P&gt;For developers, the practical opportunity is to retain more useful context and support context-heavy workloads more efficiently. Actual latency, concurrency, infrastructure requirements, and cost will depend on the application, serving configuration, traffic profile, and deployment path.&lt;/P&gt;
&lt;H4&gt;Stronger performance on agentic workloads&lt;/H4&gt;
&lt;P&gt;DeepSeek-V4.1-Flash pairs its more efficient long-context architecture with improved performance on agentic tasks. On DeepSeek’s published evaluations, V4.1-Flash improves on V4-Flash across the reported agentic benchmarks, including coding, terminal, security, tool-use, and automation scenarios.&lt;/P&gt;
&lt;img&gt;&lt;EM&gt;Figure 2. DeepSeek-V4-Flash, V4-Pro, and V4.1-Flash on agentic and reasoning benchmarks, all at maximum reasoning effort. Sorted by the gain from V4-Flash to V4.1-Flash.&lt;/EM&gt;&lt;/img&gt;
&lt;P&gt;On DeepSWE v1.1, V4.1-Flash resolves &lt;STRONG&gt;74.2% of tasks&lt;/STRONG&gt;, compared with &lt;STRONG&gt;54.4% for V4-Flash&lt;/STRONG&gt;. DeepSeek also reports improvements on Terminal-Bench 2.1, CyberGym, AutomationBench, and Agent’s Last Exam. Results can vary based on prompts, orchestration framework, context construction, deployment, and workload. Teams should evaluate model quality, latency, and cost using their own application data. Competitor or comparative figures reported by DeepSeek have not been independently verified by Microsoft.&lt;/P&gt;
&lt;H4&gt;What this means for developers&lt;/H4&gt;
&lt;P&gt;Put together, the architecture and the benchmark profile point to a specific set of workloads where DeepSeek-V4.1-Flash is a strong fit:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Repository-scale coding agents&lt;/STRONG&gt; that load large codebases, run tools in a terminal, and iterate over many steps.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Document-heavy and multimodal analysis&lt;/STRONG&gt;, including PDFs, charts, screenshots, and scanned forms, where image inputs and a 1M-token window remove the need for aggressive chunking.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Long-running automation and research agents&lt;/STRONG&gt; that accumulate tool outputs and conversation history over hundreds of turns.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Cost-sensitive, high-concurrency deployments&lt;/STRONG&gt; where the smaller KV cache translates directly into more sessions per GPU.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The reasoning effort control deserves special attention. Rather than choosing between a fast model and a thinking model, you can run one deployment and set effort per request: low for routing and extraction, high for planning and debugging. That simplifies orchestration compared with the V4 Flash and V4 Pro pairing, and it gives you a single, continuous knob to tune against quality and latency targets in Foundry evaluations.&lt;/P&gt;
&lt;H4&gt;One model, two deployment paths&lt;/H4&gt;
&lt;P&gt;As the open model ecosystem evolves, teams increasingly want flexibility not only in which model they use, but also in how they deploy and consume it. With DeepSeek-V4.1-Flash, developers can select between two paths in Microsoft Foundry based on the needs of their application.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 72.1296%; border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Deployment path&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Consider this path when you prioritize&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Direct from Azure&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;A native Azure-hosted and Microsoft-supported model offering, billed through your Azure subscription and covered by Azure service-level agreements.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Fireworks on Foundry&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Optimized open-model inference, token caching, broader Provisioned Throughput support, or custom-weight deployment capabilities on Fireworks AI's inference stack running natively in Azure.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;The right path depends on the workload. Developers should evaluate model quality, latency, throughput, deployment shape, data requirements, and cost using their own application traffic and evaluation datasets.&lt;/P&gt;
&lt;H5&gt;Other models now available through Fireworks on Foundry&lt;/H5&gt;
&lt;P&gt;Microsoft Foundry is also expanding model choice through Fireworks with additional global deployment options for leading open models. These updates give developers more flexibility to evaluate and deploy models for agentic, coding, reasoning, and long-context workloads through the same Foundry experience.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;DeepSeek-V4.1-Flash is available through Fireworks AI with PayGo Global and Provisioned Throughput deployment options.&lt;/LI&gt;
&lt;LI&gt;GLM-5.3-Flash is available through Fireworks with a Global deployment option.&lt;/LI&gt;
&lt;LI&gt;Kimi K3 now adds a Global deployment option through Fireworks, expanding beyond its previously available Data Zone deployment.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Availability, deployment options, and pricing may vary by model. Review each model card in the Microsoft Foundry model catalog for the latest details before deployment.&lt;/P&gt;
&lt;H5&gt;Pricing&lt;/H5&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 75.6481%; border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Fireworks on Foundry&lt;/td&gt;&lt;td&gt;Deployment Region&lt;/td&gt;&lt;td&gt;Input/M tokens&lt;/td&gt;&lt;td&gt;Cache/M tokens&lt;/td&gt;&lt;td&gt;Output/M tokens&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;GLM 5.3&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Standard Global&lt;/td&gt;&lt;td&gt;$1.75&lt;/td&gt;&lt;td&gt;$0.33&lt;/td&gt;&lt;td&gt;$5.50&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;GLM 5.3&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;US Data Zone&lt;/td&gt;&lt;td&gt;$2.10 &amp;nbsp;&lt;/td&gt;&lt;td&gt;$0.39 &amp;nbsp;&lt;/td&gt;&lt;td&gt;$6.60 &amp;nbsp;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;GLM 5.3 Flash&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Standard Global&lt;/td&gt;&lt;td&gt;$0.19 &amp;nbsp;&lt;/td&gt;&lt;td&gt;$0.04 &amp;nbsp;&lt;/td&gt;&lt;td&gt;$6.25 &amp;nbsp;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Kimi K3&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Standard Global&lt;/td&gt;&lt;td&gt;$3.00 &amp;nbsp;&lt;/td&gt;&lt;td&gt;$0.30 &amp;nbsp;&lt;/td&gt;&lt;td&gt;$15.00 &amp;nbsp;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Kimi K3&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;US Data Zone&lt;/td&gt;&lt;td&gt;$3.30 &amp;nbsp;&lt;/td&gt;&lt;td&gt;$0.33 &amp;nbsp;&lt;/td&gt;&lt;td&gt;$16.50 &amp;nbsp;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 22.1759%" /&gt;&lt;col style="width: 22.1759%" /&gt;&lt;col style="width: 22.1759%" /&gt;&lt;col style="width: 15.8063%" /&gt;&lt;col style="width: 17.5189%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H4&gt;DeepSeek V4.1 Flash Pricing&lt;/H4&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Deployment path&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Deployment Region&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Input/M tokens&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Output/M tokens&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Cache/M tokens&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Direct From Azure&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Global Standard&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$0.3&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$1.2&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$0.006&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Fireworks on Foundry&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;Standard Global&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$0.37&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$1.50&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;$0.007&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 20.00%" /&gt;&lt;col style="width: 20.00%" /&gt;&lt;col style="width: 20.00%" /&gt;&lt;col style="width: 20.00%" /&gt;&lt;col style="width: 20.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;H4&gt;A unified Foundry experience&lt;/H4&gt;
&lt;P&gt;Whichever deployment path teams choose, Microsoft Foundry provides a common environment for discovering models, comparing options, evaluating them against application-specific data, and managing AI development with integrated governance and observability.&lt;/P&gt;
&lt;P&gt;This lets developers focus on selecting the right model and deployment path for their scenario instead of assembling fragmented tools and separate operational environments. Model discovery, evaluation, and management stay centered in Foundry, while the consumption path adapts to the team.&lt;/P&gt;
&lt;H4&gt;Getting started&lt;/H4&gt;
&lt;P&gt;Explore DeepSeek-V4.1-Flash in the Microsoft Foundry model catalog through either deployment path:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;DeepSeek-V4.1-Flash, Direct from Azure: &lt;SPAN data-teams="true"&gt;&lt;A href="https://ai.azure.com/catalog/models/DeepSeek-V4.1-Flash?search=DeepSeek-v4.1" target="_blank" rel="noopener" aria-label="Link DeepSeek-V4.1-Flash | Model Catalog | Microsoft Foundry"&gt;DeepSeek-V4.1-Flash | Model Catalog | Microsoft Foundry&lt;/A&gt;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;DeepSeek-V4.1-Flash, Fireworks on Foundry: &lt;A href="https://ai.azure.com/catalog/models/FW-DeepSeek-V4.1-Flash?search=deepseek+v4.1" target="_blank" rel="noopener"&gt;FW-DeepSeek-V4.1-Flash | Model Catalog | Microsoft Foundry&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;You can then:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Compare the available deployment paths&lt;/LI&gt;
&lt;LI&gt;Evaluate the model using your own datasets and agent scaffolds, sweeping the reasoning effort setting&lt;/LI&gt;
&lt;LI&gt;Select the deployment option aligned with your operational and performance requirements&lt;/LI&gt;
&lt;LI&gt;Begin integrating the model into long-context, multimodal, and agentic applications&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;For the full architecture description and complete benchmark tables, see the &lt;A href="https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash" target="_blank" rel="noopener"&gt;DeepSeek-V4.1-Flash model card on Hugging Face&lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Fri, 25 Sep 2026 15:58:46 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/deepseek-v4-1-flash-is-coming-to-microsoft-foundry/ba-p/4556431</guid>
      <dc:creator>RashaudSavage</dc:creator>
      <dc:date>2026-09-25T15:58:46Z</dc:date>
    </item>
    <item>
      <title>Claude Opus 5.5 comes to Microsoft Foundry for long-running coding and knowledge work</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/claude-opus-5-5-comes-to-microsoft-foundry-for-long-running/ba-p/4558051</link>
      <description>&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;AI models are increasingly taking on work that extends far beyond a single prompt: building a feature across a codebase, investigating a complex issue, synthesizing hundreds of pages of information, or working through a multi-step business process.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;As that work gets longer, raw intelligence is only part of what matters. The model also needs to stay focused, make good decisions along the way, communicate what it is doing, and produce work that people can quickly review and use. Today, &lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;Claude Opus 5.5 is available in Microsoft Foundry&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;, bringing Anthropic’s most capable Opus model to developers and enterprises building AI applications and agents.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Claude Opus 5.5 is designed for everyday complex work. It advances Opus 5 across agentic coding, knowledge work, and long-running tasks while making it easier for people to understand what the model did, what it found, and what it needs next. &lt;/SPAN&gt;&lt;SPAN data-contrast="none"&gt;Claude Opus 5.5 also does more with fewer tokens. Lower per-token prices and much cheaper cache reads stack on top of the efficiency gains.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P aria-level="2"&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;Built for work that takes time&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:299,&amp;quot;335559739&amp;quot;:299}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Writing a function is one thing. Building a feature that touches multiple services, tracing a production issue across a large repository, or carrying a task from investigation through implementation and validation is something else entirely. Claude Opus 5.5 is designed for these longer-running workflows.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;For software development, it can work through long-running coding tasks such as &lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;building features across a codebase, debugging, refactoring, and reviewing code&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;. &lt;/SPAN&gt;&lt;SPAN data-contrast="none"&gt; It finds the root cause before changing anything, checks its work as it goes, and explains its changes in plain language, so engineers can review and trust them quickly.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;That combination becomes particularly valuable when developers use models through agentic coding environments, where a session may involve dozens of steps and run for an extended period of time. The same applies beyond software development.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;For knowledge workers, Claude Opus 5.5 can bring together information from multiple sources, work through long documents and spreadsheets, perform analysis, and help create artifacts such as &lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;memos, reports, and presentations&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;. Compared with Opus 5,&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;it&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt; produc&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;es &lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;outputs that require less editing before they are ready to share.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P aria-level="2"&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;An AI model that communicates more like a teammate&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:299,&amp;quot;335559739&amp;quot;:299}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;As agents take on more autonomous work, another challenge emerges: &lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;keeping the human in the loop without overwhelming them.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;An agent that performs 50 steps should not require someone to inspect 50 steps to understand whether the work was successful. Claude Opus 5.5 introduces improvements to agentic communication designed to make long-running work easier to follow. As it works, the model can surface the information that matters most:&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559683&amp;quot;:0,&amp;quot;335559684&amp;quot;:-2,&amp;quot;335559685&amp;quot;:1080,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="1" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;What it did&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559739&amp;quot;:0}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559683&amp;quot;:0,&amp;quot;335559684&amp;quot;:-2,&amp;quot;335559685&amp;quot;:1080,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="2" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;What it found&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559739&amp;quot;:0}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559683&amp;quot;:0,&amp;quot;335559684&amp;quot;:-2,&amp;quot;335559685&amp;quot;:1080,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="3" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;What decisions it made&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559739&amp;quot;:0}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559683&amp;quot;:0,&amp;quot;335559684&amp;quot;:-2,&amp;quot;335559685&amp;quot;:1080,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="4" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;Where it needs input from the user&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559739&amp;quot;:0}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;UL&gt;
&lt;LI aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&amp;quot;335552541&amp;quot;:1,&amp;quot;335559683&amp;quot;:0,&amp;quot;335559684&amp;quot;:-2,&amp;quot;335559685&amp;quot;:1080,&amp;quot;335559991&amp;quot;:360,&amp;quot;469769226&amp;quot;:&amp;quot;Symbol&amp;quot;,&amp;quot;469769242&amp;quot;:[8226],&amp;quot;469777803&amp;quot;:&amp;quot;left&amp;quot;,&amp;quot;469777804&amp;quot;:&amp;quot;&amp;quot;,&amp;quot;469777815&amp;quot;:&amp;quot;hybridMultilevel&amp;quot;}" data-aria-posinset="5" data-aria-level="1"&gt;&lt;SPAN data-contrast="auto"&gt;What happened at the end of a long-running task&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559739&amp;quot;:0}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;The goal is simple: spend less time decoding what the model did and more time using the result. This matters particularly for enterprise agents, where users need to understand not only the final answer but also when an agent needs clarification, encounters a constraint, or reaches a decision point that requires human judgment.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P aria-level="2"&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;Adpative&lt;/SPAN&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt; thinking&lt;/SPAN&gt; &lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:299,&amp;quot;335559739&amp;quot;:299}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Claude Opus 5.5 uses&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;adaptive thinking&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;, automatically determining how much reasoning a task requires. Rather than turning thinking on or off or manually specifying a thinking-token budget, developers use &lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;effort&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt; to influence how much work the model should put into a request. This allows the model to adapt its reasoning to the task at hand—from relatively straightforward requests to complex problems that require deeper analysis.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;For developers building agents, this can reduce the amount of application logic needed to decide when and how a model should reason.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P aria-level="2"&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;Designed for long-running agent architectures&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:299,&amp;quot;335559739&amp;quot;:299}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Long-running agents create challenges beyond model intelligence. Conversations can exceed context limits. Tools available to an agent can change. Applications may need to compact earlier context while preserving the model's understanding of the work already completed.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Alongside Claude Opus 5.5, Anthropic is introducing beta API capabilities designed for these scenarios, including &lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;asynchronous compaction, keep-tail compaction, and changing tools during a conversation while preserving thinking and prompt caching&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;. These capabilities can help agent developers maintain continuity across longer tasks without rebuilding the state of the application every time context or available tools change.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Combined with Microsoft Foundry, developers can use Claude Opus 5.5 as part of broader agent systems that connect models with enterprise data, tools, evaluation, and operational workflows.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P aria-level="2"&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;Expanded safeguards for more capable models&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:299,&amp;quot;335559739&amp;quot;:299}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;As model capabilities increase, Anthropic is also expanding the safeguards applied to Claude Opus 5.5.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Claude Opus 5.5 is the first Opus model to use safety classifiers like those introduced with Claude Fable 5.1 in areas including &lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;cybersecurity, biology, AI development, and distillation.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;For common developer, educational, and knowledge-work scenarios, customers can continue using the model for tasks such as identifying software vulnerabilities or learning about biological concepts. Certain requests that Anthropic identifies as higher-risk or dual-use may be handled by another Claude model with the appropriate safeguards.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;This reflects an increasingly important part of deploying more capable models: advancing what models can do while applying safeguards appropriate to the capabilities they introduce.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;Pricing&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;BR /&gt;&lt;table border="1" style="width: 100.001%; height: 104.046px; border-width: 1px;"&gt;&lt;colgroup&gt;&lt;col style="width: 20%" /&gt;&lt;col style="width: 20%" /&gt;&lt;col style="width: 20%" /&gt;&lt;col style="width: 20%" /&gt;&lt;col style="width: 20%" /&gt;&lt;/colgroup&gt;&lt;tbody&gt;&lt;tr style="height: 34.6821px;"&gt;&lt;td style="height: 34.6821px;"&gt;Model&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;Deployment Offers&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;Input/M Tokens&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;Output/M Tokens&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;Availability&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 34.6821px;"&gt;&lt;td style="height: 34.6821px;"&gt;Claude Opus 5.5&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;Global Standard, US DataZone&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;
&lt;P&gt;$4&lt;/P&gt;
&lt;P&gt;Cache Hit - $0.20&lt;/P&gt;
&lt;P&gt;Cache Write - $5&lt;/P&gt;
&lt;P&gt;Cache Write (1 hr) - $8&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;$20&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;GA, Hosted on Azure&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 34.6821px;"&gt;&lt;td style="height: 34.6821px;"&gt;Claude Opus 5.5 (Long Context)&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;Global Standard, US DataZone&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;
&lt;P&gt;$4&lt;/P&gt;
&lt;P&gt;Cache Hit - $0.20&lt;/P&gt;
&lt;P&gt;Cache Write - $5&lt;/P&gt;
&lt;P&gt;Cache Write (1 hr) - $8&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;$20&lt;/td&gt;&lt;td style="height: 34.6821px;"&gt;GA, Hosted on Azure&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P aria-level="2"&gt;&lt;STRONG&gt;&lt;SPAN data-contrast="none"&gt;&lt;SPAN data-ccp-parastyle="heading 2"&gt;Build with Claude Opus 5.5 in Microsoft Foundry&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;134245418&amp;quot;:true,&amp;quot;134245529&amp;quot;:true,&amp;quot;335559738&amp;quot;:299,&amp;quot;335559739&amp;quot;:299}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-contrast="auto"&gt;Choosing a model is only the beginning of putting AI into production.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;Microsoft Foundry gives developers a unified place to discover models, build and evaluate AI applications and agents, connect them with enterprise data and tools, and operate those systems in production.&lt;/SPAN&gt; &lt;SPAN data-contrast="auto"&gt;As models become capable of taking on more complete units of work, the question is shifting from &lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;Can the model answer this prompt?&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt; to &lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;Can I trust it to carry the work forward?&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt; &amp;nbsp;&lt;/SPAN&gt;&lt;SPAN data-contrast="auto"&gt;Claude Opus 5.5 represents another step in that direction: stronger performance on complex work, more adaptive reasoning, and clearer communication between people and the AI systems working alongside them.&lt;/SPAN&gt;&lt;SPAN data-ccp-props="{&amp;quot;335559738&amp;quot;:240,&amp;quot;335559739&amp;quot;:240}"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;A class="lia-external-url" href="https://ai.azure.com/catalog/models/claude-opus-5-5?search=opus+" target="_blank"&gt;Claude Opus 5.5 is available today &lt;/A&gt;&amp;nbsp;in Microsoft Foundry.&lt;/P&gt;</description>
      <pubDate>Tue, 22 Sep 2026 16:47:05 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/claude-opus-5-5-comes-to-microsoft-foundry-for-long-running/ba-p/4558051</guid>
      <dc:creator>amar_badal</dc:creator>
      <dc:date>2026-09-22T16:47:05Z</dc:date>
    </item>
    <item>
      <title>Demystifying Dynamic AI Routing: Exploring the model router in Foundry Models with RouteLab</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/demystifying-dynamic-ai-routing-exploring-the-model-router-in/ba-p/4557388</link>
      <description>&lt;P data-path-to-node="1"&gt;As enterprise generative AI applications mature, engineering teams face a continuous balancing act between cost, latency, and response quality. Hard-coding a frontier model for every request ensures high quality but maximizes token costs and latency. Conversely, defaulting to smaller, faster models risks degradation in accuracy for complex reasoning tasks.&lt;/P&gt;
&lt;P data-path-to-node="2"&gt;To solve this architectural challenge, developers can use the&amp;nbsp;&lt;A class="lia-external-url" href="https://aka.ms/trymodelrouter" target="_blank" rel="noopener"&gt;&lt;STRONG&gt;model router in Foundry Models&lt;/STRONG&gt;&lt;/A&gt;—an intelligent, real-time routing engine. By combining this capability with testing environments like &lt;STRONG data-path-to-node="2" data-index-in-node="185"&gt;RouteLab&lt;/STRONG&gt;, developers can eliminate the guesswork of model selection, visually evaluate routing decisions on custom datasets, and build highly optimized AI systems.&lt;/P&gt;
&lt;H2 data-path-to-node="3"&gt;&lt;STRONG&gt;What is the model router in Foundry Models?&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P data-path-to-node="4"&gt;The model router is a purpose-built machine-learning model that sits between your application and a configured pool of large language models (LLMs). Rather than relying on rigid, hard-coded rules, the router analyzes the complexity, required reasoning capabilities, and context of each incoming prompt in real time to select the most suitable LLM.&lt;/P&gt;
&lt;P data-path-to-node="5"&gt;By targeting a single deployment, developers gain access to three distinct routing modes:&lt;/P&gt;
&lt;UL data-path-to-node="6"&gt;
&lt;LI&gt;
&lt;P&gt;&lt;STRONG data-path-to-node="6,0,0" data-index-in-node="0"&gt;Balanced (Default):&lt;/STRONG&gt; Evaluates all models within a narrow quality margin and selects the most cost-effective option for general-purpose scenarios.&lt;/P&gt;
&lt;/LI&gt;
&lt;LI&gt;
&lt;P&gt;&lt;STRONG data-path-to-node="6,1,0" data-index-in-node="0"&gt;Cost:&lt;/STRONG&gt; Broadens the acceptable quality band to aggressively minimize token costs—ideal for high-volume, budget-sensitive workloads.&lt;/P&gt;
&lt;/LI&gt;
&lt;LI&gt;
&lt;P&gt;&lt;STRONG data-path-to-node="6,2,0" data-index-in-node="0"&gt;Quality:&lt;/STRONG&gt; Always routes to the highest-performing model for that specific prompt, ignoring cost implications for complex reasoning tasks.&lt;/P&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2 data-path-to-node="7"&gt;&lt;STRONG&gt;Getting Started: Running RouteLab Locally&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P data-path-to-node="8"&gt;Before diving into the visual evaluations, you need to spin up RouteLab in your local environment. The model-router-playground repository is designed for a frictionless developer experience right out of the box.&lt;/P&gt;
&lt;H6 data-path-to-node="8"&gt;&lt;STRONG&gt;Part 1: Setting Up the Environment&lt;/STRONG&gt;&lt;/H6&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="9,0,0" data-index-in-node="0"&gt;Clone the Repository and Open in VS Code:&lt;/STRONG&gt; Start by cloning the repository to your local machine. Open the cloned folder directly inside Visual Studio Code (VS Code) so that all project assets, scripts, and dependencies are loaded in your editor workspace.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Launch the Script:&lt;/STRONG&gt; Open the integrated terminal in VS Code and execute the setup script by running&lt;/LI&gt;
&lt;/OL&gt;
&lt;LI-CODE lang=""&gt;git clone https://github.com/Pandey-Vikas/model-router-playground.git
cd model-router-playground
./start.ps1.&lt;/LI-CODE&gt;
&lt;P&gt;This kicks off the environment initialization, installs required dependencies, and prepares the backend routing services.&lt;/P&gt;
&lt;img /&gt;
&lt;H6&gt;&lt;STRONG&gt;Part 2: Interactive Authentication and Launch&lt;/STRONG&gt;&lt;/H6&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="9,2,0" data-index-in-node="0"&gt;Complete the 6-Step Azure Login and Launch the UI:&lt;/STRONG&gt; Once the script executes, it presents a guided, 6-step interactive authentication workflow directly in your terminal. This wizard safely connects your local setup with your Azure credentials, respects enterprise access boundaries, and wires up your designated Microsoft Foundry Model Router resource. Upon completing step 6, RouteLab automatically opens in your default browser, presenting the full testing harness.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&amp;nbsp;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;H2 data-path-to-node="7"&gt;&lt;STRONG&gt;Hands-On with RouteLab: The Developer Playground&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P data-path-to-node="8"&gt;Understanding how dynamic routing behaves in practice is critical before deploying it to production. The model-router-playground repository provides a practical, UI-driven approach to exploring this architecture through its &lt;STRONG data-path-to-node="8" data-index-in-node="224"&gt;RouteLab&lt;/STRONG&gt; interface.&lt;BR /&gt;&lt;BR /&gt;As seen in the RouteLab dashboard below, the environment is designed to demystify the routing process without requiring custom evaluation scripts. Here is how it empowers development teams:&lt;/P&gt;
&lt;H3 data-path-to-node="10"&gt;1. Interactive Manual Evaluation- "&lt;EM&gt;Supports Bring your own Dataset&lt;/EM&gt;"&lt;/H3&gt;
&lt;P data-path-to-node="11"&gt;One of the biggest questions developers have about dynamic routing is, &lt;EM data-path-to-node="20" data-index-in-node="71"&gt;"Which model will it pick for my specific prompt?"&lt;/EM&gt; RouteLab’s central &lt;STRONG data-path-to-node="20" data-index-in-node="141"&gt;Manual Evaluation&lt;/STRONG&gt; interface allows you to select your desired mode (&lt;STRONG&gt;Balanced, Cost, or Quality&lt;/STRONG&gt;) and test the router. &lt;STRONG data-path-to-node="20" data-index-in-node="258"&gt;Crucially, this is where the "Bring Your Own Data" USP truly shines.&lt;/STRONG&gt; Instead of relying solely on generic examples or typing prompts one by one, you can seamlessly import your custom dataset (CSV or JSONL) directly into the Manual UI. By running your proprietary, domain-specific questions through the interface, you gain instant, line-by-line transparency into how the router handles your exact enterprise workloads—immediately displaying the selected underlying model, token usage, and latency for your own data.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;To get a comprehensive view of the router's decision-making process, you can &lt;STRONG&gt;execute &lt;/STRONG&gt;your imported dataset &lt;STRONG&gt;three separate times&lt;/STRONG&gt;—once for &lt;STRONG&gt;each routing mode&lt;/STRONG&gt;. After completing these runs, simply navigate to the &lt;STRONG data-path-to-node="21" data-index-in-node="208"&gt;Analytics&lt;/STRONG&gt; tab and launch the &lt;STRONG data-path-to-node="21" data-index-in-node="237"&gt;Compare&lt;/STRONG&gt; button. This powerful feature allows you to evaluate all three executions side-by-side, providing a clear, empirical comparison of cost savings, latency overhead, and model selections across the Balanced, Cost, and Quality modes for your specific prompts.&lt;/P&gt;
&lt;img /&gt;&lt;img /&gt;&lt;img /&gt;&lt;img /&gt;
&lt;H3 data-path-to-node="22"&gt;2. Quick Testing with the Routing Range&lt;/H3&gt;
&lt;P data-path-to-node="23"&gt;To rapidly build intuition on how the router behaves, the UI features a &lt;STRONG data-path-to-node="23" data-index-in-node="72"&gt;"Test the routing range"&lt;/STRONG&gt; capability. With a single click—such as the "Run 15 (quick)" or "Run 30 (full)" buttons—developers can fire off a spectrum of predefined prompts from direct instructions to complex synthesis. This allows you to watch the router actively adapt its model selection across a diverse ladder of complexities.&lt;/P&gt;
&lt;H3 data-path-to-node="24"&gt;3. Auto Evaluation at Scale: Validating Custom Workloads&lt;/H3&gt;
&lt;P data-path-to-node="25"&gt;While manual testing builds intuition line-by-line, production readiness requires scale. RouteLab’s &lt;STRONG data-path-to-node="25" data-index-in-node="100"&gt;Auto Evaluation&lt;/STRONG&gt; engine allows developers to run bulk assessments. It comes pre-loaded with bundled datasets (like mixed-prompts.jsonl or java_custom.jsonl) for quick benchmarking, but &lt;STRONG data-path-to-node="25" data-index-in-node="284"&gt;its ultimate USP is executing automated evaluations on your imported custom datasets.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P data-path-to-node="26"&gt;To make formatting effortless, &lt;STRONG data-path-to-node="26" data-index-in-node="31"&gt;you can download a sample .jsonl file from the available bundled datasets directly within the UI, providing a perfect template to construct your own custom datasets.&lt;/STRONG&gt; By uploading your populated proprietary files, teams can automate the execution of hundreds of their own prompts. This ensures your routing strategy and cost projections are validated against the real-world datasets your application will handle in production, bridging the gap between generic benchmarks and enterprise reality.&lt;/P&gt;
&lt;P data-path-to-node="27"&gt;&lt;STRONG data-path-to-node="27" data-index-in-node="0"&gt;Under the Hood: The Code Behind Auto Evaluation&lt;/STRONG&gt; Beyond the UI experience, the complete underlying code for this Auto Evaluation engine is available in its own dedicated repository: &lt;STRONG data-path-to-node="27" data-index-in-node="181"&gt;&lt;A href="https://github.com/microsoft-foundry/Model-Router-Auto-Evaluation" target="_blank" rel="noopener" data-hveid="0" data-ved="0CAAQ_4QMahgKEwi77pbrifWWAxUAAAAAHQAAAAAQiws"&gt;microsoft-foundry/Model-Router-Auto-Evaluation&lt;/A&gt;&lt;/STRONG&gt;. By exploring this repository, developers can see exactly how the batch-processing, routing logic, and metric aggregations are programmatically implemented. This open-code approach makes it incredibly easy for engineering teams to lift and shift the evaluation framework directly into their own enterprise CI/CD pipelines or custom testing harnesses.&lt;/P&gt;
&lt;img /&gt;
&lt;P data-path-to-node="27"&gt;After successful completion of Auto Evaluation, you click the open button Click on the open button under Past Evaluation to get the dashboard.&lt;/P&gt;
&lt;img /&gt;&lt;img /&gt;&lt;img /&gt;
&lt;H3 data-path-to-node="28"&gt;4. Deep Dives with Analytics and Logs&lt;/H3&gt;
&lt;P data-path-to-node="29"&gt;Data drives deployment decisions. The top navigation of the RouteLab UI seamlessly transitions users from Evaluation into &lt;STRONG data-path-to-node="29" data-index-in-node="122"&gt;Analytics&lt;/STRONG&gt; and &lt;STRONG data-path-to-node="29" data-index-in-node="136"&gt;Logs&lt;/STRONG&gt;. Instead of parsing raw JSON telemetry, developers can visualize the aggregate results of their custom and bundled datasets:&lt;/P&gt;
&lt;UL data-path-to-node="30"&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="30,0,0" data-index-in-node="0"&gt;Cost Distribution:&lt;/STRONG&gt; See the exact token savings achieved by routing simpler tasks to smaller models.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="30,1,0" data-index-in-node="0"&gt;Latency Overhead:&lt;/STRONG&gt; Measure the actual response times to validate that the router's overhead is negligible.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="30,2,0" data-index-in-node="0"&gt;Model Utilization:&lt;/STRONG&gt; View visual breakdowns of how often frontier models are invoked versus specialized models under different routing modes.&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2 data-path-to-node="31"&gt;The Future is Dynamic&lt;/H2&gt;
&lt;P data-path-to-node="32"&gt;The era of static, one-size-fits-all LLM deployments are ending. By leveraging the model router and visually validating your strategy through RouteLab—using both bundled and custom datasets—your team can confidently deploy AI architectures that are highly responsive, scalable, and remarkably cost-efficient.&lt;/P&gt;
&lt;P data-path-to-node="33"&gt;Ready to explore intelligent routing? Check out the &lt;A href="https://github.com/Pandey-Vikas/model-router-playground" target="_blank" rel="noopener" data-hveid="0" data-ved="0CAAQ_4QMahgKEwi77pbrifWWAxUAAAAAHQAAAAAQjAs"&gt;model-router-playground&lt;/A&gt; on GitHub today and start optimizing your enterprise AI workloads!&lt;/P&gt;
&lt;H2 data-path-to-node="34"&gt;References &amp;amp; Resources&lt;/H2&gt;
&lt;UL data-path-to-node="35"&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="35,0,0" data-index-in-node="0"&gt;Microsoft Learn Documentation:&lt;/STRONG&gt; &lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-router" target="_blank" rel="noopener" data-hveid="0" data-ved="0CAAQ_4QMahgKEwi77pbrifWWAxUAAAAAHQAAAAAQrws"&gt;Model router in Foundry Models | Learn&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="35,1,0" data-index-in-node="0"&gt;RouteLab Developer Playground:&lt;/STRONG&gt; &lt;A href="https://github.com/Pandey-Vikas/model-router-playground" target="_blank" rel="noopener" data-hveid="0" data-ved="0CAAQ_4QMahgKEwi77pbrifWWAxUAAAAAHQAAAAAQsAs"&gt;model-router-playground on GitHub&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-path-to-node="35,2,0" data-index-in-node="0"&gt;Auto Evaluation Source Code:&lt;/STRONG&gt; &lt;A href="https://github.com/microsoft-foundry/Model-Router-Auto-Evaluation" target="_blank" rel="noopener" data-hveid="0" data-ved="0CAAQ_4QMahgKEwi77pbrifWWAxUAAAAAHQAAAAAQsQs"&gt;Model-Router-Auto-Evaluation on GitHub&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;</description>
      <pubDate>Wed, 30 Sep 2026 13:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/demystifying-dynamic-ai-routing-exploring-the-model-router-in/ba-p/4557388</guid>
      <dc:creator>vikaspandey</dc:creator>
      <dc:date>2026-09-30T13:00:00Z</dc:date>
    </item>
    <item>
      <title>A2A Endpoints and A2A Tool in Microsoft Foundry agents</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/a2a-endpoints-and-a2a-tool-in-microsoft-foundry-agents/ba-p/4557115</link>
      <description>&lt;P&gt;Agent-to-agent collaboration just became much easier to build in Microsoft Foundry. The A2A Tool and incoming A2A endpoints support A2A protocol version 1.0, which is generally available. The earlier a2a_preview tool type and protocol version 0.3 remain available in preview for existing integrations. Hosted Agents can also consume A2A tools through Foundry Toolboxes exposed over MCP.&lt;/P&gt;
&lt;P&gt;If you have been following earlier previews,&amp;nbsp;&lt;SPAN data-streamdown="strong"&gt;A2A Endpoints&lt;/SPAN&gt;&amp;nbsp;were previously known as the&amp;nbsp;&lt;SPAN data-streamdown="strong"&gt;A2A API head&lt;/SPAN&gt;. The idea is simple:&lt;/P&gt;
&lt;UL data-streamdown="unordered-list"&gt;
&lt;LI data-streamdown="list-item"&gt;&lt;SPAN data-streamdown="strong"&gt;A2A Endpoint&lt;/SPAN&gt;: expose a Foundry agent so an external agent can discover and invoke it.&lt;/LI&gt;
&lt;LI data-streamdown="list-item"&gt;&lt;SPAN data-streamdown="strong"&gt;A2A Tool&lt;/SPAN&gt;: let a Foundry agent invoke another A2A-compatible agent.&lt;/LI&gt;
&lt;LI data-streamdown="list-item"&gt;&lt;SPAN data-streamdown="strong"&gt;Hosted Agents&lt;/SPAN&gt;: use the A2A Tool through a Foundry Toolbox exposed as an MCP endpoint.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Together, these capabilities make it possible to build multi-agent systems where agents can specialize, publish their skills, and securely collaborate across service boundaries.&lt;/P&gt;
&lt;H2 data-streamdown="heading-2"&gt;Why this matters&lt;/H2&gt;
&lt;P&gt;Until now, many multi-agent patterns required custom APIs, one-off adapters, or orchestration logic tightly coupled to a specific framework. A2A gives us a standardized protocol for agent-to-agent communication.&lt;/P&gt;
&lt;P&gt;That means one agent can ask another agent for help without needing to know its implementation details. The caller discovers the remote agent’s capabilities through an&amp;nbsp;&lt;SPAN data-streamdown="strong"&gt;agent card&lt;/SPAN&gt;, sends a task through the A2A protocol, and receives a response it can incorporate into the conversation.&lt;/P&gt;
&lt;P&gt;For Foundry-hosted A2A endpoints, discovery is authenticated. The agent-card URLs and protocol endpoint require Microsoft Entra ID authentication and are not publicly accessible.&lt;/P&gt;
&lt;P&gt;For example:&lt;/P&gt;
&lt;UL data-streamdown="unordered-list"&gt;
&lt;LI data-streamdown="list-item"&gt;A support agent can call a billing agent.&lt;/LI&gt;
&lt;LI data-streamdown="list-item"&gt;A research agent can call a data-analysis agent.&lt;/LI&gt;
&lt;LI data-streamdown="list-item"&gt;A Foundry-hosted enterprise agent can call a specialized agent running outside Foundry.&lt;/LI&gt;
&lt;LI data-streamdown="list-item"&gt;A Hosted Agent can call a Foundry Toolbox that wraps an A2A connection.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2 data-streamdown="heading-2"&gt;Architecture 1: External agent calls a Foundry agent through an A2A Endpoint&lt;/H2&gt;
&lt;P&gt;In this pattern, a Foundry agent is exposed as an A2A endpoint. An external agent discovers the agent card and invokes it using the A2A protocol.&lt;/P&gt;
&lt;img&gt;An external agent discovers a Foundry agent through its agent card, invokes it using A2A protocol v1.0, and receives a response generated with Foundry tools, data, and models.&lt;/img&gt;
&lt;P&gt;The Foundry agent processes the request using its configured models, instructions, tools, and enterprise data sources, and then returns the response to the calling agent.&lt;/P&gt;
&lt;P&gt;A Foundry prompt agent must support the Responses protocol before it can be exposed through an incoming A2A endpoint.&lt;/P&gt;
&lt;P&gt;The Foundry agent exposes an A2A base URL in the following format:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;https://{account}.services.ai.azure.com/api/projects/{project}/agents/{agent}/endpoint/protocols/a2a&lt;/LI-CODE&gt;
&lt;P&gt;The version-specific agent-card URLs follow these patterns:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;https://{account}.services.ai.azure.com/api/projects/{project}/agents/{agent}/endpoint/protocols/a2a/agentCard/v1.0
&lt;/LI-CODE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;LI-CODE lang=""&gt;https://{account}.services.ai.azure.com/api/projects/{project}/agents/{agent}/endpoint/protocols/a2a/agentCard/v0.3&lt;/LI-CODE&gt;
&lt;P&gt;For new integrations, target&amp;nbsp;&lt;SPAN data-streamdown="strong"&gt;A2A protocol v1.0&lt;/SPAN&gt;, which is GA. Foundry also supports v0.3 for existing preview integrations, but v1.0 is the recommended path forward.&lt;/P&gt;
&lt;P&gt;Foundry serves both protocol versions through the same A2A base path. The caller selects the version through agent-card negotiation, the A2A-Version header, or the a2a-version query parameter. Production clients should explicitly negotiate or request version 1.0 rather than relying on the default behavior.&lt;/P&gt;
&lt;P&gt;Foundry’s incoming A2A v1.0 endpoint uses JSON-RPC. Incoming A2A v0.3 supports JSON-RPC and HTTP+JSON, but v0.3 remains in preview. Only text modality is supported, and streaming responses are not supported.&lt;/P&gt;
&lt;H2 data-streamdown="heading-2"&gt;Architecture 2: Foundry agent calls another agent using the A2A Tool&lt;/H2&gt;
&lt;P&gt;The reverse pattern is just as important. A Foundry agent can use the&amp;nbsp;&lt;SPAN data-streamdown="strong"&gt;A2A Tool&lt;/SPAN&gt; to call another A2A-compatible endpoint.&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;A RemoteA2A project connection stores the remote A2A base URL and its authentication configuration. The A2A Tool references the connection and specifies the A2A protocol version to use.&lt;/P&gt;
&lt;P&gt;Foundry resolves the default agent-card path automatically and negotiates the A2A protocol version. You do not need to configure send_credentials_for_agent_card for a Foundry agent target.&lt;/P&gt;
&lt;P&gt;First, set the active Foundry project. Then create the RemoteA2A connection.&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;PROJECT_ENDPOINT="https://{account}.services.ai.azure.com/api/projects/{project}"

azd ai project set "$PROJECT_ENDPOINT"

azd ai connection create my-a2a-connection \
  --kind remote-a2a \
  --target "https://{account}.services.ai.azure.com/api/projects/{project}/agents/{agent}/endpoint/protocols/a2a" \
  --auth-type agentic-identity \
  --audience "https://ai.azure.com"&lt;/LI-CODE&gt;
&lt;P&gt;The target Foundry project or agent must grant the calling agent identity the Foundry Agent Consumer role, or another role that contains the required endpoint permissions. When assigning the role, use the calling identity’s Microsoft Entra object or principal ID, not its application or client ID.&lt;/P&gt;
&lt;P&gt;A role assignment can be created at either the target project scope or the individual agent scope:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;az role assignment create \
  --assignee-object-id "{calling-agent-principal-object-id}" \
  --assignee-principal-type "ServicePrincipal" \
  --role "eed3b665-ab3a-47b6-8f48-c9382fb1dad6" \
  --scope "{target-project-or-agent-resource-id}"&lt;/LI-CODE&gt;
&lt;P&gt;Agent identities and managed identities are represented as service principals in Microsoft Entra ID, which is why ServicePrincipal is used as the principal type.&lt;/P&gt;
&lt;H2 data-streamdown="heading-2"&gt;Architecture 3: Hosted Agent uses A2A through a Foundry Toolbox&lt;/H2&gt;
&lt;P&gt;Hosted Agents can use A2A by attaching a Foundry Toolbox. The toolbox contains an A2A Toolbox Tool and exposes it through an MCP-compatible endpoint.&lt;/P&gt;
&lt;P&gt;The complete interaction flow is:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;The Hosted Agent connects to the Foundry Toolbox over MCP.&lt;/LI&gt;
&lt;LI&gt;The Hosted Agent authenticates to the toolbox using Microsoft Entra ID.&lt;/LI&gt;
&lt;LI&gt;The toolbox invokes its configured A2A Toolbox Tool.&lt;/LI&gt;
&lt;LI&gt;The A2A Toolbox Tool references a RemoteA2A project connection.&lt;/LI&gt;
&lt;LI&gt;The RemoteA2A connection identifies the destination and downstream authentication configuration.&lt;/LI&gt;
&lt;LI&gt;The toolbox invokes the remote agent through the A2A protocol.&lt;/LI&gt;
&lt;LI&gt;The result is returned to the Hosted Agent through the toolbox’s MCP endpoint.&lt;/LI&gt;
&lt;/OL&gt;
&lt;img&gt;Hosted Agents securely invoke remote A2A agents through a Foundry Toolbox and RemoteA2A project connection.&lt;/img&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;This creates a two-hop architecture: Hosted Agent to Foundry Toolbox over MCP, followed by Foundry Toolbox to the remote agent over A2A.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;The toolbox MCP endpoint used by the Hosted Agent follows this format:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;https://{account}.services.ai.azure.com/api/projects/{project}/toolboxes/{toolbox-name}/versions/{version}/mcp?api-version=v1&lt;/LI-CODE&gt;
&lt;P&gt;&lt;EM&gt;A version-specific toolbox endpoint can also be used when the Hosted Agent must remain pinned to an immutable toolbox version.&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;The Hosted Agent authenticates to the toolbox using Microsoft Entra ID and the following scope:&lt;/P&gt;
&lt;LI-CODE lang=""&gt;https://ai.azure.com/.default&lt;/LI-CODE&gt;
&lt;P&gt;The Hosted Agent authenticates to the Foundry Toolbox MCP endpoint using Microsoft Entra ID. For direct token acquisition, use the https://ai.azure.com/.default scope.&lt;/P&gt;
&lt;P&gt;Downstream credentials, secrets, managed identity settings, and OAuth configuration should remain in the RemoteA2A project connection rather than being embedded in the Hosted Agent’s code.&lt;/P&gt;
&lt;H2 data-streamdown="heading-2"&gt;Authentication is the design decision&lt;/H2&gt;
&lt;P&gt;The most important architecture choice is authentication. A2A supports different patterns depending on whether the calling agent should act as itself or on behalf of the user.&lt;/P&gt;
&lt;P&gt;RemoteA2A connection authentication options&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 100%; height: 336.236px; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr style="height: 34.9006px;"&gt;&lt;th style="height: 34.9006px;"&gt;Auth type&lt;/th&gt;&lt;th style="height: 34.9006px;"&gt;Use when&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr style="height: 34.9006px;"&gt;&lt;td style="height: 34.9006px;"&gt;none&lt;/td&gt;&lt;td style="height: 34.9006px;"&gt;The remote endpoint does not require auth. Rare for production.&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 34.9006px;"&gt;&lt;td style="height: 34.9006px;"&gt;custom-keys&lt;/td&gt;&lt;td style="height: 34.9006px;"&gt;The endpoint expects an API key, PAT, bearer token, or custom header.&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 94.858px;"&gt;&lt;td style="height: 94.858px;"&gt;oauth2&lt;/td&gt;&lt;td style="height: 94.858px;"&gt;
&lt;P&gt;Each user authorizes access through an OAuth 2.0 consent flow. Use this when the downstream action must preserve the user’s individual permissions.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 34.9006px;"&gt;&lt;td style="height: 34.9006px;"&gt;user-entra-token&lt;/td&gt;&lt;td style="height: 34.9006px;"&gt;The user’s Microsoft Entra identity should flow to the remote service.&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 34.9006px;"&gt;&lt;td style="height: 34.9006px;"&gt;project-managed-identity&lt;/td&gt;&lt;td style="height: 34.9006px;"&gt;All agents in the project should share the project identity.&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 66.875px;"&gt;&lt;td style="height: 66.875px;"&gt;agentic-identity&lt;/td&gt;&lt;td style="height: 66.875px;"&gt;
&lt;P&gt;For service-to-service calls using an agent identity, service principal, or managed identity, use as the role-assignment principal type.&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;For&amp;nbsp;&lt;SPAN data-streamdown="strong"&gt;incoming A2A Endpoints on Foundry agents&lt;/SPAN&gt;, authentication is stricter:&lt;/P&gt;
&lt;UL data-streamdown="unordered-list"&gt;
&lt;LI data-streamdown="list-item"&gt;Incoming A2A requires&amp;nbsp;&lt;SPAN data-streamdown="strong"&gt;Microsoft Entra ID authentication&lt;/SPAN&gt;.&lt;/LI&gt;
&lt;LI data-streamdown="list-item"&gt;Key-based and unauthenticated incoming access are not supported.&lt;/LI&gt;
&lt;LI data-streamdown="list-item"&gt;The caller must have&amp;nbsp;&lt;SPAN data-streamdown="strong"&gt;Foundry Agent Consumer&lt;/SPAN&gt;&amp;nbsp;or another role that grants endpoint access.&lt;/LI&gt;
&lt;LI data-streamdown="list-item"&gt;Access can be granted at the project scope or individual agent scope.&lt;/LI&gt;
&lt;LI data-streamdown="list-item"&gt;Calls can use either&amp;nbsp;&lt;SPAN data-streamdown="strong"&gt;on-behalf-of user identity&lt;/SPAN&gt;&amp;nbsp;or a&amp;nbsp;&lt;SPAN data-streamdown="strong"&gt;service identity&lt;/SPAN&gt;&amp;nbsp;such as an agent identity, service principal, or managed identity.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Use OAuth or user identity passthrough when the remote action must respect each user’s permissions.&lt;/P&gt;
&lt;H2 data-streamdown="heading-2"&gt;What you configure&lt;/H2&gt;
&lt;P&gt;To expose a Foundry agent as an A2A Endpoint, you configure two things:&lt;/P&gt;
&lt;OL data-streamdown="ordered-list"&gt;
&lt;LI data-streamdown="list-item"&gt;An&amp;nbsp;&lt;SPAN data-streamdown="strong"&gt;agent card&lt;/SPAN&gt;, describing the agent’s capabilities.&lt;/LI&gt;
&lt;LI data-streamdown="list-item"&gt;The&amp;nbsp;&lt;SPAN data-streamdown="strong"&gt;A2A protocol&lt;/SPAN&gt;&amp;nbsp;on the agent endpoint.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;The relevant portion of the agent PATCH request body looks like this:&lt;/P&gt;
&lt;LI-CODE lang="json"&gt;{ "agent_card": { "description": "A specialist agent that answers questions about invoices.", "version": "1.0", "skills": [ { "id": "invoice-lookup", "name": "Invoice lookup", "description": "Finds and summarizes invoice status." } ] }, "agent_endpoint": { "protocol_configuration": { "responses": {}, "a2a": {} } } }&lt;/LI-CODE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Agent cards for prompt agents in Foundry can also be configured directly in the portal.&amp;nbsp; Then another agent can call it via Tools using the agents A2A endpoint.&lt;/P&gt;
&lt;img&gt;A2A Endpoint for an agent in Foundry&lt;/img&gt;&lt;img&gt;
&lt;P&gt;Agent card configuration for a Foundry agent&lt;/P&gt;
&lt;/img&gt;
&lt;H2 data-streamdown="heading-2"&gt;Final thoughts&lt;/H2&gt;
&lt;P&gt;A2A Endpoints and the A2A Tool make Foundry agents composable. Instead of building one large agent that knows everything, you can build focused agents that publish capabilities and collaborate securely.&lt;/P&gt;
&lt;P&gt;Use&amp;nbsp;&lt;SPAN data-streamdown="strong"&gt;A2A Endpoints&lt;/SPAN&gt;&amp;nbsp;when you want other agents to call your Foundry agent. Use the&amp;nbsp;&lt;SPAN data-streamdown="strong"&gt;A2A Tool&lt;/SPAN&gt;&amp;nbsp;when your Foundry agent needs to call another agent. For Hosted Agents, use a&amp;nbsp;&lt;SPAN data-streamdown="strong"&gt;Foundry Toolbox&lt;/SPAN&gt;&amp;nbsp;to bring A2A into the hosted runtime cleanly.&lt;/P&gt;
&lt;P&gt;Learn more:&lt;/P&gt;
&lt;UL data-streamdown="unordered-list"&gt;
&lt;LI data-streamdown="list-item"&gt;A2A authentication:&amp;nbsp;&lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/agent-to-agent-authentication" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/agent-to-agent-authentication&lt;/A&gt;&lt;/LI&gt;
&lt;LI data-streamdown="list-item"&gt;A2A Tool:&amp;nbsp;&lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/tools/agent-to-agent" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/tools/agent-to-agent&lt;/A&gt;&lt;/LI&gt;
&lt;LI data-streamdown="list-item"&gt;A2A Endpoint: &lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/enable-agent-to-agent-endpoint" target="_blank" rel="noopener"&gt;https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/enable-agent-to-agent-endpoint&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;</description>
      <pubDate>Fri, 18 Sep 2026 19:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/a2a-endpoints-and-a2a-tool-in-microsoft-foundry-agents/ba-p/4557115</guid>
      <dc:creator>Christian_Coello</dc:creator>
      <dc:date>2026-09-18T19:00:00Z</dc:date>
    </item>
    <item>
      <title>Announcing the Playwright Workspaces Remote MCP Server for Agentic Browser Automation</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/announcing-the-playwright-workspaces-remote-mcp-server-for/ba-p/4555698</link>
      <description>&lt;P&gt;An agent can analyze a request, choose the right next step, and call every available API—then stall when the final action exists only in a web interface.&lt;/P&gt;
&lt;P&gt;AI agents are increasingly capable of reasoning, planning, and working with APIs. But many real business processes still depend on websites: business portals, admin consoles, support tools, forms, dashboards, and applications without an API for every action. That last mile can stop an otherwise capable agent.&lt;/P&gt;
&lt;P&gt;The &lt;STRONG&gt;Playwright Workspaces remote Model Context Protocol (MCP) server&lt;/STRONG&gt;, now in preview closes this gap. It gives MCP-compatible agents a managed remote browser and a comprehensive set of browser automation tools optimized for AI agents—without requiring teams to install Playwright, package browser binaries, or operate browser infrastructure in the agent environment.&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-olk-copy-source="MessageBody"&gt;The remote MCP server uses Streamable HTTP—a transport protocol that allows real-time, bidirectional communication between your agent and the browser session—and is delivered as a&amp;nbsp; &lt;A href="https://learn.microsoft.com/en-us/azure/app-testing/playwright-workspaces/overview-what-is-microsoft-playwright-workspaces" target="_blank" rel="noopener"&gt;Playwright Workspaces&lt;/A&gt; feature. &lt;/SPAN&gt;Agents can create a browser session, navigate and inspect pages, interact with elements, upload files, review console and network activity, capture screenshots, and close the session when the workflow is complete. We are exposing definitive set of 22 tools which can help you complete browser flows successfully with agent interactions.&lt;BR /&gt;&lt;STRONG&gt;&lt;BR /&gt;&lt;/STRONG&gt;MCP config with GitHub Copilot CLI:&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;&lt;LI-CODE lang="python"&gt;MCPTool(
    server_label="playwright-browser-automation",
    server_url=mcp_server_url,
    headers={"x-api-key": playwright_service_access_token},
    require_approval="never",
)&lt;/LI-CODE&gt;
&lt;P&gt;&lt;BR /&gt;Configure Tool approval Foundry:&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Why this release matters&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Enterprise workflows rarely live in one API.&lt;/P&gt;
&lt;P&gt;The reality is:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Critical tasks often end in a web interface.&lt;/LI&gt;
&lt;LI&gt;Multi-step workflows require both reasoning and interaction.&lt;/LI&gt;
&lt;LI&gt;Website structure and timing can change between runs.&lt;/LI&gt;
&lt;LI&gt;Agent environments might not allow browser packages or local processes.&lt;/LI&gt;
&lt;LI&gt;Teams need visibility and control when automation reaches an unexpected state.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;The Playwright Workspaces remote MCP server closes this gap by exposing browser automation through an open protocol. Your agent uses standard MCP tools while Playwright Workspaces provides the remote browser infrastructure.&lt;/P&gt;
&lt;P&gt;This makes it easier to add browser execution to an existing agentic workflow and keep the browser lifecycle explicit, observable, and bounded.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;What's new&lt;/STRONG&gt;&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;&amp;nbsp;A hosted remote MCP endpoint : &lt;/STRONG&gt;Connect an MCP-compatible client to a workspace-scoped HTTPS endpoint &lt;SPAN data-olk-copy-source="MessageBody"&gt;(a URL unique to your Playwright Workspace that handles all MCP communication)&lt;/SPAN&gt;. The service is hosted and managed, so there is no local MCP server process to install or maintain.&lt;/LI&gt;
&lt;/OL&gt;
&lt;OL start="2"&gt;
&lt;LI&gt;&lt;STRONG&gt;&amp;nbsp;Browser tools designed for agentic workflows : &lt;/STRONG&gt;The preview exposes 22 tools across the browser lifecycle. They cover navigation, accessibility-based page inspection, interaction, tabs, file upload, screenshots, browser diagnostics. See the product documentation for the complete tool and parameter reference. Accessibility snapshots give agents a structured way to understand a page and select action targets. For focused tasks, browser_find can locate relevant snapshot nodes without returning the entire page representation.&lt;/LI&gt;
&lt;/OL&gt;
&lt;OL start="3"&gt;
&lt;LI&gt;&lt;STRONG&gt;&amp;nbsp;An observe-act-verify automation loop : &lt;/STRONG&gt;Reliable browser agents shouldn't execute a long sequence of clicks without checking the result. The tool set supports a practical loop which helps agents recover when a page rerenders, a selector becomes stale, a dialog opens, or an action completes differently than expected.&lt;BR /&gt;&lt;BR /&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Step&lt;/th&gt;&lt;th&gt;Action&lt;/th&gt;&lt;th&gt;Tool Example&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;1. Observe&lt;/td&gt;&lt;td&gt;Understand the current page state&lt;/td&gt;&lt;td&gt;browser_snapshot&amp;nbsp;or&amp;nbsp;browser_find&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;2. Act&lt;/td&gt;&lt;td&gt;Execute on a snapshot reference or selector&lt;/td&gt;&lt;td&gt;browser_click,&amp;nbsp;browser_type&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;3. Wait&lt;/td&gt;&lt;td&gt;Pause for expected page condition&lt;/td&gt;&lt;td&gt;browser_wait&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;4. Verify&lt;/td&gt;&lt;td&gt;Confirm resulting state before continuing&lt;/td&gt;&lt;td&gt;browser_snapshot&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;5. Close&lt;/td&gt;&lt;td&gt;End the browser session when complete&lt;/td&gt;&lt;td&gt;browser_close&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;OL start="4"&gt;
&lt;LI&gt;&lt;STRONG&gt;&amp;nbsp;Support for Microsoft Foundry, GitHub Copilot CLI and other agent builder surfaces : &lt;/STRONG&gt;&lt;SPAN data-olk-copy-source="MessageBody"&gt;You can connect the remote server to various agent builder surfaces and use cloud browsers for capabilities including session isolation and private website support. VNet injection lets agents access internal websites that aren't exposed to the public internet—useful for enterprise admin consoles and internal tools.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;OL start="5"&gt;
&lt;LI&gt;&lt;STRONG&gt;&amp;nbsp;Visibility using LiveView when you need investigation: &lt;/STRONG&gt;A created session can return a Live View URL when Live View is available. This gives developers and operators a way to observe browser activity while evaluating or troubleshooting an agentic flow. The server also exposes bounded console-message, network-request, snapshot, and screenshot results. These signals help identify navigation failures, page errors, timing issues, and unexpected application state.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;STRONG&gt;Choose the right Playwright experience&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Playwright supports several development and automation models. Choose the experience that matches your workload:&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;What you need&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;STRONG&gt;Recommended experience&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Give an MCP-compatible agent a hosted browser through a remote endpoint&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/app-testing/playwright-workspaces/how-to-playwright-workspaces-remote-mcp" target="_blank" rel="noopener"&gt;Playwright Workspaces Remote MCP Server&lt;/A&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Add browser automation to Microsoft Foundry through its integrated toolbox experience&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/tools/browser-automation?pivots=python" target="_blank" rel="noopener"&gt;Browser Automation Tool in Microsoft Foundry&lt;/A&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Bring your remote CDP browsers with preferred choice of framework&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;A class="lia-external-url" href="https://github.com/microsoft-foundry/foundry-samples/tree/main/samples/python/hosted-agents/agent-framework/responses/14-browser-automation-agent" target="_blank" rel="noopener"&gt;Hosted agent Sample in Foundry&lt;/A&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;
&lt;P&gt;Run an existing Playwright test suite at scale across managed cloud browsers&lt;/P&gt;
&lt;/td&gt;&lt;td&gt;
&lt;P&gt;&lt;A class="lia-external-url" href="https://learn.microsoft.com/en-us/azure/app-testing/playwright-workspaces/overview-run-playwright-tests-at-scale" target="_blank" rel="noopener"&gt;Playwright Workspaces test execution&lt;/A&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 50.00%" /&gt;&lt;col style="width: 50.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;The Remote MCP Server is designed for agent-driven, tool-based browser workflows. It complements rather than replaces the Playwright SDK, OSS playwright-cli, or the integrated Browser Automation Tool in Microsoft Foundry.&lt;BR /&gt;&lt;BR /&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Get Started&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG data-olk-copy-source="MessageBody"&gt;Build your endpoint URL&lt;/STRONG&gt;: Take your Playwright Workspaces API endpoint, replace .api with .mcp and append /mcp to create the remote MCP server URL. Example: &lt;A href="https://{region}.mcp.playwright.microsoft.com/playwrightworkspaces/{workspaceId}/mcp" target="_blank" rel="noopener" data-linkindex="4"&gt;&amp;nbsp;https://{region}.mcp.playwright.microsoft.com/playwrightworkspaces/{workspaceId}/mcp&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG data-olk-copy-source="MessageBody"&gt;Configure authentication&lt;/STRONG&gt;: Add the x-api-key header with your Playwright Workspaces access token. Example : set: x-api-key: &amp;lt;your-access-token&amp;gt;&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Connect &lt;/STRONG&gt;your agent to MCP&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;TIP :&lt;/STRONG&gt; &amp;nbsp;Exercise the Tools approval flow as present in Foundry, Copilot CLI– easier to get started&lt;BR /&gt;&lt;BR /&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;What this unlocks&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The Remote MCP Server enables agents to complete workflows across APIs and websites, combine automation with human approval, inspect dynamic applications, and collect page, console, network, and screenshot evidence. Instead of stopping when a process reaches a browser, an agent can take a controlled action, verify the result, and continue toward the outcome.&lt;BR /&gt;&lt;BR /&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG data-olk-copy-source="MessageBody"&gt;Next Steps&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Try it now&lt;/STRONG&gt;: Follow the&amp;nbsp;&lt;A href="https://nam06.safelinks.protection.outlook.com/?url=https%3A%2F%2Flearn.microsoft.com%2F...&amp;amp;data=05%7C02%7Cnandinim%40microsoft.com%7C9a3d3a858df04e2e2a1208df1315cdac%7C72f988bf86f141af91ab2d7cd011db47%7C1%7C0%7C639250656534444075%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&amp;amp;sdata=bwNrOuRLdLNjABANq6VnJHLAkwRODFvsvU2O%2B7fc8hE%3D&amp;amp;reserved=0" target="_blank" rel="noopener" data-auth="NotApplicable" data-linkindex="4"&gt;MCP Quickstart&lt;/A&gt;&amp;nbsp;to connect your first agent in under 10 minutes&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Explore the tools&lt;/STRONG&gt;: Review the&amp;nbsp;&lt;A href="https://nam06.safelinks.protection.outlook.com/?url=https%3A%2F%2Flearn.microsoft.com%2F...&amp;amp;data=05%7C02%7Cnandinim%40microsoft.com%7C9a3d3a858df04e2e2a1208df1315cdac%7C72f988bf86f141af91ab2d7cd011db47%7C1%7C0%7C639250656534456004%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&amp;amp;sdata=PQgv%2FC7t3Y8%2BHcYwYcVwiEO204vPnotmxPr1dCwCFOo%3D&amp;amp;reserved=0" target="_blank" rel="noopener" data-auth="NotApplicable" data-linkindex="5"&gt;complete tool reference&lt;/A&gt;&amp;nbsp;to see all 22 browser automation capabilities&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;See it in action&lt;/STRONG&gt;: Check out our&amp;nbsp;&lt;A class="lia-external-url" href="https://github.com/Azure/playwright-workspaces/tree/main/samples/remote-mcp" target="_blank" rel="noopener" data-auth="NotApplicable" data-linkindex="6"&gt;sample agents repository&lt;/A&gt;&amp;nbsp;for working examples with Foundry and GitHub Copilot CLI&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Get help&lt;/STRONG&gt;: Join the conversation on&amp;nbsp;&lt;A href="https://nam06.safelinks.protection.outlook.com/?url=https%3A%2F%2Fdiscordapp.com%2F...&amp;amp;data=05%7C02%7Cnandinim%40microsoft.com%7C9a3d3a858df04e2e2a1208df1315cdac%7C72f988bf86f141af91ab2d7cd011db47%7C1%7C0%7C639250656534480952%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&amp;amp;sdata=rr%2BWLBT0ykf1aBOlFr5lQTiHZEUZngWKtjlk3i7LoXU%3D&amp;amp;reserved=0" target="_blank" rel="noopener" data-auth="NotApplicable" data-linkindex="7"&gt;Discord&lt;/A&gt;&amp;nbsp;or submit feedback via our&amp;nbsp;&lt;A href="https://nam06.safelinks.protection.outlook.com/?url=https%3A%2F%2Faka.ms%2Fpww%2Ffeedback&amp;amp;data=05%7C02%7Cnandinim%40microsoft.com%7C9a3d3a858df04e2e2a1208df1315cdac%7C72f988bf86f141af91ab2d7cd011db47%7C1%7C0%7C639250656534493162%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&amp;amp;sdata=3Xum%2BfeixDfC3h0%2Btc8wm8dZ9n1H%2F%2BpudugyZgyLreY%3D&amp;amp;reserved=0" target="_blank" rel="noopener" data-auth="NotApplicable" data-linkindex="8"&gt;feedback channel&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;</description>
      <pubDate>Thu, 24 Sep 2026 14:41:55 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/announcing-the-playwright-workspaces-remote-mcp-server-for/ba-p/4555698</guid>
      <dc:creator>NandiniMuralidharan</dc:creator>
      <dc:date>2026-09-24T14:41:55Z</dc:date>
    </item>
    <item>
      <title>Practical Local AI for Enterprise Architects: A Workload Placement Approach</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/practical-local-ai-for-enterprise-architects-a-workload/ba-p/4556429</link>
      <description>&lt;P&gt;Enterprise AI discussions often begin with a use case and end with a deployment constraint. A clinical document assistant may have clear value, yet protected health information cannot leave an approved environment. A factory assistant may improve operations, yet the facility cannot depend on continuous connectivity. A financial analysis workload may require dedicated execution, predictable operating costs, or strict data residency.&lt;/P&gt;
&lt;P&gt;These conditions do not invalidate the business case. They define the workload execution location and operating pattern.&lt;/P&gt;
&lt;P&gt;For enterprise architects, the central question is no longer whether AI belongs in the cloud or on local infrastructure. The more useful question is this: which execution location best satisfies the business, data, security, performance, and operational requirements of each workload?&lt;/P&gt;
&lt;P&gt;This article presents a practical workload placement approach for local and contained AI. It also clarifies how Windows ML, Foundry Local, Microsoft Foundry, Azure Machine Learning, Azure Local, and GPU-enabled Azure Kubernetes Service can contribute to one governed architecture.&lt;/P&gt;
&lt;H2&gt;Architecture constraints are placement requirements&lt;/H2&gt;
&lt;P&gt;Many AI projects stall between prototype and production. The model may perform well, the user experience may be useful, and the business sponsor may support the investment. The project still stops because the production environment introduces requirements that were absent during experimentation.&lt;/P&gt;
&lt;P&gt;Common constraints include:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Regulation: Sensitive information must remain within an approved environment.&lt;/LI&gt;
&lt;LI&gt;Data residency: Inference must occur inside a defined geographic, organizational, or technical boundary.&lt;/LI&gt;
&lt;LI&gt;Security: The workload requires dedicated, private, or isolated execution.&lt;/LI&gt;
&lt;LI&gt;Connectivity: The solution must continue to operate when network access is limited or unavailable.&lt;/LI&gt;
&lt;LI&gt;Performance: The experience depends on consistent, low-latency inference.&lt;/LI&gt;
&lt;LI&gt;Cost: The operating model must be understandable, measurable, and predictable.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Architects should treat these statements as architecture inputs. A constraint can change the execution plane, the infrastructure choice, the governance pattern, or the operational model. It should not automatically eliminate the use case.&lt;/P&gt;
&lt;P&gt;This distinction changes the conversation. Instead of asking whether the organization can use AI under restrictive conditions, the architecture team can ask which approved deployment pattern satisfies those conditions.&lt;/P&gt;
&lt;H3&gt;Think in terms of an AI deployment spectrum&lt;/H3&gt;
&lt;P&gt;Enterprise AI can execute across a spectrum of locations:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Device: Laptops and workstations can run local inference close to the user and the application.&lt;/LI&gt;
&lt;LI&gt;Edge: Factory systems, field equipment, and other edge environments can support local decisions when latency or connectivity matters.&lt;/LI&gt;
&lt;LI&gt;Private infrastructure: Azure-managed or customer-controlled infrastructure can provide dedicated capacity, isolated networking, and approved data boundaries.&lt;/LI&gt;
&lt;LI&gt;Cloud: Managed models and elastic services can provide broad capability, rapid innovation, and global scale.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;These locations are not competing ideologies. They are architecture options. A single enterprise may use all four, from workstation testing and disconnected facilities to governed private infrastructure and elastic cloud services. The appropriate location depends on the workload, not on a universal preference for local or cloud execution.&lt;/P&gt;
&lt;P&gt;Local AI expands placement choice. It does not replace cloud AI.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;EM&gt;Figure 1&lt;/EM&gt;&lt;/STRONG&gt;&lt;EM&gt;. AI can run across device, edge, private infrastructure, and cloud environments. The workload requirements determine the placement.&lt;/EM&gt;&lt;/P&gt;
&lt;H2&gt;Clarify the Microsoft AI portfolio&lt;/H2&gt;
&lt;P&gt;Several Microsoft technologies support this spectrum, but they operate at different architectural layers. Clear role definitions help architecture teams avoid treating related products as interchangeable.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;EM&gt;Windows ML: the local execution engine&lt;/EM&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Windows ML supports efficient AI execution on Windows devices. It provides the execution capability that applications use to execute models on available device hardware. It is not the enterprise governance platform, the private infrastructure layer, or the large-scale GPU orchestration layer.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;EM&gt;Foundry Local: local model delivery and management&lt;/EM&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Foundry Local helps developers discover, manage, and execute models locally. It supports the path from local model selection to application integration and prompt validation. Windows ML executes models, while Foundry Local delivers and manages the local model experience.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;EM&gt;Microsoft Foundry: the enterprise control center&lt;/EM&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Microsoft Foundry supports the enterprise lifecycle for AI applications and agents. Platform teams can use it to build, evaluate, govern, and operate AI assets across a controlled environment. This is where local experimentation can become an enterprise asset.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;EM&gt;Azure Machine Learning: custom model lifecycle and MLOps&lt;/EM&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Azure Machine Learning supports training, deployment, and management for machine learning models. It is appropriate when the organization requires custom MLOps, managed experimentation, model registration, deployment controls, and lifecycle management.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;EM&gt;Azure Local: private, Azure-managed infrastructure&lt;/EM&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Azure Local provides Azure-managed infrastructure within an on-premises or approved private environment. Foundry Local focuses on models and local development. Azure Local focuses on private infrastructure.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;EM&gt;GPU-enabled AKS: scalable execution&lt;/EM&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;GPU-enabled Azure Kubernetes Service provides a Kubernetes execution layer for training and inference at scale. Platform engineering teams can use it when a workload requires higher throughput, clustered GPU capacity, or large-scale inference.&lt;/P&gt;
&lt;P&gt;The roles are complementary. Windows ML executes. Foundry Local supports local model delivery. Microsoft Foundry governs the enterprise lifecycle. Azure Machine Learning manages custom model operations. Azure Local provides private infrastructure. GPU-enabled AKS supports scalable execution.&lt;/P&gt;
&lt;P&gt;Match the workload to the execution pattern&lt;/P&gt;
&lt;P&gt;The following framework provides a starting point for architecture discussions.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table class="lia-background-color-16 lia-border-style-solid" border="1" style="width: 57.8704%; height: 413px; border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr style="height: 39px;"&gt;&lt;td class="lia-background-color-17" style="height: 39px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Workload requirement&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-background-color-17" style="height: 39px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Recommended pattern&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 39px;"&gt;&lt;td class="lia-background-color-22" style="height: 39px;"&gt;
&lt;P&gt;Fast experimentation&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-background-color-22" style="height: 39px;"&gt;
&lt;P&gt;Foundry Local&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 67px;"&gt;&lt;td class="lia-background-color-16" style="height: 67px;"&gt;
&lt;P&gt;Lowest latency on Windows or low transaction volume&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-background-color-16" style="height: 67px;"&gt;
&lt;P&gt;Windows ML with Foundry Local&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 67px;"&gt;&lt;td class="lia-background-color-22" style="height: 67px;"&gt;
&lt;P&gt;Disconnected, sovereign, or private operation&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-background-color-22" style="height: 67px;"&gt;
&lt;P&gt;Foundry Local with Azure Local&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 67px;"&gt;&lt;td style="height: 67px;"&gt;
&lt;P&gt;Enterprise agents and centralized governance&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 67px;"&gt;
&lt;P&gt;Microsoft Foundry&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 67px;"&gt;&lt;td class="lia-background-color-22" style="height: 67px;"&gt;
&lt;P&gt;Custom MLOps and model lifecycle management&lt;/P&gt;
&lt;/td&gt;&lt;td class="lia-background-color-22" style="height: 67px;"&gt;
&lt;P&gt;Azure Machine Learning&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 67px;"&gt;&lt;td style="height: 67px;"&gt;
&lt;P&gt;Large-scale inference or high transaction volume&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 67px;"&gt;
&lt;P&gt;GPU-enabled AKS&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;This framework does not replace a workload assessment. It provides the assessment with a structured focus.&lt;/P&gt;
&lt;P&gt;Architecture teams should evaluate data sensitivity, latency, transaction volume, connectivity, availability, identity, observability, model lifecycle, and operational ownership. They should also decide whether the workload must continue during a network interruption and whether infrastructure must remain within a defined boundary.&lt;/P&gt;
&lt;P&gt;The result may be a hybrid pattern. Local inference can serve a device or facility, while centralized services manage identity, policy, evaluation, telemetry, model registration, and approved updates.&lt;/P&gt;
&lt;H1&gt;Move from a local prototype to a governed asset&lt;/H1&gt;
&lt;P&gt;Local development can shorten the path to a useful proof of concept. A developer can select a model, run inference, test prompts, refine the user experience, and evaluate output without first building a complete enterprise platform.&lt;/P&gt;
&lt;P&gt;Production introduces a different set of operational responsibilities. The organization must package the model and application, register approved assets, define identity and access controls, capture telemetry, manage versions, perform evaluations, monitor health, and establish operational ownership.&lt;/P&gt;
&lt;P&gt;The workload does not necessarily need to change during this transition. The operational model does.&lt;/P&gt;
&lt;P&gt;Consider an assistant that processes a sensitive document. A small team may begin with local inference on a workstation because the source material must remain local. If the assistant expands from five users to five thousand, the user scenario can remain the same. What changes is the execution plane and the platform around it.&lt;/P&gt;
&lt;P&gt;The enterprise version may require Microsoft Foundry, an AI gateway, API management, Microsoft Entra ID, centralized observability, approved telemetry, and GPU-enabled AKS. The architecture adds shared governance, observability, scalability, and operational controls while preserving the original workload requirement.&lt;/P&gt;
&lt;P&gt;This is the governed handoff. The prototype becomes a repeatable enterprise asset without losing the placement characteristics that made the use case viable.&lt;/P&gt;
&lt;H3&gt;Use a contained AI reference architecture&lt;/H3&gt;
&lt;P&gt;A contained AI platform can be described through four connected layers.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Experience&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Applications, copilots, and agents provide the user experience. This layer should remain as independent as practical from the selected execution plane so the organization can change placement without redesigning the entire application.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Local AI orchestration&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Windows ML and Foundry Local support execution on devices and at the edge. This layer is useful when processing must remain close to the user, source system, or facility.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Platform and governance&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Microsoft Foundry, an AI gateway, API management, and Microsoft Entra ID provide shared controls. This layer can centralize identity, policy, safety practices, routing, evaluation, and lifecycle governance.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Execution and data&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;Azure Machine Learning, GPU-enabled AKS, Azure Local, and approved data services provide the execution capacity and data access pattern. The selected combination depends on scale, isolation, lifecycle, and infrastructure requirements.&lt;/P&gt;
&lt;P&gt;Observability spans every layer. Identity, policy, safety, telemetry, health, logging, and evaluations should not be added after deployment. They are part of the production architecture.&lt;/P&gt;
&lt;P&gt;The guiding principle is simple: use one governed pattern, then vary the execution location according to the workload.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;EM&gt;Figure 2&lt;/EM&gt;&lt;/STRONG&gt;&lt;EM&gt;. A shared governance layer can apply identity, policy, safety, evaluation, and observability controls across several execution environments.&lt;/EM&gt;&lt;/P&gt;
&lt;H2&gt;Treat economics as a workload-specific decision&lt;/H2&gt;
&lt;P&gt;Local AI can support a more predictable operating model when dedicated capacity is appropriate. Cloud services can provide elasticity, managed capabilities, and broad model choice. Neither option produces universal savings.&lt;/P&gt;
&lt;P&gt;Economic analysis should reflect the customer workload. Architects should compare utilization, capacity, transaction volume, hardware lifecycle, software operations, support, networking, availability, and governance overhead. A dedicated environment with low utilization may be inefficient. A stable, high-utilization workload may benefit from reserved or dedicated capacity. A variable workload may benefit from consumption-based cloud services.&lt;/P&gt;
&lt;P&gt;The right economic question is not whether local AI is cheaper. It is whether the selected placement produces an explainable operating model for the workload.&lt;/P&gt;
&lt;H1&gt;Revisit use cases that were previously blocked&lt;/H1&gt;
&lt;P&gt;Local and contained execution patterns can reopen scenarios that were difficult to approve under a cloud-only assumption. Examples include clinical document extraction, factory-floor assistance, disconnected government operations, air-gapped intelligence workloads, offline field guidance, and restricted-data analysis.&lt;/P&gt;
&lt;P&gt;The business value in these scenarios may have existed for years. The missing element was an acceptable deployment model.&lt;/P&gt;
&lt;P&gt;Architects can now separate the use case from the execution constraint. They can preserve the required boundary while still applying enterprise identity, governance, observability, and lifecycle practices.&lt;/P&gt;
&lt;H3&gt;A practical next step&lt;/H3&gt;
&lt;P&gt;Start with one blocked or constrained workload.&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Select a use case with clear business value.&lt;/LI&gt;
&lt;LI&gt;Define the required data, network, security, regulatory, and operational boundary.&lt;/LI&gt;
&lt;LI&gt;Identify the appropriate execution location across device, edge, private infrastructure, and cloud.&lt;/LI&gt;
&lt;LI&gt;Map the workload to the Microsoft technologies that serve each architecture layer.&lt;/LI&gt;
&lt;LI&gt;Define the governed handoff from prototype to production.&lt;/LI&gt;
&lt;LI&gt;Validate the operating model, including identity, evaluation, telemetry, support, and cost.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;This process turns an abstract local-versus-cloud debate into a concrete architecture decision.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;EM&gt;Figure 3&lt;/EM&gt;&lt;/STRONG&gt;&lt;EM&gt;. The AI workload can remain consistent as the operating model matures from an individual prototype to a governed enterprise platform.&lt;/EM&gt;&lt;/P&gt;
&lt;H2&gt;Conclusion&lt;/H2&gt;
&lt;P&gt;The future of enterprise AI is workload placement.&lt;/P&gt;
&lt;P&gt;Organizations need governed AI systems that can run wherever the business, data, and risk model require. Some workloads belong in elastic cloud services. Others require local inference, private infrastructure, isolated execution, or disconnected operation. Many will combine these patterns.&lt;/P&gt;
&lt;P&gt;Enterprise architects can support that diversity without creating separate governance models for every location. Build locally when the workload requires it. Govern centrally across the lifecycle. Scale through the execution plane that fits the demand.&lt;/P&gt;
&lt;P&gt;The most useful question is not, “Should this AI run locally or in the cloud?” It is, “What placement allows this workload to operate securely, reliably, and at enterprise scale?”&lt;/P&gt;</description>
      <pubDate>Wed, 16 Sep 2026 14:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/practical-local-ai-for-enterprise-architects-a-workload/ba-p/4556429</guid>
      <dc:creator>arronhoffer</dc:creator>
      <dc:date>2026-09-16T14:00:00Z</dc:date>
    </item>
    <item>
      <title>Your Agent Found the Right Schema. Then Ignored It.</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/your-agent-found-the-right-schema-then-ignored-it/ba-p/4553775</link>
      <description>&lt;H3 data-line="113"&gt;The assumption nobody tests&lt;/H3&gt;
&lt;P data-line="115"&gt;Most agent evaluations quietly assume something that never happens in production: that the model is handed exactly the right information, and nothing else.&lt;/P&gt;
&lt;P data-line="119"&gt;We changed that one assumption — and watched a model drop from&amp;nbsp;&lt;STRONG&gt;58% to 18%&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P data-line="121"&gt;Retrieval was not the problem. The correct information was sitting in the prompt. The model just didn't use it.&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P data-line="119"&gt;&lt;STRONG&gt;TL;DR&lt;/STRONG&gt; — Standard agent evaluations hand the model exactly the right schema. Real retrieval does not. Changing only that cost 40 points of accuracy, and neither a larger model nor better prompting recovered it. A $150 fine-tuning job won two-thirds of it back and beat prompted&amp;nbsp;gpt-5&amp;nbsp;by 28 points at a twenty-fifth of the per-turn cost. Steps 2–4 of the runbook below reproduce the gap on your own registry and need no fine-tuning.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P data-line="131"&gt;We measured this on schemas, but if you build agents that retrieve&amp;nbsp;&lt;EM&gt;tools&lt;/EM&gt;,&amp;nbsp;&lt;EM&gt;functions&lt;/EM&gt;, or&amp;nbsp;&lt;EM&gt;MCP servers&lt;/EM&gt;&amp;nbsp;into a prompt, the shape of the problem is the same. More on that at the end.&lt;/P&gt;
&lt;P data-line="135"&gt;Here is the setup. In schema-guided dialogue agents, the model is given a&amp;nbsp;&lt;EM&gt;schema&lt;/EM&gt;&amp;nbsp;— the list of fields it is allowed to fill — and must return a filled state. Benchmarks put the one correct schema in the prompt and report strong numbers: 86.4% and 74.7% on SGD [2, 3, 4]. But a deployed assistant has no idea which schema is correct. It retrieves candidates from a registry, so the prompt ends up holding the right schema&amp;nbsp;&lt;EM&gt;plus&lt;/EM&gt;&amp;nbsp;whatever else the retriever dragged in.&lt;/P&gt;
&lt;P data-line="142"&gt;So we changed exactly one thing:&amp;nbsp;&lt;STRONG&gt;how the schema reaches the prompt.&lt;/STRONG&gt;&amp;nbsp;Same dataset, same metric, same model. We call the realistic condition&amp;nbsp;&lt;STRONG&gt;RetDist&lt;/STRONG&gt;&amp;nbsp;(&lt;EM&gt;retrieval + distractors&lt;/EM&gt;) — BM25 top-1 plus two random distractor schemas from the same registry, shuffled. Note what this does&amp;nbsp;&lt;EM&gt;not&lt;/EM&gt;&amp;nbsp;guarantee: BM25 ranks the correct schema first on only 65.4% of turns, so on about a third of them the right answer is not in the prompt at all. That matters later.&lt;/P&gt;
&lt;P data-line="149"&gt;The score we measure is joint goal accuracy (JGA): the fraction of turns where&amp;nbsp;&lt;EM&gt;every&lt;/EM&gt;&amp;nbsp;field is exactly right. It is unforgiving on purpose — one wrong field fails the turn, which is also how a user experiences it.&lt;/P&gt;
&lt;P data-line="153"&gt;So if a turn needs three fields and the model gets two of them right, that turn scores zero. This matters for reading the numbers below: 58 does&amp;nbsp;&lt;STRONG&gt;not&lt;/STRONG&gt;&amp;nbsp;mean "58% of fields correct" — it means 58% of turns were&amp;nbsp;&lt;EM&gt;completely&lt;/EM&gt;&amp;nbsp;correct. Slot-level accuracy on the same runs is far higher, and far less honest about what a user would experience.&lt;/P&gt;
&lt;P data-line="159"&gt;&lt;STRONG&gt;58.0% with the gold schema. 18.3% under RetDist.&lt;/STRONG&gt; Same model —&amp;nbsp;gpt-4.1-mini, same weights, same dialogues, same metric. Only the prompt assembly changed.&lt;/P&gt;
&lt;img&gt;&lt;STRONG&gt;Figure 1. The only thing that changes is how the schema reaches the prompt.&lt;/STRONG&gt;&lt;EM&gt;Standard evaluation (A) hands gpt-4.1-mini the one correct schema, so the task reduces to filling in fields. A deployed assistant (B) retrieves candidates from the 21-service test registry, so the correct schema arrives alongside confusable ones. Same dialogue, same weights, same metric — 58.0 JGA becomes 18.3.&lt;/EM&gt;&lt;/img&gt;
&lt;H3 data-line="170"&gt;It is not a retrieval problem&lt;/H3&gt;
&lt;P data-line="172"&gt;The obvious explanation is that retrieval simply missed. It didn't.&lt;/P&gt;
&lt;P data-line="174"&gt;We split every turn by whether the correct schema was actually in the prompt. On the turns where it&amp;nbsp;&lt;STRONG&gt;was present&lt;/STRONG&gt;&amp;nbsp;— retrieval did its job —&amp;nbsp;gpt-4.1-mini&amp;nbsp;still only reached&amp;nbsp;&lt;STRONG&gt;24.0%&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P data-line="174"&gt;&amp;nbsp;&lt;/P&gt;
&lt;img&gt;&lt;STRONG&gt;Figure 2. Retrieval is not the bottleneck.&lt;/STRONG&gt;&lt;EM&gt;Restricted to the 27,085 turns where the correct schema was verifiably present in the prompt, the base model still reaches only 24.0 JGA. Fine-tuning on the same distribution lifts the same stratum to 62.7.&lt;/EM&gt;&lt;/img&gt;
&lt;P data-line="185"&gt;So the model isn't failing to &lt;EM&gt;find&lt;/EM&gt;&amp;nbsp;the schema. It's failing to&amp;nbsp;&lt;EM&gt;follow&lt;/EM&gt;&amp;nbsp;it. It falls back on field names memorised during pre-training instead of reading the schema in front of it. We call this&amp;nbsp;&lt;STRONG&gt;Schema-Following Collapse&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P data-line="189"&gt;This is why better retrieval would not have saved us. Prior robustness work on this benchmark varied the schema's&amp;nbsp;&lt;EM&gt;wording&lt;/EM&gt; [5]; the failure here survives even perfect retrieval.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H3 data-line="193"&gt;Bigger models don't fix it&lt;/H3&gt;
&lt;img&gt;&lt;STRONG&gt;Figure 3. Scale and prompt engineering do not close the gap.&lt;/STRONG&gt;&lt;EM&gt;Every prompted configuration — Llama-3.3-70B, gpt-5, and gpt-5 again with few-shot examples and strict decoding — lands inside the same 15–19 JGA band, alongside the un-tuned gpt-4.1-mini. Only fine-tuning leaves it. All bars are SGD-200 (n=1,810) under RetDist; the dashed line is the gold-schema ceiling.&lt;/EM&gt;&lt;/img&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P data-line="203"&gt;Prompt engineering moved&amp;nbsp;gpt-5&amp;nbsp;by 2.8 points. Everything stayed in a 15–19% band. This is not a capability problem you can scale or prompt your way out of.&lt;/P&gt;
&lt;P data-line="206"&gt;&amp;nbsp;&lt;/P&gt;
&lt;BLOCKQUOTE&gt;
&lt;P data-line="212"&gt;&lt;STRONG&gt;Two caveats on that chart, because it names names.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P data-line="214"&gt;&lt;STRONG&gt;This is a prompting stress test, not a capability ranking.&lt;/STRONG&gt;&amp;nbsp;All three models see one fixed deployment prompt. Give&amp;nbsp;gpt-5&amp;nbsp;a different prompt, a schema-selection tool, or permission to ask a clarifying question and the number changes. The finding is that no prompted configuration escaped the band — not that any particular model is weak.&amp;nbsp;gpt-5&amp;nbsp;scoring 16.6 here says something about the task, not about&amp;nbsp;gpt-5.&lt;/P&gt;
&lt;P data-line="221"&gt;&lt;STRONG&gt;And no, we did not fine-tune them.&lt;/STRONG&gt; At the time of these runs&amp;nbsp;gpt-5&amp;nbsp;exposed no supervised fine-tuning path on Foundry — it is reasoning-mode and fixed at temperature 1.0 — and fine-tuning a 70B model costs orders of magnitude more than the $150 that makes this recipe worth writing about. So this compares a fine-tuned small model against prompted large ones. That is the choice a team on a budget actually faces, but it is not a claim about what those models could do if they were trained.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;H3 data-line="229"&gt;What did work: train on the noise&lt;/H3&gt;
&lt;P data-line="231"&gt;&lt;STRONG&gt;RAFT (Retrieval-Augmented Fine-Tuning)&lt;/STRONG&gt;&amp;nbsp;[1], originally proposed for document question answering, trains a model with the evidence it should use alongside realistic distractors it should ignore. Unlike RAG, which adds retrieval at inference time, RAFT changes the&amp;nbsp;&lt;EM&gt;training&lt;/EM&gt;&amp;nbsp;distribution, so the model learns how to behave after retrieval has already happened — including how to behave when retrieval brought back the wrong thing.&lt;/P&gt;
&lt;P data-line="238"&gt;We adapt that idea from retrieved&amp;nbsp;&lt;EM&gt;documents&lt;/EM&gt;&amp;nbsp;to retrieved&amp;nbsp;&lt;EM&gt;schemas&lt;/EM&gt;. We call the recipe&amp;nbsp;&lt;STRONG&gt;SchemaRAFT&lt;/STRONG&gt;: build the fine-tuning examples to look like the deployment prompt, including both the correct schema and confusable alternatives. This is a task-specific application of RAFT, not a new training algorithm.&lt;/P&gt;
&lt;img&gt;&lt;STRONG&gt;Figure 4. The recipe.&lt;/STRONG&gt;&lt;EM&gt;Every training prompt is assembled the way the serving prompt is assembled — the correct schema plus two confusable distractors. One in five examples contains no correct schema at all and is labelled with an empty state, which is what teaches the model to abstain rather than guess.&lt;/EM&gt;&lt;/img&gt;
&lt;UL data-line="252"&gt;
&lt;LI data-line="252"&gt;19,300 chat-format examples (~2,100 prompt tokens, ~40 completion tokens each)&lt;/LI&gt;
&lt;LI data-line="253"&gt;each example: 1 gold schema + 2 random distractors, exactly as at serving time&lt;/LI&gt;
&lt;LI data-line="254"&gt;&lt;STRONG&gt;20% of examples omit the gold schema entirely and are labelled&amp;nbsp;{}&lt;/STRONG&gt;&amp;nbsp;— this teaches the model to withhold output rather than invent slots&lt;/LI&gt;
&lt;LI data-line="256"&gt;3 epochs, ~124M training tokens, one fine-tuning job on Microsoft Foundry&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="258"&gt;Here is one training example, trimmed. The system message carries the schemas, the user message carries the dialogue, and the assistant message is the only part the model is scored on.&lt;/P&gt;
&lt;P data-line="262"&gt;&lt;STRONG&gt;system&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang=""&gt;You are a dialogue state tracker for task-oriented conversations.
...
Rules:
- Output ONLY a valid JSON object on a single line.
- Use only slot names that appear in the schemas below.
- If no relevant slots have been mentioned, output: {}

Service schemas (one is relevant; others may be distractors):
Service: Hotels_4
  Accommodation searching and booking portal
  Slots:
    - location [free-text]: City or town where the hotel is located
    - number_of_rooms [categorical] (values: 1, 2, 3): Rooms to reserve
--------------------------------------------------
Service: Restaurants_2
  Restaurant search and reservation service
  Slots:
    - city [free-text]: City where the restaurant is located
    - cuisine [free-text]: Type of food served
    - party_size [categorical] (values: 1, 2, 3, 4): Number of seats
    ...
--------------------------------------------------
Service: Buses_3
  ...&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;user&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang=""&gt;Conversation history:
User: I need a table for two in San Francisco.
System: What kind of food?
User: Italian, somewhere moderate. Around 7 pm tomorrow.

Extract the current belief state for service: Restaurants_2&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;assistant&lt;/STRONG&gt;&amp;nbsp;— this is the label, and the only tokens in the loss&lt;/P&gt;
&lt;LI-CODE lang="json"&gt;{"city": "san francisco", "cuisine": "italian", "party_size": "2",
 "price_range": "moderate", "time": "19:00"}&lt;/LI-CODE&gt;
&lt;P data-line="309"&gt;Two things are worth noticing.&lt;/P&gt;
&lt;P data-line="311"&gt;&lt;STRONG&gt;The model is told which service it needs.&lt;/STRONG&gt;&amp;nbsp;The name sits at the end of the user message and matches one of the three schema headers verbatim.&amp;nbsp;&lt;EM&gt;Choosing&lt;/EM&gt;&amp;nbsp;the right schema is string matching. The only real task is to then stay inside it — and that is what collapses. A base model reaches for&amp;nbsp;location&amp;nbsp;from&amp;nbsp;Hotels_4&amp;nbsp;because "San Francisco" and hotels co-occur constantly in pre-training, even though&amp;nbsp;Restaurants_2&amp;nbsp;is named right there and has a&amp;nbsp;city&amp;nbsp;field. The gap is not a search failure. It is a grounding failure, and we made the search step as easy as it could possibly be.&lt;/P&gt;
&lt;P data-line="320"&gt;&lt;STRONG&gt;Nothing in the loss punishes using a distractor's field.&lt;/STRONG&gt;&amp;nbsp;There is no auxiliary objective, no reranker, no constrained decoding. It is ordinary next-token prediction on the assistant message. The model sees ~19,300 examples in which three schemas are present and the answer only ever draws on one of them, and that is the entire training signal.&lt;/P&gt;
&lt;P data-line="326"&gt;The inference prompt is the same object; only the assembly changes. At serving time BM25 chooses the schemas instead of the data builder, so the correct one is present 65.4% of the time rather than 80%, and the assistant message is generated instead of given.&lt;/P&gt;
&lt;P data-line="331"&gt;Result on&amp;nbsp;gpt-4.1-mini&amp;nbsp;— the same model from Figure 1, now fine-tuned:&amp;nbsp;&lt;STRONG&gt;18.3 → 45.0 JGA&lt;/STRONG&gt;, recovering about two-thirds of the gap to the gold-schema ceiling. On gold-in-prompt turns it goes 24.0 → 62.7.&lt;/P&gt;
&lt;P data-line="335"&gt;We re-ran the identical recipe on&amp;nbsp;gpt-4o-mini&amp;nbsp;with no code changes and got +28.0 points (14.5 → 42.5). It's a recipe, not a lucky run.&lt;/P&gt;
&lt;H3 data-line="338"&gt;The economics&lt;/H3&gt;
&lt;P data-line="340"&gt;Measured on the same deployment, per turn:&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 67.5%; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th class="lia-align-center"&gt;&amp;nbsp;&lt;/th&gt;&lt;th class="lia-align-center"&gt;Latency&lt;/th&gt;&lt;th class="lia-align-center"&gt;Relative cost&lt;/th&gt;&lt;th class="lia-align-center"&gt;JGA&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class="lia-align-center"&gt;gpt-5, prompted&lt;/td&gt;&lt;td class="lia-align-center"&gt;4,401 ms&lt;/td&gt;&lt;td class="lia-align-center"&gt;24.4×&lt;/td&gt;&lt;td class="lia-align-center"&gt;16.6&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td class="lia-align-center"&gt;&lt;STRONG&gt;gpt-4.1-mini&amp;nbsp;+ SchemaRAFT&lt;/STRONG&gt;&lt;/td&gt;&lt;td class="lia-align-center"&gt;&lt;STRONG&gt;1,095 ms&lt;/STRONG&gt;&lt;/td&gt;&lt;td class="lia-align-center"&gt;&lt;STRONG&gt;1.0×&lt;/STRONG&gt;&lt;/td&gt;&lt;td class="lia-align-center"&gt;&lt;STRONG&gt;45.0&lt;/STRONG&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 34.8805%" /&gt;&lt;col style="width: 19.3603%" /&gt;&lt;col style="width: 21.8372%" /&gt;&lt;col style="width: 23.8945%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P&gt;Four times faster, roughly 25× cheaper per turn, and 2.7× the accuracy. The one-time training job costs about $150 and pays for itself after roughly 18,000 served turns.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H3 data-line="351"&gt;Where else this shows up&lt;/H3&gt;
&lt;P data-line="353"&gt;We measured this on dialogue state tracking, but nothing in the failure is specific to dialogue. The ingredients are generic:&amp;nbsp;&lt;STRONG&gt;several semantically similar capability descriptions retrieved into one prompt, and a model asked to commit to one of them.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P data-line="358"&gt;It is worth being precise about how close the analogy is. An SGD service schema is a name, a description, and a list of typed parameters with their own descriptions. So is an OpenAI tool definition. So is an MCP tool listing. These are the same object wearing different labels, and the model is doing the same thing with all of them: reading several near-identical interface descriptions and picking one.&lt;/P&gt;
&lt;P data-line="365"&gt;That pattern is everywhere in agent systems today — function calling with a large tool catalogue, MCP server selection, tool and plugin routing, multi-agent handoff, enterprise API registries.&lt;/P&gt;
&lt;P data-line="369"&gt;What we did&amp;nbsp;&lt;EM&gt;not&lt;/EM&gt;&amp;nbsp;test is whether the collapse reproduces when the decoding target is a tool call rather than a belief state. So treat the extension as a hypothesis, not a result. The cheap move is to measure it on your own registry — which is exactly what the runbook below does, and it needs no fine-tuning.&lt;/P&gt;
&lt;H3 data-line="374"&gt;When this won't apply to you&lt;/H3&gt;
&lt;P data-line="376"&gt;This gap needs candidates that are genuinely&amp;nbsp;&lt;STRONG&gt;confusable&lt;/STRONG&gt;. SGD defines 45 services across 20 domains with heavily overlapping field names — the slot&amp;nbsp;city&amp;nbsp;alone appears in 9 of them, and 15 services carry some variant of it (from_city,&amp;nbsp;to_city, and so on). If the schemas in your registry are clearly distinguishable, a model that retrieves the wrong one tends to notice, and you won't see a collapse this large.&lt;/P&gt;
&lt;P data-line="382"&gt;Measure before you fine-tune.&lt;/P&gt;
&lt;H3 data-line="384"&gt;What we haven't answered yet&lt;/H3&gt;
&lt;P data-line="386"&gt;We are reporting a finding, not a finished agenda. The open questions we think matter most, roughly in order of how cheaply you could settle them:&lt;/P&gt;
&lt;UL data-line="389"&gt;
&lt;LI data-line="390"&gt;&lt;STRONG&gt;Does the collapse survive a better retriever?&lt;/STRONG&gt;&amp;nbsp;We used BM25 top-1 on character 3–5-grams — deliberately weak, because that is what a registry lookup usually is. A dense retriever with a cross-encoder reranker would reduce how often the wrong schema is in the prompt, but Figure 2 says it cannot fix the turns where the right schema was already there.&lt;/LI&gt;
&lt;LI data-line="395"&gt;&lt;STRONG&gt;Does it reproduce when the output is a tool call?&lt;/STRONG&gt;&amp;nbsp;Everything here decodes a belief state. A function-call target has its own constrained decoding and its own schema-validation layer, and either could mask or amplify the effect. Untested.&lt;/LI&gt;
&lt;LI data-line="399"&gt;&lt;STRONG&gt;Can schema selection replace fine-tuning?&lt;/STRONG&gt;&amp;nbsp;Making the model pick one schema first and fill it second is a much cheaper intervention than training. It also moves the failure rather than removing it — the picking step is the same discrimination problem. Worth measuring before anyone spends $150.&lt;/LI&gt;
&lt;LI data-line="403"&gt;&lt;STRONG&gt;What is the right abstention ratio?&lt;/STRONG&gt;&amp;nbsp;20% distractor-only examples is a round number that worked, not a tuned one. Abstention also turned out to be a separate skill from schema-following: fine-tuning lifted the gold-present turns enormously and barely moved behaviour on turns where the correct schema was genuinely absent.&lt;/LI&gt;
&lt;LI data-line="408"&gt;&lt;STRONG&gt;Does one fine-tune transfer to a different registry?&lt;/STRONG&gt;&amp;nbsp;The recipe transfers across&amp;nbsp;&lt;EM&gt;base models&lt;/EM&gt;&amp;nbsp;— we re-ran it unchanged on a second one for +28.0 points. Whether a model trained on one organisation's schemas follows a different organisation's schemas is a separate question, and the answer decides whether this is a one-time cost or a per-registry cost.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-line="414"&gt;If you run the runbook on your own registry, we would like to know what number you get — especially if you get no gap at all. Negative results on the "when this won't apply" boundary are the most useful thing anyone could send us. Leave a comment on this post with how many schemas are in your registry, how confusable they are, and the two numbers from Steps 2 and 3. That is enough for us to tell whether it is the same failure mode.&lt;/P&gt;
&lt;H2 data-line="422"&gt;Try it yourself (~20 minutes)&lt;/H2&gt;
&lt;BLOCKQUOTE&gt;
&lt;P data-line="424"&gt;The step numbers match the sample's README, so you can move between the two. &lt;STRONG&gt;Steps 2–4 reproduce the gap&lt;/STRONG&gt; and need no fine-tuning; they cost only inference on about 400 turns. &lt;STRONG&gt;Step 5 is the fix&lt;/STRONG&gt;, written out so you can read the recipe without running it.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;
&lt;P data-line="428"&gt;&lt;STRONG&gt;Prerequisites:&lt;/STRONG&gt;&amp;nbsp;Python 3.10+, a Microsoft Foundry deployment of a small chat model (we used&amp;nbsp;gpt-4.1-mini), and the&amp;nbsp;&lt;STRONG&gt;Foundry User&lt;/STRONG&gt; role on that resource. The sample is keyless — it authenticates with&amp;nbsp;DefaultAzureCredential, so there is no API key to paste anywhere. The sample README has the full access walkthrough.&lt;/P&gt;
&lt;LI-CODE lang="shell"&gt;git clone https://github.com/microsoft-foundry/foundry-samples.git
cd foundry-samples/samples/python/finetuning/schemaraft-dst
python -m venv .venv; .\.venv\Scripts\Activate.ps1
pip install -r requirements.txt

Copy-Item .env.example .env   # set AOAI_ENDPOINT and DEPLOYMENT
az login&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Step 0 — confirm the checkout works.&lt;/STRONG&gt;&amp;nbsp;Offline and free: no Azure calls, no tokens. If this fails, nothing after it will work.&lt;/P&gt;
&lt;LI-CODE lang="shell"&gt;python selftest.py            # 42 offline checks&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Step 1 — get the data.&lt;/STRONG&gt;&amp;nbsp;The test split only; training data is not needed until Step 5.&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;python download_sgd.py --splits test&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Step 2 — the number papers report.&lt;/STRONG&gt;&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;python eval_jga.py --mode oracle `
    --deployment &amp;lt;your-deployment&amp;gt; --n 200 `
    --output-dir results/oracle&lt;/LI-CODE&gt;
&lt;P data-line="465"&gt;&lt;STRONG&gt;Step 3 — the number your users get.&lt;/STRONG&gt;&amp;nbsp;One flag changes.&lt;/P&gt;
&lt;LI-CODE lang="shell"&gt;python eval_jga.py --mode retrieved_distractor `
    --deployment &amp;lt;your-deployment&amp;gt; --n 200 `
    --n-schemas 1 --n-distractors 2 `
    --output-dir results/retdist&lt;/LI-CODE&gt;
&lt;P data-line="477"&gt;Compare&amp;nbsp;results/*/summary.json. Expect roughly 58 vs 18.&lt;/P&gt;
&lt;P data-line="479"&gt;&lt;STRONG&gt;Step 4 — confirm it isn't retrieval.&lt;/STRONG&gt; Score only the turns where the correct schema really made it into the prompt:&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;python goldpresent_strata.py --results results/retdist&lt;/LI-CODE&gt;
&lt;P data-line="486"&gt;They stay far below the oracle number.&lt;/P&gt;
&lt;P data-line="488"&gt;&lt;STRONG&gt;Step 5 — the fix, for reference.&lt;/STRONG&gt; Don't run this one unless you mean it: it is a real fine-tuning job (~$150, 6–10h in queue). It is here so the recipe is something you can read, not just a link.&lt;/P&gt;
&lt;LI-CODE lang="bash"&gt;python download_sgd.py --splits train
python build_sft_data.py --split train --n-distractors 2 --gold-absent-frac 0.2
python submit_finetune.py --train-file data/schemaraft_train.jsonl --epochs 3&lt;/LI-CODE&gt;
&lt;P&gt;The job writes&amp;nbsp;finetune_job.json, so&amp;nbsp;--resume&amp;nbsp;reattaches instead of paying for a second run. Expect it to sit in the queue for a few hours before training starts; you can watch both phases in the Foundry portal's fine-tuning section. Once the status reads&amp;nbsp;&lt;EM&gt;succeeded&lt;/EM&gt;, the new model appears alongside the base models when you deploy. Point --deployment at it and re-run Step 3 — same command, same data, same metric. That is the whole experiment.&lt;/P&gt;
&lt;H2 data-line="493"&gt;What to do with this&lt;/H2&gt;
&lt;P data-line="509"&gt;If you are shipping an agent that retrieves anything into its prompt — schemas, tools, MCP servers, API definitions — the cheap version of this post is one afternoon:&lt;/P&gt;
&lt;OL data-line="513"&gt;
&lt;LI data-line="513"&gt;&lt;STRONG&gt;Re-run your own eval with distractors in the prompt.&lt;/STRONG&gt;&amp;nbsp;Not a better retriever, not a bigger model — just stop handing the model the right answer. Steps 2–4 above do exactly this and cost only inference.&lt;/LI&gt;
&lt;LI data-line="516"&gt;&lt;STRONG&gt;If the number holds, you are done.&lt;/STRONG&gt;&amp;nbsp;Your registry is not confusable enough for this failure mode and you should not spend money on it.&lt;/LI&gt;
&lt;LI data-line="518"&gt;&lt;STRONG&gt;If it collapses, you now know that prompting won't save you&lt;/STRONG&gt;&amp;nbsp;— that is what Figure 3 bought you — and the fix is a supervised fine-tuning job on the same deployment you are already using.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P data-line="522"&gt;The part we would push back on is skipping step 1. The gap in this post was invisible to a standard evaluation, survived perfect retrieval, and survived gpt-5. It would have shipped.&lt;/P&gt;
&lt;P data-line="508"&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2 data-line="514"&gt;Appendix — exact setup&lt;/H2&gt;
&lt;P data-line="530"&gt;Collected here so the narrative above stays readable and the numbers above stay auditable.&lt;/P&gt;
&lt;P data-line="533"&gt;&lt;STRONG&gt;Data.&lt;/STRONG&gt;&amp;nbsp;Schema-Guided Dialogue (SGD) [2]: 45 services across 20 domains. The test-time retrieval registry is the 21 services in the test split; 15 of them never appear in training, which is the point of the benchmark. Full test split: 39,324 turns.&lt;/P&gt;
&lt;P data-line="538"&gt;&lt;STRONG&gt;Retrieval (RetDist).&lt;/STRONG&gt;&amp;nbsp;BM25 over character 3–5-grams of the schema text, top-1, plus two distractor schemas sampled uniformly from the same registry. The model never sees a retrieval score — BM25's job is selection, not ranking signal. On the test split the gold schema is top-1 for 65.4% of turns; random distractor sampling surfaces it on some of the rest, leaving 31.1% of turns where the correct schema is genuinely absent from the prompt.&lt;/P&gt;
&lt;P data-line="545"&gt;&lt;STRONG&gt;Evaluation populations.&lt;/STRONG&gt; Three, and they must not be mixed:&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Population&lt;/th&gt;&lt;th&gt;n&lt;/th&gt;&lt;th&gt;Used for&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;SGD-200&lt;/td&gt;&lt;td&gt;1,810 turns (200 dialogues)&lt;/td&gt;&lt;td&gt;Figure 3, the economics table, all frontier comparisons — frontier models are too expensive to run on the full split&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Gold-in-prompt stratum&lt;/td&gt;&lt;td&gt;27,085 turns&lt;/td&gt;&lt;td&gt;Figure 2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Full test split&lt;/td&gt;&lt;td&gt;39,324 turns&lt;/td&gt;&lt;td&gt;Fine-tuning effects only&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;col style="width: 33.33%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P data-line="553"&gt;The gold-in-prompt stratum is every turn where the correct schema reached the prompt — whether BM25 ranked it top-1 (65.4% of turns) or a random distractor draw happened to surface it. The remaining 12,239 turns (31.1%) are the ones where it is genuinely absent.&lt;/P&gt;
&lt;P data-line="558"&gt;&lt;STRONG&gt;Models.&lt;/STRONG&gt;&amp;nbsp;Base and fine-tuned:&amp;nbsp;gpt-4.1-mini. Prompted baselines:&amp;nbsp;gpt-5,&amp;nbsp;Llama-3.3-70B. All three are served from the same Foundry deployment, which is what makes the latency and relative-cost columns comparable. The prompted baselines were not fine-tuned:&amp;nbsp;gpt-5&amp;nbsp;exposed no SFT path on Foundry at the time of these runs, and fine-tuning a 70B model is far outside the cost envelope this post is about.&lt;/P&gt;
&lt;P data-line="565"&gt;&lt;STRONG&gt;Fine-tuning.&lt;/STRONG&gt;&amp;nbsp;~19,300 chat-format examples, ~2,100 prompt tokens and ~40 completion tokens each; every example carries one gold schema plus two random distractors; 20% are distractor-only and labelled&amp;nbsp;{}. Supervised fine-tuning, 3 epochs, ~124M training tokens, one job on Microsoft Foundry.&lt;/P&gt;
&lt;P data-line="570"&gt;&lt;STRONG&gt;Cost basis.&lt;/STRONG&gt;&amp;nbsp;Latency and relative cost are measured per turn on that same deployment, not quoted from a price sheet. The ~$150 figure is the one-time training job; the ~18,000-turn break-even is that cost divided by the measured per-turn saving against&amp;nbsp;gpt-5.&lt;/P&gt;
&lt;P data-line="575"&gt;&lt;STRONG&gt;Reproducibility.&lt;/STRONG&gt; Steps 2–4 of the runbook reproduce the gap and cost only inference. Step 5 reproduces the fix and is a real fine-tuning job.&lt;/P&gt;
&lt;H2 data-line="561"&gt;References&lt;/H2&gt;
&lt;OL data-line="563"&gt;
&lt;LI data-line="563"&gt;Zhang, T., Patil, S. G., Jain, N., Shen, S., Zaharia, M., Stoica, I., and Gonzalez, J. E.&amp;nbsp;&lt;A href="https://arxiv.org/abs/2403.10131" target="_blank" rel="noopener" data-href="https://arxiv.org/abs/2403.10131"&gt;&lt;STRONG&gt;RAFT: Adapting Language Model to Domain Specific RAG.&lt;/STRONG&gt;&lt;/A&gt;&amp;nbsp;COLM, 2024.&lt;/LI&gt;
&lt;LI data-line="566"&gt;Rastogi, A., Zang, X., Sunkara, S., Gupta, R., and Khaitan, P.&amp;nbsp;&lt;A href="https://arxiv.org/abs/1909.05855" target="_blank" rel="noopener" data-href="https://arxiv.org/abs/1909.05855"&gt;&lt;STRONG&gt;Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue Dataset.&lt;/STRONG&gt;&lt;/A&gt;&amp;nbsp;AAAI, 2020.&amp;nbsp;&lt;EM&gt;(the SGD benchmark used throughout)&lt;/EM&gt;&lt;/LI&gt;
&lt;LI data-line="570"&gt;Zhao, J., Gupta, R., Cao, Y., Yu, D., Wang, M., Lee, H., Rastogi, A., Shafran, I., and Wu, Y.&amp;nbsp;&lt;A href="https://arxiv.org/abs/2201.08904" target="_blank" rel="noopener" data-href="https://arxiv.org/abs/2201.08904"&gt;&lt;STRONG&gt;Description-Driven Task-Oriented Dialog Modeling.&lt;/STRONG&gt;&lt;/A&gt;&amp;nbsp;ACL, 2022.&amp;nbsp;&lt;EM&gt;(D3ST — 86.4% JGA under an oracle schema)&lt;/EM&gt;&lt;/LI&gt;
&lt;LI data-line="574"&gt;Bang, N., Lee, J., and Koo, M.-W.&amp;nbsp;&lt;A href="https://arxiv.org/abs/2305.02468" target="_blank" rel="noopener" data-href="https://arxiv.org/abs/2305.02468"&gt;&lt;STRONG&gt;Task-Optimized Adapters for an End-to-End Task-Oriented Dialogue System.&lt;/STRONG&gt;&lt;/A&gt;&amp;nbsp;Findings of ACL, 2023.&amp;nbsp;&lt;EM&gt;(TOATOD — 74.7% JGA under an oracle schema)&lt;/EM&gt;&lt;/LI&gt;
&lt;LI data-line="577"&gt;Lee, H., Gupta, R., Rastogi, A., Cao, Y., Zhang, B., and Wu, Y.&amp;nbsp;&lt;A href="https://arxiv.org/abs/2110.06800" target="_blank" rel="noopener" data-href="https://arxiv.org/abs/2110.06800"&gt;&lt;STRONG&gt;SGD-X: A Benchmark for Robust Generalization in Schema-Guided Dialogue Systems.&lt;/STRONG&gt;&lt;/A&gt;&amp;nbsp;AAAI, 2022.&amp;nbsp;&lt;EM&gt;(prior robustness probe — stylistic schema variants, still oracle delivery)&lt;/EM&gt;&lt;/LI&gt;
&lt;LI data-line="581"&gt;Asai, A., Wu, Z., Wang, Y., Sil, A., and Hajishirzi, H.&amp;nbsp;&lt;A href="https://arxiv.org/abs/2310.11511" target="_blank" rel="noopener" data-href="https://arxiv.org/abs/2310.11511"&gt;&lt;STRONG&gt;Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection.&lt;/STRONG&gt;&lt;/A&gt; ICLR, 2024.&lt;/LI&gt;
&lt;/OL&gt;</description>
      <pubDate>Tue, 15 Sep 2026 17:35:05 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/your-agent-found-the-right-schema-then-ignored-it/ba-p/4553775</guid>
      <dc:creator>shihyaolin</dc:creator>
      <dc:date>2026-09-15T17:35:05Z</dc:date>
    </item>
    <item>
      <title>Data Meets Agents</title>
      <link>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/data-meets-agents/ba-p/4553082</link>
      <description>&lt;DIV class="lia-align-justify"&gt;
&lt;H2&gt;Introduction&lt;/H2&gt;
&lt;/DIV&gt;
&lt;P class="lia-align-justify"&gt;Enterprise agents become more valuable when they can securely access the data, knowledge and business systems organisations already rely on. And, &lt;STRONG&gt;Data Meets Agents&lt;/STRONG&gt; is a way of framing a recurring architecture question we hear from customers: &lt;EM&gt;How should an agent access enterprise data that already resides across Microsoft Fabric, Azure Databricks, Snowflake, Google BigQuery, Amazon S3/Redshift, Salesforce and other platforms? How can we make that data and its insights available to agents in a secure and governed way? And are there repeatable architecture patterns and best practices that we can adopt across the organisation?&lt;/EM&gt;&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;The challenge is not simply finding a connector. It is deciding &lt;STRONG&gt;how the agent should access that data while preserving the right security, governance, semantics, freshness and system-of-record boundaries&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;Microsoft provides several ways to build agents, from &lt;STRONG&gt;Microsoft first-party agents and Microsoft 365 Agent Builder to Copilot Studio and Microsoft Foundry Agent Service&lt;/STRONG&gt;, each serving different personas and levels of extensibility.&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;This umbrella blog briefly sets that broader agent-building landscape, but focuses on &lt;STRONG&gt;Microsoft Foundry Agent Service&lt;/STRONG&gt; as a starting point and three complementary patterns for integrating enterprise data:&lt;/P&gt;
&lt;OL class="lia-align-justify"&gt;
&lt;LI&gt;&lt;STRONG&gt;Direct Tool Access through APIs/MCP&lt;/STRONG&gt; – for live access to data and business operations.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;OneLake + Fabric IQ&lt;/STRONG&gt; – for governed enterprise data, analytics and business semantics.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Foundry IQ knowledge grounding&lt;/STRONG&gt; – for reusable enterprise knowledge and retrieval.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P class="lia-align-justify"&gt;Although &lt;STRONG&gt;Azure API Management (APIM)&lt;/STRONG&gt; is not explicitly shown in the simplified architecture diagrams, we strongly recommend considering it as a &lt;STRONG&gt;runtime governance layer across these integration patterns&lt;/STRONG&gt;, where applicable, to provide consistent controls around API and MCP exposure, authentication, policy enforcement, observability and lifecycle management.&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;We’ll explore where each pattern fits, the trade-offs involved, and how they can be combined within a &lt;STRONG&gt;multi-pattern enterprise architecture&lt;/STRONG&gt;.&lt;/P&gt;
&lt;DIV class="lia-align-justify"&gt;
&lt;H2&gt;Microsoft Agent-Building Platform Landscape&lt;/H2&gt;
&lt;/DIV&gt;
&lt;P class="lia-align-justify"&gt;Microsoft provides several approaches to building agents, designed for different personas, use cases and levels of extensibility.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN lia-align-justify"&gt;&lt;table border="1" style="width: 100%; height: 295px; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr style="height: 35px;"&gt;&lt;th class="lia-align-left" style="height: 35px;"&gt;Platform / Approach&lt;/th&gt;&lt;th class="lia-align-left" style="height: 35px;"&gt;Primary audience&lt;/th&gt;&lt;th class="lia-align-left" style="height: 35px;"&gt;Development model&lt;/th&gt;&lt;th class="lia-align-left" style="height: 35px;"&gt;Typical use cases&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr style="height: 59px;"&gt;&lt;td class="lia-align-left" style="height: 59px;"&gt;&lt;STRONG&gt;Microsoft first-party agents&lt;/STRONG&gt;&lt;/td&gt;&lt;td class="lia-align-left" style="height: 59px;"&gt;Business users&lt;/td&gt;&lt;td class="lia-align-left" style="height: 59px;"&gt;Ready-to-use / configurable&lt;/td&gt;&lt;td class="lia-align-left" style="height: 59px;"&gt;Purpose-built agent experiences across Microsoft products&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 59px;"&gt;&lt;td class="lia-align-left" style="height: 59px;"&gt;&lt;STRONG&gt;Microsoft 365 Agent Builder&lt;/STRONG&gt;&lt;/td&gt;&lt;td class="lia-align-left" style="height: 59px;"&gt;Information workers&lt;/td&gt;&lt;td class="lia-align-left" style="height: 59px;"&gt;No-code / declarative&lt;/td&gt;&lt;td class="lia-align-left" style="height: 59px;"&gt;Lightweight agents grounded in Microsoft 365 knowledge&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 59px;"&gt;&lt;td class="lia-align-left" style="height: 59px;"&gt;&lt;STRONG&gt;Microsoft Copilot Studio&lt;/STRONG&gt;&lt;/td&gt;&lt;td class="lia-align-left" style="height: 59px;"&gt;Makers and developers&lt;/td&gt;&lt;td class="lia-align-left" style="height: 59px;"&gt;Low-code with pro-code extensibility&lt;/td&gt;&lt;td class="lia-align-left" style="height: 59px;"&gt;Enterprise agents, workflows and business process automation&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 83px;"&gt;&lt;td class="lia-align-left" style="height: 83px;"&gt;&lt;STRONG&gt;Microsoft Foundry Agent Service&lt;/STRONG&gt;&lt;/td&gt;&lt;td class="lia-align-left" style="height: 83px;"&gt;Developers, AI engineers and architects&lt;/td&gt;&lt;td class="lia-align-left" style="height: 83px;"&gt;Pro-code with declarative options&lt;/td&gt;&lt;td class="lia-align-left" style="height: 83px;"&gt;Custom agentic applications, complex orchestration and deep enterprise integration&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P class="lia-align-justify"&gt;Across this ecosystem,&amp;nbsp;&lt;STRONG&gt;Agent 365 provides the cross-platform control plane for governing the enterprise agent estate&lt;/STRONG&gt;, bringing together agent identity, inventory, security, compliance, lifecycle management and observability.&lt;/P&gt;
&lt;DIV class="lia-align-justify"&gt;
&lt;H2&gt;Microsoft Intelligence Layer&lt;/H2&gt;
&lt;/DIV&gt;
&lt;P class="lia-align-justify"&gt;Alongside the agent-building platforms, Microsoft provides intelligence capabilities that help agents understand and reason over different types of context.&lt;/P&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN lia-align-justify"&gt;&lt;table border="1" style="width: 100%; border-width: 1px;"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Intelligence&lt;/th&gt;&lt;th&gt;Primary role&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Work IQ&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Organisational context from Microsoft 365, including people, communications, meetings and files&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Fabric IQ&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Business and analytical context through governed data, semantic models, entities and relationships&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Foundry IQ&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;Enterprise knowledge grounding, retrieval and reusable knowledge across multiple sources&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;STRONG&gt;Web IQ&lt;/STRONG&gt;&lt;/td&gt;&lt;td&gt;External context and knowledge from the web&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 17.8087%" /&gt;&lt;col style="width: 82.2531%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;P class="lia-align-justify"&gt;For &lt;STRONG&gt;Data Meets Agents&lt;/STRONG&gt;, however, the focus is not the IQ capabilities themselves. Our focus is the enterprise data that &lt;STRONG&gt;already resides across different data platforms and operational systems&lt;/STRONG&gt; — such as Azure Databricks, Snowflake and other cloud data platforms, as well as CRM, ticketing and line-of-business applications so focus will be Fabric IQ and Foundry IQ from intelligence layer.&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;&amp;nbsp;&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;For this umbrella blog, we briefly establish broader agent building landscape and Microsoft IQ before focusing on &lt;STRONG&gt;Microsoft Foundry Agent Service&lt;/STRONG&gt; and the three &lt;STRONG&gt;Data Meets Agents&lt;/STRONG&gt; integration patterns. The next section focuses on&amp;nbsp;&lt;STRONG&gt;Microsoft Foundry Agent Service&lt;/STRONG&gt; and three architectural patterns for enabling agents to access that existing enterprise data.&amp;nbsp;&lt;/P&gt;
&lt;DIV class="lia-align-justify"&gt;
&lt;H2&gt;Why Microsoft Foundry Agent Service?&lt;/H2&gt;
&lt;/DIV&gt;
&lt;P class="lia-align-justify"&gt;&lt;STRONG&gt;Microsoft Foundry Agent Service&lt;/STRONG&gt; is positioned for organisations that need greater control over &lt;STRONG&gt;agent orchestration, enterprise integration and how agents execute and interact with tools, data and other agents&lt;/STRONG&gt;. Developers can build sophisticated agents and multi-agent orchestration using &lt;STRONG&gt;Microsoft Agent Framework or other code-based agent frameworks&lt;/STRONG&gt;, and deploy that application logic to Foundry as &lt;STRONG&gt;Hosted Agents&lt;/STRONG&gt;, retaining control over their agent and orchestration logic. Foundry provides the managed &lt;STRONG&gt;control plane and runtime capabilities&lt;/STRONG&gt; around those agents, including &lt;STRONG&gt;deployment and hosting, scaling, identity and access, endpoints, tool integration, tracing, monitoring, evaluation and lifecycle management&lt;/STRONG&gt;. This combination of &lt;STRONG&gt;framework flexibility, code-level control and a managed enterprise control plane&lt;/STRONG&gt; makes Foundry particularly relevant for complex agentic applications that need to orchestrate across multiple enterprise data platforms, knowledge sources and business systems.&lt;/P&gt;
&lt;DIV class="lia-align-justify"&gt;
&lt;H2&gt;Enterprise Data Integration Patterns with Foundry Agent Service&lt;/H2&gt;
&lt;/DIV&gt;
&lt;P class="lia-align-justify"&gt;Enterprise data already resides across &lt;STRONG&gt;cloud data platforms, databases and operational systems such as CRM, ticketing and line-of-business applications&lt;/STRONG&gt;. The architectural question is therefore not simply how to connect an agent to that data, but:&amp;nbsp;&lt;STRONG&gt;Where should retrieval and semantic reasoning happen, under whose identity, and where should governance be enforced?&lt;/STRONG&gt;&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;With &lt;STRONG&gt;Microsoft Foundry Agent Service&lt;/STRONG&gt;, we frame this through three complementary patterns: &lt;STRONG&gt;Direct Tool Access, OneLake + Fabric IQ, and Foundry IQ knowledge grounding&lt;/STRONG&gt;. The right pattern depends on whether the agent needs &lt;STRONG&gt;live operational access, governed analytical semantics, or reusable enterprise knowledge&lt;/STRONG&gt; — and a single solution&lt;STRONG&gt; may combine all three.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;&amp;nbsp;&lt;/P&gt;
&lt;img /&gt;
&lt;DIV class="lia-align-justify"&gt;
&lt;H3&gt;Pattern 1 — Direct Tool Access: APIs and MCP&lt;/H3&gt;
&lt;/DIV&gt;
&lt;P class="lia-align-justify"&gt;The Foundry agent accesses an external platform at runtime through a governed API or MCP capability, keeping the external platform as the system of record.&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;&lt;STRONG&gt;Prefer when:&lt;/STRONG&gt; the agent needs live data, an exact lookup, a platform-native capability, or a controlled business action — for example querying Snowflake or Databricks, retrieving a CRM record, checking a ticket or triggering a workflow.&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;&lt;STRONG&gt;Architecture focus: &lt;/STRONG&gt;identity, least privilege, runtime governance, tool design, latency and operational reliability.&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;&lt;STRONG&gt; &lt;/STRONG&gt;&lt;/P&gt;
&lt;DIV class="lia-align-justify"&gt;
&lt;H3&gt;Pattern 2 — Unified Data Platform: OneLake + Fabric IQ&lt;/H3&gt;
&lt;/DIV&gt;
&lt;P class="lia-align-justify"&gt;Although Fabric capabilities can also be exposed to Foundry through MCP, we show this as a separate architecture pattern because the data strategy is fundamentally different.&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;Rather than having the agent directly query each external platform, enterprise data residing in Databricks, Snowflake and other cloud data platforms can be referenced, mirrored or intentionally ingested through OneLake. Fabric then provides a governed analytical and semantic layer through capabilities such as Fabric Data Agents, ontologies and Power BI semantic models, which Foundry can consume.&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;&lt;STRONG&gt;Fabric is the governed data and semantic layer between agents and heterogeneous enterprise data.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;&lt;STRONG&gt;Prefer when: &lt;/STRONG&gt;agents need shared analytics, governed business semantics, measures, entities and relationships across enterprise data, rather than direct access to individual source systems.&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;&lt;STRONG&gt;Architecture focus: &lt;/STRONG&gt;reference versus replication, semantic ownership, data freshness, Fabric governance and reuse across AI and analytics.&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;&lt;STRONG&gt; &lt;/STRONG&gt;&lt;/P&gt;
&lt;DIV class="lia-align-justify"&gt;
&lt;H3&gt;Pattern 3 — Foundry IQ Knowledge Grounding&lt;/H3&gt;
&lt;/DIV&gt;
&lt;P class="lia-align-justify"&gt;Enterprise knowledge is exposed through Foundry IQ knowledge bases, providing reusable retrieval and grounding across indexed and remote knowledge sources.&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;&lt;STRONG&gt;Prefer when: &lt;/STRONG&gt;agents need to discover, retrieve and reason over trusted organisational knowledge with grounding and citations, rather than perform analytical queries or operational transactions.&lt;/P&gt;
&lt;P class="lia-align-justify"&gt;&lt;STRONG&gt;Architecture focus: &lt;/STRONG&gt;knowledge curation, retrieval quality, permissions, citations, reuse and knowledge governance.&lt;/P&gt;
&lt;DIV class="lia-align-justify"&gt;
&lt;H3&gt;&lt;STRONG data-olk-copy-source="MessageBody"&gt;Quick Pattern Selection Guide&lt;/STRONG&gt;&amp;nbsp;&lt;/H3&gt;
&lt;/DIV&gt;
&lt;DIV class="styles_lia-table-wrapper__h6Xo9 styles_table-responsive__MW0lN"&gt;&lt;table border="1" style="width: 1059px; height: 394.333px; border-width: 1px;"&gt;&lt;tbody&gt;&lt;tr style="height: 47.6354px;"&gt;&lt;td style="height: 47.6354px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Decision Criteria&lt;/STRONG&gt;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 47.6354px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Direct as a Tool&lt;/STRONG&gt;&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 47.6354px;"&gt;
&lt;P&gt;&lt;STRONG&gt;OneLake&amp;nbsp;+ Fabric IQ&lt;/STRONG&gt;&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 47.6354px;"&gt;
&lt;P&gt;&lt;STRONG&gt;Foundry IQ&lt;/STRONG&gt;&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 47.6354px;"&gt;&lt;td style="height: 47.6354px;"&gt;
&lt;P&gt;Data location&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 47.6354px;"&gt;
&lt;P&gt;Remains in source&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 47.6354px;"&gt;
&lt;P&gt;Replicated or referenced&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 47.6354px;"&gt;
&lt;P&gt;Indexed or remotely accessed&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 87.1354px;"&gt;&lt;td style="height: 87.1354px;"&gt;
&lt;P&gt;Freshness&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 87.1354px;"&gt;
&lt;P&gt;Runtime source result&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 87.1354px;"&gt;
&lt;P&gt;Depends on integration/query mechanism&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 87.1354px;"&gt;
&lt;P&gt;Depends on index freshness or remote source&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 82.1354px;"&gt;&lt;td style="height: 82.1354px;"&gt;
&lt;P&gt;Best for&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 82.1354px;"&gt;
&lt;P&gt;Live query and controlled operation&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 82.1354px;"&gt;
&lt;P&gt;Shared analytics and business semantics&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 82.1354px;"&gt;
&lt;P&gt;Reusable grounding and unstructured knowledge&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 82.1354px;"&gt;&lt;td style="height: 82.1354px;"&gt;
&lt;P&gt;Main control point&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 82.1354px;"&gt;
&lt;P&gt;MCP server + source permissions&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 82.1354px;"&gt;
&lt;P&gt;Source + Fabric governance&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 82.1354px;"&gt;
&lt;P&gt;Knowledge base + source permissions&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;tr style="height: 47.6562px;"&gt;&lt;td style="height: 47.6562px;"&gt;
&lt;P&gt;Can&amp;nbsp;combine?&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 47.6562px;"&gt;
&lt;P&gt;Yes&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 47.6562px;"&gt;
&lt;P&gt;Yes&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;td style="height: 47.6562px;"&gt;
&lt;P&gt;Yes&amp;nbsp;&lt;/P&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;colgroup&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;col style="width: 25.00%" /&gt;&lt;/colgroup&gt;&lt;/table&gt;&lt;/DIV&gt;
&lt;DIV class="lia-align-justify"&gt;
&lt;H2&gt;Conclusion&lt;/H2&gt;
&lt;/DIV&gt;
&lt;P class="lia-align-justify"&gt;&lt;STRONG&gt; &lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;The three patterns are complementary.&lt;/STRONG&gt; A single Foundry agent — or a &lt;STRONG&gt;multi-agent solution orchestrated using Microsoft Agent Framework or other agent frameworks and deployed to Foundry&lt;/STRONG&gt; — may use &lt;STRONG&gt;Foundry IQ for enterprise knowledge, Fabric IQ for governed analytical context, and Direct Tool Access for live queries or business actions&lt;/STRONG&gt;. The architectural decision should be made &lt;STRONG&gt;per data source and use case&lt;/STRONG&gt;, rather than standardising every enterprise integration on a single approach.&lt;/P&gt;
&lt;P&gt;&lt;SPAN data-olk-copy-source="MessageBody"&gt;This post establishes the architectural framework; follow-up posts will provide implementation details for each pattern.&lt;/SPAN&gt;— &lt;STRONG&gt;Direct Tool Access with APIs/MCP, Unified Data Platform with OneLake + Fabric IQ, and Foundry IQ knowledge grounding&lt;/STRONG&gt; — covering architecture, identity, governance and key design trade-offs.&lt;/P&gt;
&lt;P&gt;We’ll also explore &lt;STRONG&gt;data-source-specific architectures&lt;/STRONG&gt;, looking at how enterprise data already residing in platforms such as &lt;STRONG&gt;Azure Databricks, Snowflake, Microsoft Fabric and other data platforms and business systems&lt;/STRONG&gt; can be integrated with &lt;STRONG&gt;Microsoft Foundry Agent Service&lt;/STRONG&gt;. Where relevant and supported, we’ll also explore how these approaches apply to &lt;STRONG&gt;Copilot Studio, Microsoft first-party agents and Microsoft 365 Agent Builder&lt;/STRONG&gt;, helping architects choose the right &lt;STRONG&gt;platform, integration pattern and data architecture&lt;/STRONG&gt; for each use case.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;A special thanks to &lt;A class="lia-internal-link lia-internal-url lia-internal-url-user" href="https://techcommunity.microsoft.com/users/lima/3657645" target="_blank" rel="noopener" data-lia-auto-title="Li Ma" data-lia-auto-title-active="0"&gt;Li Ma&lt;/A&gt; and &lt;A class="lia-internal-link lia-internal-url lia-internal-url-user" href="https://techcommunity.microsoft.com/users/shilfi/3651281" target="_blank" rel="noopener" data-lia-auto-title="Shilfi Gafur" data-lia-auto-title-active="0"&gt;Shilfi Gafur&lt;/A&gt; for reviewing and contributing to this work. I’m also looking forward to collaborating with them on the upcoming Data Meets Agents blogs as we take a deeper dive into these patterns and data-source-specific architectures.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;STRONG data-olk-copy-source="MessageBody"&gt;Stay Connected&lt;/STRONG&gt;&amp;nbsp;- 📝&amp;nbsp;&lt;STRONG&gt;Coming soon&lt;/STRONG&gt;: Deep-dive blogs on each pattern (subscribe to be notified) - 💬&amp;nbsp;&lt;STRONG&gt;Questions?&lt;/STRONG&gt;: Drop a comment below&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;STRONG data-olk-copy-source="MessageBody"&gt;Get Started&lt;/STRONG&gt; - 🚀 ⚡&amp;nbsp;Try now: &lt;A href="https://learn.microsoft.com/en-us/azure/foundry/agents/quickstarts/quickstart-hosted-agent?pivots=azd" target="_blank"&gt;Quickstart: Deploy your first hosted agent - Microsoft Foundry | Microsoft Learn&lt;/A&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://techcommunity.microsoft.com/blog/microsoft-security-blog/microsoft-copilot-studio-vs-microsoft-foundry-building-ai-agents-and-apps/4483160" target="_blank" rel="noopener"&gt;Microsoft Copilot Studio vs. Microsoft Foundry: Building AI Agents and Apps | Microsoft Community Hub&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/foundry/" target="_blank" rel="noopener"&gt;Microsoft Foundry documentation | Microsoft Learn&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://azure.microsoft.com/en-us/blog/agent-factory-building-your-first-ai-agent-with-the-tools-to-deliver-real-world-outcomes/?msockid=346decbeee6c6b19175afb21eff66acc" target="_blank"&gt;Agent Factory: Building your first AI agent with the tools to deliver real-world outcomes | Microsoft Azure Blog&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://learn.microsoft.com/en-us/training/browse/" target="_blank" rel="noopener"&gt;Browse all training - Training | Microsoft Learn&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://learn.microsoft.com/en-us/fabric/onelake/onelake-overview" target="_blank" rel="noopener"&gt;OneLake, the unified data lake - Microsoft Fabric | Microsoft Learn&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/what-is-foundry-iq?tabs=portal" target="_blank" rel="noopener"&gt;What is Foundry IQ? - Microsoft Foundry | Microsoft Learn&lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Mon, 14 Sep 2026 16:00:00 GMT</pubDate>
      <guid>https://techcommunity.microsoft.com/t5/microsoft-foundry-blog/data-meets-agents/ba-p/4553082</guid>
      <dc:creator>Yeliz_Kilinc</dc:creator>
      <dc:date>2026-09-14T16:00:00Z</dc:date>
    </item>
  </channel>
</rss>

