microsoft foundry
119 TopicsBeyond the Trace: The Science of Insight Quality
Co-authors and reviewers: Morteza Ziyadi, Hanchi Wang, Han Che, Billy Hu, Sean Gayler, Nishal Dsilva, Avinav Jami, Ankit Singhal, Augustus Arthur TL;DR. Insights in Foundry turns recurring agent behavior into evidence-linked findings developers can review and act on. We evaluate trace linkage, finding quality, and detection of known issues using labeled benchmarks, LLM-judge assessment, and controlled end-to-end tests. What is an Insight? An Insight is a reviewable finding about recurring agent behavior. It brings together an explanation, supporting trace evidence, and a possible next step, helping developers investigate a pattern rather than inspect each execution in isolation. Depending on the available evidence and supported configuration, an Insight can include: Part of an Insight What it gives the reviewer Title and description The recurring behavior and an evidence-based explanation of a possible cause. Linked traces The broader group of traces associated with the finding. Highlighted traces Representative examples to open and check against the explanation. Category, severity, and status Context for triaging the finding, not a substitute for a risk assessment. Agent version and recency Which version is represented and when the finding was created. Proposed action or fix An investigation or improvement path; concrete proposals depend on the configuration. Figure 1. An Insight detail view in Microsoft Foundry, cropped from public Microsoft Learn documentation. The description, evidence, and proposed fix support human review. This UI illustration is not a benchmark result; the source link provides a full-size view. UI source and field definitions: Insights in Foundry documentation. From trace search to an actionable review queue Production agents can generate thousands of traces containing model calls, tool calls, latency, token usage, errors, and final responses. Observability shows what happened, while evaluations test criteria a team already knows to measure. The harder problem is discovering repeated behavior that the team did not know to predefine. Insights in Foundry analyzes traces from the Application Insights resource connected to a Foundry project and organizes recurring behavior into reviewable Insights. Depending on the evidence and supported configuration, an Insight can include representative traces, affected agent versions, severity, an explanation of a possible cause, and a recommended next action. Concrete prompt or code proposals are available only for supported agent types and configurations; other Insights provide general investigation or remediation guidance. Insights in Foundry is now available in public preview. It analyzes production traces to surface recurring behaviors and regressions with supporting evidence and suggested areas for investigation or improvement. Developers remain in control: review the cited traces and validate proposed changes through normal evaluation and deployment practices. During public preview, we will continue expanding the experience based on customer feedback. To make this concrete, Figure 2 maps 220 inputs from TraceElephant, a public agent-trace benchmark, to seven generated Insights. Figure 3 examines one linked trace. Figure 2. Recorded links from 220 public TraceElephant inputs to seven generated Insights in the September 17 benchmark. The 80-input finding includes the looping trace examined in Figure 3. Titles are editorial paraphrases; the layout is editorial, not a product screenshot or a validation of each diagnosis. Figure 3. A public TraceElephant trace contains 54 model calls; steps 26, 30, 34, 38, 42, 46, and 50 repeat an identical extraction plan. The September 17 benchmark run of the production pipeline links this trace to a finding about missing progress-aware termination. Wording is paraphrased; this is an illustration, not product UI. The proposed intervention still needs evaluation. Three complementary ways to evaluate Insight quality Insight quality involves which inputs a finding links, how well it explains the evidence, and whether it identifies known issues in controlled tests. We combine labeled trace evaluation, unlabeled finding assessment, and controlled end-to-end testing. These are complementary sources of evidence, not mutually exclusive dataset categories. Evidence Question Assessment / results Labeled trace evaluation Do Insights link inputs that reference annotations mark as failing? Reference labels; trace precision and trace recall (%). Unlabeled finding assessment Are generated findings grounded and useful? Eight-dimension LLM judge; mean score (1-5). Controlled end-to-end tests Do Insights identify known injected issues? Known issues and healthy baselines; detections, misses, unsupported findings, duplicates, and unscored cases. Here, 'unlabeled' means reference failure labels are not used for the evaluation; the sources can still contain annotations or reference answers. Public versus internal describes where data comes from, not how it is evaluated. LLM judges can also assess findings from labeled or controlled scenarios. Labeled evaluation: Do Insights link the right traces? Trace precision: among unique benchmark inputs linked by at least one generated Insight, the fraction marked as failing. This measures discrimination only when a dataset includes healthy inputs, and it does not validate the diagnosis. Trace recall: the percentage of failure-labeled benchmark inputs linked by at least one generated Insight. It does not measure how many distinct failure modes were discovered. We evaluated the production pipeline on the same 746 benchmark inputs (653 failure-labeled) on September 15, 16, and 17, 2026. Repeating these inputs does not create 2,238 distinct examples. The six benchmark corpus slices come from public agent-trace research datasets with human reference annotations: AgentRx (Tau-bench retail), AgentRx Magentic-One, TRAIL, AgentErrorBench, TraceElephant, and the MAST-Data human subset. AgentErrorBench is the annotated failure-trajectory dataset released with AgentDebug, a framework for detecting and recovering from agent failures. The two AgentRx slices come from the same public release. The benchmark uses selected and normalized inputs from these releases, not a representative sample of customer production traffic. For each dataset, the chart shows daily generated-Insight counts and trace recall. The table reports equal-weight means of daily trace precision and recall. These datasets support ongoing development; results depend on the dataset and model used. Labeled results Figure 4. Generated Insight count versus trace recall for all six public corpus slices. Each point is one daily run; labels retain values where markers overlap. Recall measures linkage to failure-labeled inputs, not diagnosis correctness. Three-day mean precision and recall follow in the table. Dataset Inputs/day Failure-labeled/day Linked/day range Mean trace precision Mean trace recall AgentRx (Tau-bench retail) 102 29 43-45 37.8% 57.5% AgentRx Magentic-One 58 44 54-57 75.3% 94.7% TRAIL 148 143 139-140 96.9% 94.6% AgentErrorBench 200 200 194-196 100.0%** 97.3% TraceElephant 220 220 218-220 100.0%** 99.7% MAST-Data (human subset) 18 17 14-16 93.3% 82.4% ** When every input is failure-labeled, 100% trace precision cannot measure false-positive control. What the labeled results show High precision can reflect the corpus base rate. AgentErrorBench and TraceElephant contain only failure-labeled inputs, so their measured precision is 100% by construction. TRAIL, AgentRx Magentic-One, and the MAST subset also contain mostly failures; their precision values are close to their underlying failure-label prevalence. Read precision beside trace recall and each dataset's label prevalence rather than as a standalone quality score. AgentRx (Tau-bench retail) is the clearest mixed-traffic stress test. Only 29 of 102 inputs were failure-labeled. Across the three days, trace precision was 43.2%, 37.8%, and 32.6%, averaging 37.8%; trace recall was 65.5%, 58.6%, and 48.3%, averaging 57.5%. Daily linked counts were 44, 45, and 43, including 19, 17, and 14 labeled failures respectively. These daily rates are averaged before rounding. Some linked inputs without failure labels could contain issues outside the reference labels, such as cost or latency, but this benchmark cannot confirm that. Trace precision and trace recall measure whether Insights link inputs that reference annotations mark as failing. They do not establish whether an explanation is correct or a proposed action is useful. Our unlabeled evaluation examines those qualities using an eight-dimension LLM judge. Unlabeled evaluation: Are the findings grounded and useful? For this evaluation, a separate LLM judge scores generated Insights against the supplied input evidence, without using reference failure labels to compute precision or recall. These scores are automated assessments produced by an LLM judge. They are not human ratings and do not establish that a proposed action will improve the agent. This evaluation covers PUPA, FailSafeQA, tau2-bench, synthetic scenarios, and an internal production dataset for the monitoring dashboard agent compiled from Microsoft employee usage. These sources are normalized into benchmark inputs; not every input is a complete execution trace. The eight-dimension judge rubric Each generated Insight receives a score from 1 to 5 on eight dimensions. The rubric separates the quality of the explanation, the importance of the issue, and the usefulness of the suggested action: Dimension Question the judge considers Actionability Does the Insight tell a developer what to investigate, evaluate, or change next? Specificity Does it name concrete behavior, tools, or prompt elements rather than offer generic advice? Novelty Does it surface a pattern beyond an obvious dashboard signal? This is an estimate, not a measurement of what a team already knows. Correctness Is the claim supported by the supplied evidence? Severity calibration Does the assigned severity match the evidenced impact? Impact How consequential does the underlying issue appear, rather than just how well it is described? Fix specificity Does the proposed fix identify an exact asset or behavior to change? Fix applicability Is that proposed change plausibly on target for the issue? For example, a concrete next step earns a stronger actionability rating than generic advice. Severity calibration and impact are deliberately separate: a minor issue can have a well-calibrated severity label without being high-impact. Fix applicability is judged from the proposal and evidence; the fix is not executed as part of this scoring. Unlabeled results Results compiled from two recent benchmark runs cover 4,496 inputs and 40 generated Insights. Every emitted Insight in these selected snapshots was scored with the same judge rubric. Each Insight's overall score is the mean of its eight dimension scores; the table then averages those overall scores within each dataset. We do not combine datasets into one quality score. Dataset Inputs Insights scored Mean judge score (1-5) PUPA 901 5 3.53 FailSafeQA 1,101 11 3.57 Monitoring dashboard agent (internal) 1,342 4 3.88 tau2-bench 1,112 16 2.84 Synthetic scenarios 40 4 3.81 These are per-dataset snapshots, not a controlled comparison of dataset difficulty or pipeline versions. Small Insight counts and grading choices, such as grouped versus individual judging, make these protocol-specific diagnostics rather than stable absolute ratings. Automated judge scores are not precision or recall, human ratings, or measured improvements after applying a fix. The dimension-level scores make the weak spots more concrete. Fix specificity was the lowest-scoring dimension in four of the five snapshots, with means from 2.00 to 2.64. For tau2-bench, correctness was lowest at 2.25: a signal to review whether claims are supported by the supplied evidence. Making a next step more concrete and grounding a diagnosis more carefully are different improvement targets. What the public findings look like The examples below connect generated findings to concrete input evidence: an arithmetic inconsistency, a question/answer mismatch, and a payment allocation that exceeds the available balance. They are selected illustrations, not a representative sample of the 40 scored Insights. Figure 5. Three illustrative findings, with one checked input per finding. Titles and evidence summaries are editorial paraphrases. PUPA is a normalized QA record; FailSafeQA normalization pairs a perturbed question with the original answer, creating the displayed mismatch. Neither is a new agent execution. Tau2-bench shows a recorded tool interaction. These examples are not a representative quality sample. Public input sources: PUPA; FailSafeQA; tau2-bench. Controlled end-to-end tests: Can Insights find known issues? Dataset benchmarks are complemented by tests with healthy baseline agents and versions containing predefined defects. The framework generates traffic, runs Insights on the resulting traces, and compares the findings with the known issues and supporting evidence. It separately tracks detections, misses, unsupported findings, and duplicates. Cases with incomplete evidence remain unscored rather than counting as passes or misses. The September 23 hosted-agent report, used here as a concrete illustration, covered finance, travel, and support-ticket scenarios. Of 12 expected issues, 11 were scorable and 10 were detected. All three scorable healthy baselines had no confirmed unsupported findings. The figure shows the full outcome counts, including the missed and unscored cases. This end-to-end framework is separate from the 40-input synthetic-scenarios dataset in the unlabeled results. Figure 6. Controlled tests exercise healthy baselines and known defects through the Insights pipeline. The September 23 report records 10 detections among 11 scorable expected issues, one additional unscored issue, and no confirmed unsupported findings on three scorable baselines. No confirmed unsupported findings (noise) or duplicate findings were reported in the evaluated subset. This illustrative snapshot is not a representative product-wide rate. A broader view of quality Trace selection and rubric review are part of a broader quality framework. Category agreement is an exploratory benchmark diagnostic, not an assessment of the portal's category labels. Controlled scenarios also check for unsupported or duplicate findings. Figure 7. Five complementary quality questions. The labeled results above measure trace precision and trace recall; the unlabeled results assess findings with an LLM judge. Category agreement is an internal diagnostic, while controlled scenarios check noise and duplication. How we track quality over time Labeled benchmarks, unlabeled assessment, and controlled tests feed recurring quality reports and human investigation. Test inventories and assessment conditions can change, so daily scores are not automatically a comparable improvement trend. The reported results do not establish longitudinal recurrence or deduplication performance across runs. In the portal, Give feedback lets users report incorrect findings, categories, severity, grouping, or duplication. Start with scenarios whose expected failures and healthy controls are known. Track detection, trace precision, categorization, severity, distinctness, evidence grounding, and action quality instead of relying on one score. Review changes to data, models, prompts, and evaluation contracts as changes to the measurement system. Investigate weak results and newly reported failure patterns rather than tuning only for an aggregate. Keep human review in the loop, because benchmark labels and automated graders cannot determine business impact or remediation correctness. How practitioners should interpret an Insight Validate each Insight before acting on it. Microsoft Learn recommends this review sequence: Confirm the affected workflow, agent version, category, severity, and time. Open the highlighted traces and verify that the cited behavior is present. Compare problematic examples with healthy traces to test whether the grouping and likely cause are plausible. Decide whether the issue belongs to the agent, a tool, a model endpoint, a data source, or the platform. Turn confirmed recurring behavior into evaluation coverage, an optimization objective, owner routing, or a monitored no-action decision. An empty result does not prove that an agent is healthy, and a large linked-trace count does not prove business impact. Quality depends on representative traces, complete telemetry, a supported analysis model, and careful human validation. Get started Start with the Insights in Foundry documentation for prerequisites, portal steps, evidence-review guidance, SDK examples, pricing considerations, and preview limitations. Prerequisites include a connected Application Insights resource, recent representative traces, a supported GPT-5-or-newer Judge model deployment, and the required role assignments. Insight generation, including scheduled runs, uses your model deployment and can incur model charges. See the current documentation for the latest supported agents, models, regions, limits, pricing, and UI guidance. Python samples: on-demand Insights and scheduled Insights. Closing thoughts Agent quality work often starts with a simple question: what is repeatedly going wrong that we did not know to test? Insights in Foundry is designed to help teams answer that question with evidence-linked findings and a reviewable next step. Measuring those findings requires explicit definitions, diverse data, honest treatment of weak results, and human judgment where automated metrics stop. Labeled trace evaluation, unlabeled rubric assessment, and controlled end-to-end tests answer complementary questions. Newly observed patterns can then inform the next round of evaluation and review.180Views0likes0CommentsEvaluate More, Spend Less: Batching Microsoft Foundry Evaluators for Efficient Evaluation
Authors: Salma Elshafey, Ali Mahmoudzadeh, Kayla Ames, Ahmad Qardahji, Vivek Bhadauria, Morteza Ziyadi, April Kwong A single agent trajectory may need to be evaluated for groundedness, coherence, instruction following, task completion, and correct tool use. In a conventional pipeline, each evaluator receives the same messages, tool calls, tool results, and tool definitions in a separate model call. As trajectories grow, evaluation repeatedly pays to process the same context. Can one LLM judge call apply five or six evaluators without losing the signal that makes evaluation useful? This study tests a simple design: send the shared context once and ask one composite evaluator to score several criteria in the same call. In summary: On 100-row quality and tool use samples, composite evaluation required 5-6x fewer model calls, reduced total input-token use by 61.25–71.07%, completion tokens by 46.63–63.89%, and measured run wall time by 35.74–46.10%. Across the tested workloads, frontier judge models maintained stable quality across several high-value dimensions. The results support composite evaluation as a cost-efficient default, with targeted individual evaluation for workload-sensitive rubric criteria. Why Evaluation Becomes Expensive Evaluating a single conversation rarely involves just one evaluation dimension. A typical assessment examines multiple dimensions of quality, such as whether the response is coherent, follows instructions, completes the user's task, remains grounded in available evidence, and uses tools correctly. Because each dimension is typically evaluated through a separate prompt and model call, the same conversation trajectory, tool calls, tool outputs, and tool definitions are processed repeatedly. As a result, cost grows with both the number of evaluation criteria and the amount of shared context, making long agent trajectories especially expensive to evaluate. The Composite-Evaluator Approach Rather than evaluating each dimension independently, composite evaluation scores multiple criteria within a single judge-model call, allowing the shared context to be processed once instead of repeatedly. We built two composite evaluators. The Output Quality Evaluator The Output Quality Evaluator scores six Microsoft Foundry evaluators in one LLM call: Evaluator Question Fluency Is the response clear and well formed? Coherence Is it logically consistent with the conversation? Intent Resolution Did the assistant understand the user's goal? Task Adherence Did it follow the user's instructions? Groundedness Are its claims supported by the available evidence? Task Completion Did it complete the requested task? Tool Use Quality Evaluator The Tool Use Quality Evaluator scores five Microsoft Foundry evaluators in one LLM call: Evaluator Question Tool Call Accuracy Was the tool call correct overall? Tool Call Success Did the invocation succeed? Tool Input Accuracy Were the arguments correct? Tool Output Utilization Did the response use the tool output correctly? Tool Selection Was the appropriate tool selected? Each composite prompt includes the shared trajectory and tool definitions once, followed by the full definition, rating scale, and applicability rules for every evaluator. It instructs the judge to assess each evaluator independently so that one verdict does not bias another. The judge returns structured results for each criterion, including a score, rationale, and applicability status. For multi-turn evaluations, it can also identify the earliest turn where a failure occurred. How We Validated Evaluator Quality The study compares individual and composite evaluators across: Two modes: Single-Turn and Multi-Turn. Nine judge models: GPT-4o, GPT-5.4, GPT-5.4 mini, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, DeepSeek V4 Flash, DeepSeek V4 Pro, and Grok 4.1 Fast Reasoning. Multiple datasets, as shown in the table below Dataset Mode Rows Primary validation target Reference labels Internal quality set Single-Turn and Multi-Turn 283 Six quality evaluators Existing per-evaluator labels Internal tool use set Single-Turn and Multi-Turn 200 Five tool use evaluators Existing per-evaluator labels BFCL v4 Multi-Turn 200 Task Completion, Tool Call Accuracy, and failed-turn localization Deterministic state and tool-call checks AgentIF Multi-Turn 148 Task Adherence Constraint-level majority-vote reference FaithDial Multi-Turn 300 Groundedness Faithful versus hallucinated responses FED Multi-Turn 125 Coherence Five-annotator human scores Tau-Voice Multi-Turn 278 Task Completion Reward-derived completion labels AgentRx tau_retail Multi-Turn 29 Failed-turn localization Human failure annotations We measured several distinct properties: Input tokens, output tokens, model calls, and latency Accuracy, macro-F1, and Cohen's kappa where labels were available Agreement and correlation structure between individual and composite modes Repeatability across four evaluation runs Finding 1: Composites Cut Input Tokens by 61.25–71.07% and Latency by 35.74–46.10% We measured cost and latency on matched 100-row Single-Turn samples: one for Output Quality and one for Tool Use, comparing parallel individual evaluators with one composite evaluator per row. For each row, evaluator context includes the query and response, any tool calls and results, and the serialized tool definitions. Note: Single-turn means that the evaluator's focus is the last agent response only, but the entire conversation history is passed to the evaluator as context. Trajectory Length in the Cost Sample The log-scale distributions show that most rows clustered around a few thousand tokens, with a smaller number of substantially longer trajectories. In both samples, longer evaluator contexts produced larger absolute token savings. The Measured Cost Reduction Across the sample: Repeated context is the main saving. Total input-token volume fell from 2,058,930 to 797,843 for Output Quality, a 61.25% reduction, and from 2,036,122 to 589,097 for Tool Use, a 71.07% reduction. Uncached input fell by 57.51% and 53.29%, respectively. For Output Quality, 52.72% of composite input tokens were cached versus 56.88% across the individual suite; for Tool Use, the figures were 46.22% versus 66.69%. However, the composites still processed substantially fewer tokens overall. Call volume collapses. Output Quality fell from 600 individual evaluator calls to 100 composite calls. Tool Use fell from 494 calls to 100. In six rubric-row combinations, native applicability rules legitimately returned no score, such as when there was no tool call to evaluate; those skips explain why the individual Tool Use baseline is below 500. Measured runs finish sooner. Output Quality wall time fell from 504.41 seconds to 324.12 seconds, a 35.74% reduction, while Tool Use fell from 536.95 seconds to 289.42 seconds, a 46.10% reduction. Per sample row, that is 5.044 to 3.241 seconds for Output Quality and 5.369 to 2.894 seconds for Tool Use. Completion tokens fell by 46.63% and 63.89%, respectively. On the quality sample, the composite used 797,843 total input tokens and 36,919 completion tokens, compared with 2,058,930 and 69,180 for parallel individual evaluation. On the tool use sample, the composite used 589,097 total input tokens and 28,358 completion tokens, compared with 2,036,122 and 78,522. These are measured run totals, not prompt-length estimates. The broader experiments showed the same benefit. The full six-evaluator rubric Single-Turn study used about 68% fewer input tokens, while the 283-row Multi-Turn quality study used about 77% fewer: 28,715 rendered input tokens per row for a six-call equivalent versus 6,608 for the composite call. Exact currency savings depend on model and deployment pricing, but batching consistently removed duplicated input and round-trip overhead. Finding 2: Quality Tradeoffs Depend on Rubric and Judge Model Output Quality Rubric Benchmark Results On the internal Single-Turn quality set, composite Task Adherence accuracy improved for every judge, while Task Completion remained close to individual evaluation. External Multi-Turn benchmarks showed a more varied pattern across rubric, judge, and workload. This figure reports Δκ = κComposite - κIndividual: blue favors composite evaluation and orange favors individual evaluation. The mixed directions reinforce the need to validate the selected rubric and judge on production-like data. BFCL provided the strongest labeled Multi-Turn example for Task Completion. With GPT-5.6 Luna, the composite evaluator reached macro-F1 0.848 and κ = 0.696, compared with 0.836 and 0.672 for the individual evaluator. Across the other judges, composite and individual Task Completion remained close for Sol, Terra, and DeepSeek V4 Flash, while results for DeepSeek V4 Pro and Grok 4.1 Fast favored individual evaluation. Tool Use Quality Rubric Benchmark Results The five-rubric tool composite was evaluated on production-style tool traces. Each row carries ground truth for one source rubric, so the table reports accuracy on the available per-evaluator subsets. Judge Tool Call Accuracy Tool Call Success Tool Input Accuracy Tool Output Utilization Tool Selection gpt-4o 0.800 0.902 0.850 0.800 0.878 gpt-5.4 0.800 0.902 0.775 0.737 0.829 gpt-5.4-mini 0.750 0.902 0.725 0.763 0.756 gpt-5.6-luna 0.914 0.902 0.848 0.775 0.901 gpt-5.6-sol 0.886 0.902 0.750 0.743 0.854 gpt-5.6-terra 0.857 0.902 0.750 0.794 0.854 DeepSeek-V4-Flash 0.775 0.902 0.800 0.848 0.854 DeepSeek-V4-Pro 0.800 0.878 0.800 0.789 0.902 grok-4-1-fast-reasoning 0.946 0.902 0.925 0.745 0.823 Across the production-style corpus, the composite reached 0.75-0.95 accuracy across the five rubric criteria. The results show model-specific strengths: Grok 4.1 Fast led Tool Call Accuracy and Tool Input Accuracy, DeepSeek V4 Flash led Tool Output Utilization, and DeepSeek V4 Pro led Tool Selection. Tool Call Success remained tied at 0.902 for most judges. Luna remained among the strongest overall, but no single judge dominated each of the rubric criteria. Classification Accuracy On the truly multi-turn BFCL benchmark, Tool Call Accuracy is the only tool rubric criterion with native ground truth. GPT-5.6 Terra led the single-run panel at 0.864 macro-F1, followed by GPT-5.6 Sol at 0.858 and GPT-5.6 Luna at 0.828. Other models showed lower results, with Grok 4.1 Fast at 0.626, DeepSeek V4 Pro at 0.625, and V4 Flash at 0.521. This shows that with the correct model, the composite evaluator can classify overall tool-call correctness on full conversations as well as production-style traces. Failure Localization The same BFCL run also tested where failures occurred. GPT-5.6 Terra led at 0.78 member-any failed-turn localization: in 78% of failing conversations, the turn predicted by the Tool Call Accuracy rubric criterion matched at least one turn in BFCL's set of failing turns. It produced 22 false alarms across 75 passing conversations. GPT-5.6 Sol reached 0.75 member-any, and GPT-5.6 Luna reached 0.74. The Tool Selection channel with Luna provided a high-precision secondary signal, identifying a failing turn in 50% of failures with only 3 false alarms. To test whether localization generalizes beyond mechanically checked BFCL traces, we also used the 29-row AgentRx tau_retail set with human failure annotations. On the eight in-scope tool-mechanics failures, Sol localized 8 of 8 root causes, while Terra and Luna localized 7 of 8. Meanwhile, Grok 4.1 Fast only localized 2/8, and the two DeepSeek models 1/8. This result is directional because of the small sample. Does Batching Preserve Rubric Relationships? After testing direct accuracy, we examined whether batching also preserves how rubric verdicts relate to one another. This is supporting structural evidence, not an accuracy measure against ground truth. The heatmaps compare Spearman correlations between GPT-5.6 Luna's binarized rubric verdicts on rows scored by both modes. The tool use structure was especially stable, including the strongest relationship between Tool Call Accuracy and Tool Input Accuracy (0.76 in both modes). Quality relationships were more mixed. For example, Single-Turn Intent Resolution and Task Completion fell from 0.71 to 0.31, while Multi-Turn Task Adherence and Task Completion fell from 0.53 to 0.23. This shows that batching can preserve the broad rubric structure without preserving every relationship equally. Finding 3: Reliability Depends on Both Rubric and Judge Multi-Turn Quality Reliability We repeated the six-evaluator composite call four times on the same 200 BFCL rows. Across the nine-judge panel, DeepSeek V4 Flash had the lowest average flip rate at 3.9%. GPT-5.6 Sol and Grok 4.1 Fast followed at 5.2%, DeepSeek V4 Pro at 5.6%, GPT-5.6 Luna at 6.8%, GPT-4o at 7.4%, GPT-5.6 Terra at 7.8%, GPT-5.4 mini at 10.9%, and GPT-5.4 at 15.6%. The rubric criteria-level result is more important than the leaderboard: Fluency: 0.0% average flip rate Coherence: 2.4% Intent Resolution: 8.6% Task Adherence: 6.7% Task Completion: 11.5% Groundedness: 16.4% The same judge can be highly stable on one rubric criterion and noisy on another. Reliability should therefore be monitored at the criterion level, not inferred from a single aggregate score. Multi-Turn Tool Use Reliability The five-evaluator MT tool study showed the same need for criterion-level monitoring. Across four BFCL repeats, GPT-5.6 Terra and DeepSeek V4 Flash had the lowest average flip rate at 6.0%, GPT-5.6 Sol reached 6.1%, and GPT-5.6 Luna and GPT-4o were at 7.0%. GPT-5.6 Sol achieved the strongest mode-of-four Tool Call Accuracy kappa at 0.759, while Terra led the single-run quality/localization results. Tool Call Success was the most stable criterion, with a 4.7% average flip rate across judges. Practical Recommendations Based on the results across quality, tool-use, localization, and reliability studies, we can derive practical guidance for both judge selection and evaluator batching. Recommended Judges: The GPT-5.6 Family Across the quality, tool-use, localization, and reliability experiments, the GPT-5.6 family consistently delivered the strongest overall results. While individual benchmarks occasionally favored other models or specific GPT-5.6 variants, Luna, Sol, and Terra repeatedly ranked among the top performers and showed the most balanced performance across workloads. Sol was generally the most reliable across repeated quality and tool-use studies, Terra led several BFCL tool-use and failed-turn localization evaluations, and Luna achieved the highest average quality across the tested workloads. For workloads similar to those evaluated here, we recommend starting with GPT-5.6 judges and selecting between Sol, Terra, and Luna based on your specific priorities. When available at a lower deployment cost, Luna is a particularly attractive option because it maintained frontier-level evaluation quality while achieving the strongest average performance in our experiments. By comparison, lower-cost models often involved larger tradeoffs between quality, reliability, and benchmark performance. Exact cost savings depend on model pricing and deployment configuration, so benchmark candidate judges on production-like workloads before optimizing solely for cost. Which Evaluators to Batch These evaluators performed as well in the composite evaluator as they did individually — they're strong candidates for batching right away: Fluency and Coherence — highly stable across all nine judges and repeats Tool Call Success and Tool Selection — robust across the tested judge models Task Completion — close to individual evaluation across benchmarks; slightly stronger with GPT-5.6 Luna on BFCL Tool Call Accuracy, Tool Input Accuracy, and Tool Output Utilization — strong on production-style traces, with accuracy ranging from 0.75 to 0.95 depending on evaluator and judge Which Evaluators to Watch These evaluators showed workload- or judge-dependent results. They can still be batched, but we recommend checking the results on your own data: Task Adherence — performed well on typical conversations but struggled when a single request contained many independent constraints Intent Resolution — results varied across datasets even though agreement was strong in some cases Groundedness — the least consistent evaluator in our tests, with scores changing across repeated runs more often than any other dimension, albeit having better performance on frontier models, like GPT-5.6 Luna and GPT-5.4. If grounding accuracy is critical for your use case, consider keeping the individual evaluator as a fallback Where This Approach Fits Composite evaluation is most effective as a selective default rather than a universal replacement for every standalone evaluator. GPT-5.6 Sol offered the strongest overall balance, Terra led the single-run BFCL tool results, and Luna is the recommended cheaper high-quality option. Judge and rubric selection should therefore be validated on production-like data. The Multi-Turn comparison has one structural boundary: native individual evaluators do not exist for Multi-Turn Fluency and Intent Resolution. The composite evaluator can produce both dimensions, but a direct individual-versus-composite validation is not available for them in that setting. Finally, token savings do not translate to one fixed currency percentage. Pricing varies by model and deployment, while output length and retry behavior vary by workload. Takeaway Composite evaluators are a practical way to remove repeated context from evaluation pipelines. In the updated 100-row comparisons, they reduced total input-token use by 61.25% for six quality evaluators and 71.07% for five tool use evaluators, reducing completion tokens by 46.63–63.89%, while reducing measured run wall time by 35.74% and 46.10%, respectively, and preserving comparable quality across many dimensions. The strongest deployment pattern is selective rather than absolute: start with a model from the GPT-5.6 family; batch the evaluators that remain stable, monitor them independently, and keep focused fallbacks for the evaluators that do not. Composite evaluation is ultimately about removing unnecessary work. By evaluating multiple dimensions in a single judge call, teams can dramatically reduce repeated context processing while maintaining the evaluation coverage needed to monitor production systems. Get Started Start with the composite evaluator on a small set of known-good and known-bad agent conversations. Choose Output Quality to assess the six quality criteria or Tool Use Quality to assess the five tool use criteria. Review each criterion’s score and applicability status, then spot-check the results against individual evaluators for the failure modes that matter most to your application. The GPT-5.6 family is the recommended starting point for judge models. Re-check results when you change the judge model or the domains of conversations being evaluated. Try these evaluators in the Microsoft Foundry portal now. You can explore the docs on how to use them Agent Evaluators for Generative AI - Microsoft Foundry | Microsoft Learn.98Views0likes0CommentsYour Agent Shouldn't Wait for a Prompt: Build an Event-Driven Microsoft Foundry Routine
Most agents are still waiting in a chat window. They may be capable of classifying an incident, finding the right documentation, or recommending an owner - but nothing happens until someone remembers to ask. The hard part is no longer always the reasoning. It is noticing that work has arrived, invoking the agent securely, and keeping enough history to understand what happened. Routines in Microsoft Foundry close that gap. A routine connects a trigger - such as a schedule, timer, GitHub issue, or Microsoft Teams message - to an agent action. Microsoft Foundry queues the invocation, runs the agent, and stores a run record for later inspection. In this post, we will build an event-driven routine for a familiar developer workflow: When someone opens a GitHub issue, invoke a triage agent immediately - without waiting for a person to copy the issue into a chat. Along the way, we will look at the architecture, create the routine with the Azure Developer CLI (azd), test it, inspect its run history, and make an explicit identity decision before putting it into production. From conversational agents to event-driven agents A chat-first agent follows a request-response pattern: A user opens an interface. The user provides a prompt. The agent performs work. The interaction ends or waits for another prompt. An event-driven agent starts differently. Work in an external system becomes the prompt. A newly opened issue, for example, already contains useful context: its title, description, author, repository, labels, and timestamps. A routine can receive that event through an authorized connection and invoke an agent while the context is still fresh. That changes the agent's role. It is no longer only a place people go for answers. It becomes a participant in an operational workflow. The routine does not replace the agent. It supplies the managed automation around it: Trigger: Defines when work starts. Connection: Authenticates the event source. Action: Identifies the agent and invocation protocol. Dispatch identity: Determines whose permissions are used when the agent and its tools run. Run history: Records executions so operators can inspect outcomes. This separation is useful. The agent owns the reasoning; the routine owns when and how that reasoning begins. The scenario: triage every new GitHub issue Our example assumes that a triage agent is already deployed in a Microsoft Foundry project. When an issue opens, the agent should: Summarize the issue in two or three sentences. Classify it as a bug, feature request, documentation issue, or support question. Estimate severity and explain the evidence. Recommend an owner or team. Identify missing reproduction details. Produce a proposed response for a maintainer to review. Keeping a human review step is intentional. Event-driven does not have to mean unrestricted autonomy. A routine can automate the expensive first pass while a maintainer remains responsible for labels, assignments, and public responses. Prerequisites You need: An active Microsoft Foundry project. The Foundry User role or higher on the project. A deployed prompt agent or hosted agent. Workflow agents are not currently supported by routines. The Azure Developer CLI. A GitHub connection authorized in the Foundry project. The routines extension for azd. Install the extension and confirm that the routine commands are available: azd extension install azure.ai.routines azd ai routine --help Set the project endpoint through your active azd environment or pass it explicitly to each command: azd env set AZURE_AI_PROJECT_ENDPOINT \ "https://<account>.services.ai.azure.com/api/projects/<project>" The agent must exist before you attach a routine to it. A routine references an agent; it does not deploy one. Create the GitHub issue routine The following command creates an enabled routine named triage-on-open. Replace the placeholders with the connection and repository details from your environment: azd ai routine create triage-on-open \ --trigger github-issue \ --connection-id "<workspace-connection-id>" \ --owner "<github-owner>" \ --repository "<github-repository>" \ --issue-event opened \ --action agent-invoke \ --agent-name "triage-agent" \ --description "Triage every newly opened GitHub issue" There are two important type translations in this command: The CLI alias github-issue represents the routine trigger type github_issue. The CLI alias agent-invoke invokes the agent through the Invocations API. If your agent uses the Responses API instead, use --action agent-response. Choose the protocol that matches the deployed agent rather than treating the two actions as interchangeable. For automation that belongs in source control, define the routine as an azure.ai.routine service in azure.yaml. This makes the relationship between the agent and its trigger reproducible across environments: services: triage-agent: host: azure.ai.agent project: ./agent triage-on-open: host: azure.ai.routine uses: - triage-agent description: Triage every newly opened GitHub issue enabled: true triggers: issue-opened: type: github_issue connection_id: ${GITHUB_CONNECTION_ID} owner: ${GITHUB_OWNER} repository: ${GITHUB_REPOSITORY} issue_event: opened action: type: invoke_agent_invocations_api agent_name: triage-agent Deploy the routine after the target agent: azd deploy triage-on-open --no-prompt The uses relationship tells azd to order the agent before the routine. Deployment is idempotent: redeploying updates the named routine instead of creating duplicates. Give the agent a clear triage contract Automation magnifies ambiguity. A vague instruction that is merely inconvenient in a chat can produce inconsistent work every time an event fires. Give the triage agent a bounded contract such as: You are the first-pass triage agent for this repository. For each newly opened issue: 1. Summarize the reported behavior without adding facts. 2. Classify it as bug, feature, documentation, or support. 3. Assign severity only when the issue contains supporting evidence. 4. List missing information needed to reproduce or route the issue. 5. Recommend an owner from the approved ownership map. 6. Draft a response, but do not publish, close, label, or assign the issue. Return structured JSON that matches the triage schema. A useful output contract might look like this: { "summary": "The CLI exits when a project endpoint contains an explicit port.", "category": "bug", "severity": { "level": "medium", "reason": "The issue blocks routine creation but has a documented workaround." }, "missing_information": [ "Azure Developer CLI version", "Redacted project endpoint shape", "Full error output" ], "recommended_owner": "developer-experience", "proposed_response": "Thanks for the report. Could you share..." } Structured output gives downstream systems something predictable to validate. It also makes evaluation easier: you can test category accuracy, required-field completeness, unsupported severity claims, and whether the agent attempted a prohibited action. Test before waiting for a real event Start by checking that Foundry stored the routine you intended: azd ai routine show triage-on-open --output json Confirm: The routine is enabled. The trigger watches the correct owner and repository. issue_event is opened. The action references the intended agent. The connection ID belongs to the expected Foundry project. You can manually dispatch a routine while testing: azd ai routine dispatch triage-on-open \ --input '{"test":true,"issue":{"number":123,"title":"Test triage event"}}' The manual input is a one-time override for that dispatch. It does not replace the event payload or modify the routine's stored configuration. Inspect recent executions: azd ai routine run list triage-on-open --top 20 Then open a test issue in the watched repository and inspect the run list again. A production test should verify more than "the agent ran." Check that the correct event started the run, the agent received enough context, tool calls used the intended identity, the output matched the schema, and prohibited actions did not occur. Choose the dispatch identity deliberately Every routine uses the agent identity by default. This is usually the better fit for unattended automation because access belongs to the agent rather than to an employee's account. Use agent identity when the agent's tools authenticate with managed identity, workload identity, or keys and the agent has been granted only the permissions required for the task. Some tools require delegated user access. In that case, you can create the routine with creator identity. Creator identity means the Microsoft Entra identity of the person or service principal that creates the routine—not the agent publisher, connection creator, latest editor, or user who caused an event. This distinction has operational consequences: If the creator loses access or consent, delegated tool calls can fail. Recreating the routine as another principal changes the delegated creator identity. The event connection identity is separate from the identity used to dispatch the agent. Dispatch identity is a creation-time decision. To switch an existing routine between agent and creator identity, delete and recreate it. For a GitHub triage flow, a strong starting design is: Use a narrowly scoped project connection to receive issue events. Use agent identity for Foundry and Azure resources. Keep repository-changing actions disabled until evaluation demonstrates reliable behavior. Require human approval before posting, assigning, labeling, or closing. Identity is not a deployment detail. It is part of the automation's behavior and should be reviewed with the same care as the prompt and tool list. Make failure visible An event-driven agent can fail even when its reasoning is sound. The connection might expire. The agent might receive an unexpected payload. A tool might lose permission. The output might violate its schema. Monitor the workflow at four boundaries: Trigger: Did the expected event fire the routine exactly once? Invocation: Did Foundry invoke the intended agent and protocol? Tools: Did tool calls succeed with the intended identity and scope? Outcome: Did the output satisfy the triage contract? Use azd ai routine run list for routine execution history and your agent's Foundry observability data for traces, tool calls, latency, and failures. Preserve representative failures as evaluation cases instead of fixing each incident only in the prompt. Also design for duplicate delivery. Before taking a repository-changing action, check whether the issue and event have already been processed. Idempotency matters more once an agent can act without a person initiating each run. Know the current boundaries Before adopting routines for a regulated or business-critical workload, review the current service constraints: Routines support prompt agents and hosted agents, but not workflow agents. A recurring schedule has a minimum interval of five minutes. GitHub issue triggers support opened and closed issue events. Routines inherit the project's networking configuration and can work with virtual-network-secured projects. Routines do not currently support customer-managed key encryption. Regional availability has exceptions; verify your project's region in the current documentation. These boundaries can change. Treat the routines documentation as the source of truth when moving from a tutorial to production. What changes when agents stop waiting? The most interesting part of this design is not the GitHub trigger. It is the change in operating model. The agent begins work because the world changed, not because someone opened a chat. That makes trigger scope, identity, output contracts, run history, evaluation, and human approval part of the agent design—not infrastructure to consider later. Once the triage routine is working, the same pattern can support: A Teams message that starts support classification. A nightly backlog review. A one-time release-readiness check. A scheduled compliance summary. A hosted agent that uses the reminder tool to resume the same conversation after a long-running task. Start with one bounded event and one reversible outcome. Measure what the agent does, not merely whether it ran. Then expand its permissions only as evidence earns that autonomy. Try it next Choose the path that matches where you are: Build: Follow the Microsoft Foundry routines documentation and connect one existing agent to a schedule or event. Harden: Review dispatch identity, connection scope, idempotency, output validation, and human approval before enabling repository-changing tools. Extend: Add the reminder tool to a hosted agent that needs to continue work later. Explore: Open the Microsoft Foundry portal to inspect your project, agents, and routine runs. Your agent already knows how to do useful work. The next step is teaching it when that work should begin.170Views1like1CommentDeploying hosted agents in Foundry Agent Service via Terraform
Background Your Terraform workflow already manages your Azure infrastructure but deploying hosted agents still requires manual SDK calls or REST API scripts. This post shows you how to bring agent deployments into your existing IaC pipeline using the AzAPI provider, so your entire Foundry stack can be versioned, reviewed, and deployed together. For this post we will focus on leveraging Foundry Hosted Agents. At a high level hosted agents will abstract the overhead of managing your agents compute. Conceptually, think of taking the compute today, that might be in Azure Container Apps (ACA) or Azure Functions and moving it into Foundry. For more information can check out my previous blog on this topic. In this example Foundry runs the agent code on managed compute from a container image. For this specific example we need to use Azure Container Registry to house our code artifact. If wanting to deploy directly from source refer to my previous blog Deploying Foundry Hosted Agents from Source Why an Infrastructure as Code Approach? Hosted agents today in Foundry support SDK and REST API deployments, why should we look at leveraging an IaC provider to handle our agents today? Many organizations subscribe to an IaC strategy for managing their Azure Resources. A hosted agent, at its core, is configuring and allocating infrastructure resources behind Foundry. This would be similar to how we configure plan sizes for products like App Services. Additionally, anything defined by IaC makes it easier to version the configuration in source control and incorporate it into repeatable CI/CD pipelines. Production workflows should also account for remote state, approval controls, and image-version promotion. Another added benefit for anything under IaC is the ability to apply custom policies, Azure or otherwise, over the codebase. Prerequisites An Azure subscription with permission to create resources and role assignments. Access to a region and model deployment with sufficient quota for Foundry Hosted Agents. Azure CLI, Terraform, Git, and Docker installed. Docker must support building linux/amd64 images. Permission to push images to Azure Container Registry and create agents in the Foundry project. The provided sample creates or configures the supporting Azure Container Registry, Foundry account and project, model deployment, managed identity, project connection, and hosted-agent deployment. Hosted agents require a Linux AMD64 (linux/amd64) container image. When building from an ARM-based workstation, including Apple Silicon, explicitly target linux/amd64 and ensure Docker cross-platform emulation is available. See Microsoft’s hosted-agent container requirements. docker build --platform linux/amd64 -t <registry-name>.azurecr.io/<image-name>:<tag> To get started quickly I have a repository w/ all the prerequisites and reviewed code at simple-hosted-agent-deploy-azapi Deploy the Sample To deploy the sample: Clone the repository and reopen it in the included development container. Authenticate to Azure and select the target subscription. Copy and update the example Terraform variable files with your environment-specific values. Run the included deployment script to provision the base resources, build and push the container image, and create the hosted agent. Confirm that the generated agent version reaches an active state before invoking it. Refer to the repository README for the current commands, configuration values, and cleanup steps. Process Overview Let’s level set on what our end-to-end process may look like. In many large organizations, the agent deployment process could be decoupled from the Foundry base architecture. Components such as the Foundry account, project, connections, and model may be controlled by a centralized team. These shared components often follow a different deployment lifecycle from an individual agent. A development team may own the code. We need the code to deploy the hosted agent. Quite the chicken-and-egg problem. So let’s break it down from a day 1 perspective: Deploy Foundry Base Architecture Foundry Account Foundry Project Model Project Connection Application Insights Build and Push code to Azure Container Registry Deploy the hosted agent pointed to the image defined in Azure Container Registry For this example, I assume that Azure Container Registry is managed outside the Foundry lifecycle because many organizations prefer to consolidate container images into a smaller number of shared registries. For illustration purposes, the provided example includes the registry deployment. Day n perspective would potentially look like: Build and Push code to Azure Container Registry Deploy the hosted agent pointed to the image defined in Azure Container Registry Updating the hosted-agent configuration creates a new agent version under the existing logical agent. Foundry manages the version history, while Terraform manages the logical agent configuration represented by the azapi_data_plane_resource. After each deployment, confirm that the new version reaches an active state before directing workloads to it. Technical Details At this time, let’s take a minute and discuss some of the technical details around how the hosted agent is created in Foundry. The Foundry project and its connections are Azure Resource Manager control-plane resources. The logical agent is created through the Foundry project’s data-plane API, while Foundry manages the supporting deployment resources required to host it. This is in contrast to services represented directly as Azure Resource Manager resources, such as App Service (Microsoft.Web/sites) and Azure Container Apps (Microsoft.App/containerApps). Here is a visual that depicts the data plane vs control plane objects: The Foundry project is the agent’s logical parent and provides the data-plane endpoint through which the agent is created. The project itself is deployed as the Azure Resource Manager resource Microsoft.CognitiveServices/accounts/projects Foundry connections can be deployed as Microsoft.CognitiveServices/accounts/projects/connections. In this case, the connection is used for connecting and authenticating to Azure Container Registry. At this point, the project and connection resources can be deployed through AzAPI or Bicep. The next step requires a data-plane call. For Terraform deployments, this operation can be managed declaratively with the AzAPI provider’s azapi_data_plane_resource. Foundry also supports agent deployment through its SDKs and REST API. azapi_data_plane_resource The AzAPI provider’s azapi_data_plane_resource resource is the key component for deploying a hosted agent in Foundry Agent Service through Terraform. Per official MS Learn documentation: Some Azure services expose a separate data plane API—a service-specific HTTPS endpoint where you interact directly with the service rather than through ARM. Examples include the Key Vault secrets API at {vaultName}.vault.azure.net, the Azure AI Search index API at {searchServiceName}.search.windows.net, and the Synapse workspace pipeline API at {workspaceName}.dev.azuresynapse.net. azapi_data_plane_resource bridges this gap by enabling Terraform to manage resources on these data plane endpoints using the same AzAPI provider authentication and lifecycle model. https://learn.microsoft.com/en-us/azure/developer/terraform/concept-azapi-data-plane-framework So how would a terraform implementation for this work? First let’s create the appropriate resource type, name and parent reference, in this case the Foundry Project: resource "azapi_data_plane_resource" "hosted_agent" { type = "Microsoft.Foundry/agents@v1" name = var.agent_name # AzAPI data-plane parents use the endpoint host/path without a URI scheme. parent_id = trimprefix(var.project_endpoint, "https://") …} Now we need to pass in the dedicated parameters as part of the body. To find these parameters refer to the Foundry REST API documentation. body = { name = var.agent_name definition = { kind = "hosted" container_configuration = { image = var.image_uri } cpu = var.cpu memory = var.memory protocol_versions = [ { # The hosted container must implement this Foundry Responses contract. protocol = "responses" version = "2.0.0" } ] environment_variables = merge(var.environment_variables, { AZURE_AI_MODEL_DEPLOYMENT_NAME = var.model_deployment_name }) rai_config = { # Hosted agents require the full policy ARM ID; model deployments use its name. rai_policy_name = var.rai_policy_id } } } For the entire reusable module check out https://github.com/JFolberth/simple-hosted-agent-deploy-azapi/tree/main/simple_agent/azure/infra/modules/hosted_agent Conclusion Hosted Agents in Foundry Agent Service provides a way to run custom agent code on Foundry-managed compute while reducing the operational overhead of managing the underlying hosting infrastructure. By using the AzAPI provider’s azapi_data_plane_resource, teams can incorporate the logical agent deployment into an existing Terraform workflow alongside the Foundry project, model, connections, identity, and container registry resources it depends on. This approach is especially useful when platform and application responsibilities are separated. A platform team can manage the shared Foundry and Azure infrastructure, while application teams independently build, publish, and deploy new versions of their agent container. With the appropriate remote state, access controls, image-versioning strategy, and deployment validation in place, the same pattern can be extended into a repeatable CI/CD workflow for both initial deployment and ongoing agent updates. The linked sample demonstrates this end-to-end pattern, from provisioning the supporting Azure resources and publishing the container image through creating the hosted agent with Terraform. Use it as a starting point, then adapt its state management, access controls, artifact promotion, and validation steps to meet your organization’s production requirements.241Views0likes0CommentsEngineering Agentic Recall Controls with MCP and Microsoft Foundry
AI agents become operationally interesting when they can reach real systems. They also become operationally dangerous at exactly the same moment. Caldova Recall Control Tower is a developer demonstration built around that tension. It uses a fictional pharmaceutical recall to show how an agent can gather evidence and prepare a decision while deterministic application code retains authority over approval and inventory mutation. The implementation combines the Model Context Protocol (MCP), Microsoft Agent Framework, a Microsoft Foundry Hosted Agent, FastAPI, Microsoft Entra authentication, managed identity, and optimistic concurrency in Azure Blob Storage. Central design rule: Let the model interpret and recommend. Make ordinary code authenticate, authorize, mutate, and prove what happened. Caldova is fictional, all operational data is synthetic, and this sample is not a production recall system or a source of clinical advice. The scenario: useful reasoning, consequential action The demo starts with a temperature excursion affecting batch B-2408-AX7 of Caldova Relief 20 mg tablets. The synthetic inventory contains 2,196 units across two distribution centers and two retail stores. A useful system must establish the notice, locate every affected position, check supplier status, explain uncertainty, and recommend an action. That analysis is a good fit for specialized agents. Quarantining inventory is not. Quarantine changes operational state. It therefore needs an authenticated human, explicit authorization, a batch-scoped approval, concurrency control, idempotency, and an audit record. None of those guarantees should depend on a prompt being followed. The authenticated hosted application at the start of the fictional recall. Architecture: separate reasoning from authority The solution has two related but deliberately separate paths. Caldova separates model reasoning from application authority. Official Microsoft service icons identify Foundry Agent Service, App Service, Managed Identity, and Blob Storage. The reasoning path invokes a Hosted Agent through the Responses protocol. Four agents run in a fixed sequence: triage, inventory impact, supplier/compliance, and supervisor. The first three have narrow read-only tools. The supervisor has no tools and synthesizes the accumulated context into a decision brief. The authority path remains in the web application. It validates the EasyAuth identity claims, checks an approver allowlist, binds approval to the caller, batch, action, and current demo generation, and only then calls deterministic domain code. State is stored per actor in Blob Storage and updated with ETag match conditions so concurrent writes fail instead of silently overwriting each other. There is also a deterministic localhost demo. It reuses synthetic domain fixtures and demonstrates MCP contracts, approval, quarantine, and replay without a model or cloud account. It is useful for development, but its typed approver name and in-memory state are not production identity or durable compliance evidence. Building a narrow MCP surface MCP standardizes how an AI application discovers and calls external tools. It does not remove the need to design those tools carefully. Caldova exposes small, typed operations such as get_recall_notice , locate_inventory , and get_supplier_status . Inputs are constrained with Pydantic, and tool annotations tell clients that these operations are read-only and closed-world: @mcp.tool( title="Locate affected inventory", annotations=ToolAnnotations( read_only_hint=True, open_world_hint=False, ), ) def locate_inventory(batch_id: BatchId) -> dict[str, Any]: return STORE.locate_inventory(batch_id) The mutation tool is separately marked destructive and requires a batch-scoped approval token. Those annotations improve discovery and planning, but they are metadata, not an authorization boundary. The real check occurs inside quarantine_batch , below the model and below the tool description. The Hosted Agent does not receive the mutation tools at all. Its specialists call only three read operations through an isolated MCP stdio subprocess. This is stronger than asking an all-powerful agent to "please remain read-only": capability is constrained by construction. The subprocess boundary also keeps MCP v2 dependencies isolated from the Foundry hosting environment. Each call has a timeout, bounded concurrency, structured JSON handling, and a generic failure response that does not leak subprocess details. Fixed workflows beat vague autonomy for this case Multi-agent does not have to mean dynamic routing. Caldova uses an explicit sequence because the business dependency is explicit: validate the notice before locating inventory, locate inventory before checking supplier implications, and synthesize only after all three specialist outputs exist. return ( WorkflowBuilder(start_executor=triage, output_from=[supervisor]) .add_edge(triage, inventory) .add_edge(inventory, compliance) .add_edge(compliance, supervisor) .build() .as_agent() ) This topology is easier to test and reason about than an unconstrained planner. Each specialist has one job and one tool allowlist. Full context is passed where synthesis requires it, while the supervisor remains tool-free. Request isolation matters too. A hosted process can serve concurrent users, so workflow state must not leak between requests. The sample creates a fresh workflow agent for each request context rather than reusing mutable agent state globally. What the live hosted run showed The hosted application completed a read-only assessment for the synthetic batch. The resulting brief reported: 2,196 affected units across four locations. A high-risk inbound temperature excursion. Supplier acknowledgement, a 36-hour replacement estimate, and a drafted credit note. Unknown transit temperature details, excursion duration, stability impact, final supplier disposition, and potentially issued stock. A recommendation to hold or quarantine stock, explicitly stating that no quarantine had occurred. A required human approval before any inventory restriction. The live Hosted Agent decision brief. Transient response and correlation identifiers are masked; the operations rail is excluded because it contains actor-scoped audit data. The screenshot also shows an important truthfulness choice: the UI says Hosted workflow trace unavailable. The application does not invent stage completion or tool-call evidence when the hosted endpoint does not return trustworthy trace data. The answer can be displayed, but it must not be presented as proof of an internal execution path. Approval is a protocol, not a button The hosted web path uses App Service authentication with Microsoft Entra. The application accepts the injected principal only on the configured App Service host, validates tenant and object identifiers, applies a user allowlist, and performs an additional approver check for mutation requests. State-changing calls also require the expected origin and an application request header. Approval is then bound to five facts: The authenticated actor. The current demo generation. The affected batch. The quarantine action. A ten-minute validity window until first use. Resetting the demo creates a new generation, invalidating old handles. Consuming an approval does not make replay unsafe: the same bound handle can repeat the same quarantine operation, but domain code changes only positions that are not already quarantined. A second call reports an idempotent replay with zero additional positions changed. This is the difference between a human-in-the-loop interface and a human-authorized system. A modal dialog provides user experience; identity binding and deterministic policy provide control. Durable state needs concurrency semantics The hosted application stores each actor's synthetic session in a separate Blob object. A load returns both JSON state and its ETag. A save uses IfNotModified semantics; if another request updated the same state first, Azure Storage rejects the stale write and the API returns a conflict. conditions = ( {"etag": etag, "match_condition": MatchConditions.IfNotModified} if etag else {} ) await blob.upload_blob( json.dumps(state), overwrite=etag is not None, **conditions, ) Without that condition, two browser requests could both read the same approval state and overwrite one another using last-writer-wins behavior. Agent systems do not get a concurrency exemption: ordinary distributed-systems rules still apply. Fail closed, and make the failure legible The analysis adapter accepts only HTTPS Foundry endpoints with the expected path, uses a managed-identity token for https://ai.azure.com/.default , disables redirects, and enforces bounded connect and overall timeouts. It accepts only a completed assistant response with non-empty output text. If the endpoint times out, returns partial output, returns malformed data, or becomes unavailable, the application clears the analysis lease and reports that no inventory changed. It does not substitute a local answer and label it as hosted. Approval remains locked until a new hosted analysis succeeds. This can feel strict during a demo, but it protects provenance. A degraded fallback is useful only when the UI and audit model can identify it accurately. What is proven, and what is not The sample provides useful evidence for several engineering claims: Typed MCP tools reject malformed input. Hosted specialists receive read-only capabilities only. Approval is checked below the model and bound to identity and session state. Quarantine is idempotent in the synthetic domain. Blob ETags prevent stale session writes. Empty, partial, failed, and timed-out hosted responses fail closed. Local tests cover domain, MCP, workflow, API, and repository-hygiene behavior. It does not prove that the sample is a production recall platform. The scenario is synthetic. The local audit log is not tamper-evident. The hosted UI currently lacks trustworthy per-stage and per-tool trace rendering. Deployment-specific RBAC, EasyAuth configuration, telemetry access, model behavior, load characteristics, costs, and recovery procedures require validation in each environment. Evaluation evidence also expires. Golden cases and evaluator configuration are useful assets, but historical results are not a current release certificate. Re-run evaluations against the deployed agent version and inspect failures before making quality claims. Try the pattern Start with the deterministic path before provisioning cloud resources: Set-Location (git rev-parse --show-toplevel) py -3.13 -m venv caldova-recall-control/.venv ./caldova-recall-control/.venv/Scripts/python.exe -m pip install ` -r caldova-recall-control/requirements-ui.txt ./caldova-recall-control/.venv/Scripts/python.exe -m uvicorn ` control_tower_api:app ` --app-dir caldova-recall-control/src ` --host 127.0.0.1 ` --port 8091 Then inspect the MCP server over stdio: ./caldova-recall-control/.venv/Scripts/python.exe ` caldova-recall-control/scripts/inspect_mcp.py The inspector discovers the real tool schemas, reads the synthetic inventory, rejects malformed input and an unapproved mutation, and confirms the stock remains unchanged. Only after that local contract is understood should you configure a Foundry project, deployment identity, model deployment, and Hosted Agent. Engineering takeaways The most reusable lesson in Caldova is not the number of agents. It is the placement of authority. Give each model the smallest useful toolset. Prefer explicit workflow topology when the business process is known. Treat tool annotations as descriptive metadata, not access control. Bind consequential approval to authenticated identity, resource, action, and session generation. Put mutation and idempotency in deterministic domain code. Use optimistic concurrency for durable web state. Preserve provenance by failing closed instead of silently changing execution paths. Show only evidence the system actually captured. Agents are excellent at turning fragmented evidence into an actionable brief. Reliable systems make sure the brief and the action remain two different things. References Caldova source repository Model Context Protocol introduction Microsoft Agent Framework workflow capabilities Deploy a Hosted Agent in Microsoft Foundry Configure Microsoft Entra authentication for Azure App Service Manage concurrency in Azure Blob StorageMicrosoft Foundry Hosted Agents and MCP in Practice: Building Fibey Field Ops
An agent can produce a convincing answer while the system around it is still difficult to deploy, authorize, debug, and recover. For AI engineers and developers, that is often the real gap between a promising prototype and an application people can depend on. Fibey Field Ops makes that gap concrete. It is a synthetic fiber-operations assistant built with Microsoft Foundry Hosted Agents, Model Context Protocol (MCP), and Azure Container Apps. This walkthrough follows one field-service task through the implementation, then examines the deployment and operational decisions behind it. Introduction: the application is more than the model Imagine a technician preparing a fiber work order. Before leaving the depot, they need the job details, available parts, relevant procedures, and network status. Those facts belong to different systems. A useful assistant must retrieve them, combine them, and explain what is missing without inventing an answer. The challenge is not simply selecting a capable model. It is establishing reliable contracts between the model, its tools, the hosting platform, and the application. Fibey demonstrates those contracts with one hosted agent, five instruction skills, eleven operational tools, and five supporting Container Apps. There is an important qualification: Fibey is a protected, synthetic-data demonstration, not a production-ready field-service system. Its gateway mappings and work orders remain in memory, and it does not enforce per-user session ownership or approval of operational writes. Those limitations are useful teaching material rather than details to hide. The Fibey repository contains the implementation, infrastructure, documentation, and presentation deck. Repository access depends on its permissions. 1. Separate reasoning, instructions, and tool execution MCP is a standard interface for discovering and invoking tools. In Fibey, the agent connects to one Microsoft Foundry Toolbox MCP endpoint. The toolbox exposes capabilities backed by an inventory MCP server, a work-orders OpenAPI service, and a Search knowledge base. That common interface does not make the underlying systems identical. OpenAPI still describes an HTTP API, inventory still implements MCP, and knowledge retrieval still depends on indexed documents. The toolbox centralizes the agent-facing integration and connection configuration while preserving those implementation choices. Four terms describe different responsibilities: Concept Responsibility in Fibey Agent Classifies the request, loads instructions, selects tools, and constructs the response Skill An instruction document for a task such as inventory lookup or field briefing Tool An operation with an advertised input schema and a result Toolbox The curated MCP surface and references to downstream connections Fibey's five skills cover inventory lookup, work-order management, knowledge retrieval, work-order preparation, and field briefings. The last two coordinate several tools; they do not create additional agents. The configured toolbox exposes eleven operational tools directly: six inventory/status operations, four work-order operations, and one knowledge-retrieval operation. The agent uses their actual names and schemas. Discovery wrappers such as tool_search and call_tool are relevant only when a toolbox exposes them; they are not mandatory steps before every call. This distinction matters when debugging. A missing wrapper is not necessarily a broken integration. A skill mentioning a capability is also not proof that the runtime can invoke it. The current tool schema is the executable contract. 2. Follow the hosted Azure architecture Microsoft Foundry hosts the agent container and exposes its endpoint. Azure Container Apps (ACA) hosts the application services around it. The deployment source of truth is azure.yaml , which declares the GPT-5.4-mini model deployment and the hosted agent's Responses 2.0.0 protocol. The design keeps the browser-facing application separate from agent execution and backend integration. That makes it easier to inspect each boundary, but it does not make the whole deployment private or remove the need for application authorization. Treat GPT-5.4-mini as the sample's configured baseline, not a claim that it is optimal for every workload. When evaluating another model, measure tool-selection accuracy, schema compliance, grounded answers, latency, and cost per completed task rather than choosing from a fluent demo response alone. The engineering architecture expands this presentation view with resource ownership, identities, configuration, and telemetry paths. The application request path The browser signs in through Microsoft Entra ID at the UI's ACA authentication boundary. An explicit allowlist restricts access to the intended user. Nginx serves the React application and proxies chat requests to the internal FastAPI gateway. The gateway invokes the Foundry-hosted agent using its managed identity. Server-sent events (SSE), a streaming HTTP format, carry answer text, tool activity, citations, failures, and completion back to the UI. Gateway and dashboard ingress are internal to the ACA environment. Inventory and work orders have external ingress so the toolbox can reach them, but their operational endpoints require separate API keys. External accessibility is not anonymous access, and internal ingress is not a complete network-isolation strategy. The knowledge and status paths The knowledge pipeline starts with eight Markdown documents in the repository. A setup script uploads them to a private Blob Storage container, configures a Search data source and indexer, verifies ingestion, and creates the knowledge source and knowledge base used by Foundry IQ. This is not an embedding pipeline assembled implicitly by the chat application. The current configuration uses minimal reasoning and extractive retrieval, with defaults of three output documents and 6,000 output tokens. Its knowledge-base configuration and MCP endpoint use preview APIs, which need lifecycle and support review before production adoption. Network status takes a different path. Inventory's get_network_status tool reads the configured internal HTML dashboard. It is a narrow HTTP fetch, not browser automation, arbitrary website navigation, or a real operational clearance. Azure Container Registry supplies container images. ACA logs go to Log Analytics. Hosted code enables OpenTelemetry, the standard instrumentation framework for traces and metrics, but the repository does not provision Application Insights. The project's linked trace destination must be configured and verified separately. 3. Trace a work-order briefing through the implementation The most useful demonstration starts with one concrete request: prepare a technician for a job. The field-briefing skill provides the instructions for combining work-order data, inventory, procedures, and status without pretending that one backend contains everything. Use a request such as the following in a prepared synthetic environment. The exact wording and tool order may vary; the important evidence is which operations succeeded and which facts support the answer. Brief me on WO-007, including stock, relevant procedures, safety, and network status. The intended flow is: Load the field-briefing instructions and retrieve WO-007. Identify required parts and use a batch stock check when several parts need checking. Combine procedure and safety questions into one focused knowledge retrieval. Read the configured synthetic status dashboard through inventory MCP. Produce a briefing grounded in successful results, with source references and explicit gaps. Batching is a useful engineering choice, not a claim of a measured performance improvement. One batch stock call avoids unnecessary repeated requests. Combining related retrieval questions can also reduce duplicate context and tool traffic. This screenshot was captured from an authenticated deployment. The displayed briefing identifies an unavailable connector kit and available test equipment. It illustrates a specific synthetic response, not current stock or a benchmark. Use the supported hosted integration The relevant implementation is the hosted entrypoint. It uses FoundryToolbox from agent_framework_foundry_hosting , a FoundryChatClient , a skills provider, and ResponsesHostServer.run_async() . The hosting helper does more than attach a bearer token. It authenticates MCP requests and forwards the hosted runtime's per-request call ID. Replacing it with a generic transport can lose context that the platform expects. The entrypoint also closes credentials and clients when execution exits. Hosted skill discovery prefers published skills when available and retains bundled instructions as a fallback in auto mode. Explicit mcp mode fails if published skills cannot be loaded; file mode uses the bundled documents. This makes the fallback intentional rather than silently running without the task instructions. Distinguish history from compute affinity The gateway maintains two hosted mappings. previous_response_id links conversation history, while agent_session_id preserves affinity to the hosted compute session. Losing one is not equivalent to losing the other. Both mappings are held in gateway memory, which is why it remains at one replica. A UUID (universally unique identifier) is a conversation handle, not proof of ownership. Reset clears the local mappings but does not reset work orders or delete the old remote compute session. There is also a privacy distinction between storage and telemetry. Gateway requests use stored Responses history, while the agent's model-call options use store: false . Sensitive tracing being disabled does not mean all conversation persistence is disabled. 4. Run locally without confusing development and cloud boundaries Local development is useful for inspecting the gateway and agent without rebuilding the hosted image. It still calls a real Foundry project, model, and toolbox. It is not an offline simulation, and tool writes can affect the configured synthetic backend. Use Python 3.12+, uv, Node.js 24 LTS, and an authorized Azure developer identity. The commands below run from the repository root in PowerShell. Keep the existing dependency locks and copy .env.example only when creating a new local configuration. uv sync --frozen Copy-Item .env.example .env Set FOUNDRY_PROJECT_ENDPOINT , FOUNDRY_MODEL , and TOOLBOX_MCP_URL in the ignored .env . Use your environment's actual values. Backend API keys belong in Foundry connections, not in the browser or prompts. Start the gateway: az login uv run --frozen uvicorn fibey.gateway.api_server:app --host 127.0.0.1 --port 8080 In another terminal, start the UI: Set-Location ui npm ci npm run dev Open http://localhost:5173 . Vite forwards /api to the gateway on port 8080. These development servers do not reproduce the cloud Entra boundary, and a cloud toolbox cannot reach your workstation's localhost . Inspect the API progressively With the local gateway running, start with a health request in a separate PowerShell terminal: $base = "http://127.0.0.1:8080" Invoke-RestMethod "$base/api/health" Next, create a UUID conversation and submit one synthetic request: $session = [guid]::NewGuid().ToString() $body = @{ message = "Show me WO-007."; session_id = $session } | ConvertTo-Json -Compress $response = Invoke-WebRequest "$base/api/chat" -Method Post -ContentType "application/json" -Body $body $response.Headers["X-Session-Id"] $response.Content This prints the completed SSE body rather than animating the stream. Reuse the same session for a follow-up by extracting a small helper: function Invoke-FibeyTurn { param([string] $Message, [string] $SessionId) $payload = @{ message = $Message; session_id = $SessionId } | ConvertTo-Json -Compress $result = Invoke-WebRequest "$base/api/chat" -Method Post ` -ContentType "application/json" -Body $payload return $result.Content } Invoke-FibeyTurn -Message "What parts does that work order need?" -SessionId $session The helper uses $base and $session from the preceding examples. It makes continuity explicit without hiding the API contract. The browser remains the better place to watch incremental text and activity. See local development for reset behavior and supporting-service details. 5. Deploy artifacts and infrastructure together Fibey's initial deployment is intentionally staged. A container image, its target port, its health probes, its registry association, and its runtime permissions must agree. Successfully building an image does not establish that agreement. The first supporting-infrastructure pass creates placeholder apps on port 80. Once all five real images have been published, a second pass applies the images with their application ports and probes. Plain azd deploy of a supporting service is not the initial port-switch mechanism for this sample. The deployment guide gives the complete sequence and prerequisites: Select the intended environment and provision the Foundry layer. Configure project aliases, the Entra application, the allowed user, and separate API keys. Provision placeholder supporting apps and verify scoped access and registry identity associations. Publish all five supporting images and confirm every image setting is populated. Provision supporting infrastructure again to apply the matching images, ports, and probes. Run knowledge setup, then toolbox setup. Deploy fibey-agent and perform protected-path acceptance checks. For example, this publication step comes after placeholders and access checks, not at the start of an unconfigured environment: azd publish status-dashboard azd publish inventory-mcp azd publish work-orders-api azd publish gateway azd publish ui Only after all five SERVICE_*_IMAGE_NAME settings exist should the next azd provision infra apply them. Keep those settings: an empty value selects a placeholder again. This is where DevOps becomes tangible. Review source, Bicep, locks, schemas, skills, and configuration together; record accepted image references, model configuration, toolbox version, and hosted-agent version. GitOps adds controlled reconciliation of reviewed desired state. The repository supplies the ingredients, not an existing CI/CD pipeline or GitOps controller. Toolbox promotion deserves the same discipline. The unversioned consumer endpoint follows the published default version. A version-specific developer endpoint can test a candidate before promotion. Updating a shared connection or default can affect consumers without rebuilding their images, so Git history alone is not a rollback mechanism. 6. Treat governance, observability, and scale as separate concerns Role-based access control (RBAC) grants an identity permission at a resource scope. The control plane creates and configures Azure resources; the data plane performs application work such as invoking agents, querying Search, or reading blobs. The identities used across those operations are not interchangeable. In particular, a successful UI login is not automatic end-user identity passthrough to every tool, and resource provisioning permissions do not establish all runtime permissions. Boundary Identity or credential Browser to UI Entra user, application registration, and user allowlist Gateway to Foundry Gateway managed identity with project-scoped invocation access Hosted agent to model and toolbox Foundry-provided runtime agent identity Toolbox to inventory and work orders Separate API keys stored in project connections Toolbox to Search knowledge base Foundry project managed identity Search indexer to documents Search managed identity with Blob reader access Image delivery adds another boundary: each supporting app needs both AcrPull and a registry association selecting its identity. The Foundry project identity used for infrastructure operations is also distinct from the hosted agent's runtime identity. Approval must exist outside the prompt Fibey's synthetic work-order writes run without an enforced human approval round trip. That is appropriate to understand in a disposable demo and inappropriate to conceal when discussing production. The hosted toolbox documentation is explicit: approval metadata alone does not block tools/call . The runtime must pause, collect a decision, and resume or reject the exact proposed operation. A prompt saying "ask first" is not equivalent to that control. Real writes need server-side authorization, approval tied to the arguments, idempotency, and a durable audit record. If a write's response is uncertain, read its state before retrying. Treat tool results as untrusted data rather than instructions. Debug the boundary that failed The activity sidebar makes tool use visible, but it is not a durable audit log or access to the model's hidden reasoning. Combine it with timestamps, request/session identifiers, ACA logs, and configured hosted traces. Keep sensitive message tracing off unless a controlled investigation explicitly requires it. Verify that a synthetic trace reaches the intended sink; enabled instrumentation alone does not prove ingestion. Symptom First checks UI returns 502 Gateway revision, target port, fixed Nginx upstream, SNI, and certificate trust Agent or tool returns 401/403 Identity, token audience, role scope, and downstream connection credential Knowledge retrieval fails Indexer completion, project Search role, API version, and advertised input schema Briefing receives 429 Model quota, concurrency, retrieval budgets, and repeated tool calls Stream ends early Terminal response event, timeout, malformed SSE, and transport failure The gateway must surface failed or truncated streams rather than treating partial text as a completed action. A final stream terminator after an error is not application success. Scale only after identifying state and capacity limits Foundry manages hosted runtime capabilities, but it does not externalize Fibey's gateway mappings or in-memory work orders. Adding replicas before redesigning that state would undermine continuity and consistent updates. Model throughput, hosted compute, downstream API capacity, Search, and ACA are separate constraints. Measure latency, errors, tokens, tool counts, and recovery under concurrent load. Registry storage/builds, model inference, compute, Search, Blob Storage, and telemetry all have costs; using a managed platform does not remove the need for budgets. 7. Evaluate the result before adding more agents The practical result is an inspectable workflow that combines heterogeneous systems into a grounded response. We can demonstrate inventory lookup, a synthetic work order, cited knowledge, a fixed status fetch, and a combined briefing through one agent-facing toolbox. That is functional evidence, not a production certification or benchmark. This post does not establish latency percentiles, cost per successful task, throughput, or availability under load. Those measurements should be acceptance criteria for an adaptation, not numbers inferred from a screenshot. Multi-agent coordination is a possible next design, not deployed Fibey behavior. A coordinator could delegate read-only inventory and procedure tasks while a restricted specialist handles approved writes. Each handoff would need typed inputs/results, correlation IDs, deadlines, cancellation, and budgets. Sharing a toolbox alone does not implement that coordination. Specialists also introduce additional failure paths, identity decisions, and state. Start with them only when task complexity, ownership, or isolation requirements justify the cost. The current five skills already provide modular instructions without creating five independently operated agents. 8. Summary: reuse the engineering boundaries Fibey's reusable pattern is the separation of model reasoning, task instructions, tool contracts, hosted execution, and operational controls. MCP gives the agent a common tool interface; Foundry supplies managed hosting and integration capabilities. The application team remains responsible for the guarantees around its data and actions. Start with one workflow and verify each boundary. Then add durable state, per-user authorization, enforced approvals, secret rotation, appropriate networking, evaluation, and recovery before introducing real operational data or more autonomous behavior. For a practical starting point, use the repository, follow the deployment guide, and reproduce the session walkthrough with synthetic data. The presentation deck provides the same architecture and demo sequence for a team discussion. References The repository describes what this sample implements. Microsoft documentation describes the surrounding platform capabilities, responsibilities, and supported integration contracts. Read both before adapting the solution, especially where preview APIs, identity behavior, or approval enforcement affect your requirements. Fibey Field Ops repository Fibey engineering architecture Fibey deployment guide Fibey toolbox integration Microsoft Foundry hosted agents Use a toolbox with a hosted agent Official Python hosted-agent toolbox sampleChoosing a real-time voice architecture on Microsoft Foundry: three enterprise patterns
A practical comparison of three real-time voice architectures on Microsoft Foundry, including implementation tradeoffs and four enterprise release gates for residency, networking, retrieval authorization, and tool credentials.786Views1like0CommentsBuild an AI-assisted support email workflow with Power Automate and Microsoft Foundry
Difficulty: Intermediate A support email arrives with a product name, an error code, and a country. Before anyone can help, someone has to interpret the request, find the right team, record the case, and decide what to tell the customer. In this tutorial, you build that support email workflow with Microsoft 365, Power Automate, Azure Functions, and Microsoft Foundry. AI extracts the request details and recommends a support team. The flow checks the recommendation against a SharePoint catalog, asks a person to approve it, and creates an editable Outlook acknowledgement draft. The central design decision is who can do what: the model proposes a routing key; the flow validates it; a person approves draft creation. No customer-facing send action is included. What you will build Mina manages support operations at Aster Imaging, a fictional imaging-equipment company. Claire, a fictional distributor in France, reports error E42 on a NovaScan X2 after installing software version 4.2.1. When Mina approves the validated recommendation, the workflow produces: Input or decision Recorded result Claire's support email One case linked to its source message ID Validated AI recommendation Technical Support, NovaScan X2, France, EU Distributor Support Mina selects Approve Approval outcome and an editable acknowledgement draft Missing information, a risk flag, rejection, or timeout A case held in Needs Review The recorded approved case shows the review outcome and successful automation state: All names, products, organizations, and test messages are synthetic. This is a reference implementation for a lab, based on public support-form patterns, rather than a description of a company's internal process. Who this is for This walkthrough is for Power Platform makers and developers who can already create a cloud flow and a SharePoint list, and want to connect AI to a workflow with explicit review boundaries. You will use the Azure portal and a few terminal commands to deploy the supplied Function project. In this tutorial Architecture and resources Tutorial scope and boundaries Prerequisites Download the sample Step 1: Prepare the mailbox and SharePoint lists Step 2: Deploy and connect the AI classifier Step 3: Build and validate the support email flow Step 4: Add human approval, create a draft, and evaluate the workflow Troubleshooting Production considerations Clean up the lab Architecture and resources Workflow architecture The completed workflow follows this path: When Claire's message arrives in the shared mailbox, Power Automate first checks the message ID to avoid creating a duplicate case. It then sends the subject and body to the AI classifier. The classifier identifies the request as technical support, extracts NovaScan X2 and France , and recommends EU Distributor Support . Power Automate does not accept that recommendation without verification. It checks the structured response, confidence band, and risk flags, and confirms that the recommended routing key matches an active entry in the SupportTeams list. If the checks pass, the workflow records the case in SharePoint and asks Mina to review the recommendation. If Mina approves it, Power Automate creates an acknowledgement draft in Outlook. It does not send the message. If the recommendation has low confidence, contains a risk flag, fails validation, or is rejected by Mina, the case remains in Needs Review . Resource names To keep the steps consistent and easy to follow, this tutorial uses the resource names below. You can use different names in your environment, but keep track of the corresponding values as you work through the steps. Resource Name Power Platform environment RW Development Solution RW Global Support Intake Publisher prefix rw SharePoint site RW Lab SharePoint lists Cases , SupportTeams Shared mailbox RW Lab Support Main flow RW - Global Support Intake - v1 Azure resource group Your lab resource group; referenced as <RESOURCE_GROUP> Microsoft Foundry resource Your globally unique resource; referenced as <FOUNDRY_RESOURCE_NAME> Microsoft Foundry project The project opened in the Foundry portal; referenced as <FOUNDRY_PROJECT_NAME> Model deployment gpt-5.4-mini Azure Function App Your globally unique app; referenced as <FUNCTION_APP_NAME> Tutorial scope and boundaries This tutorial builds the workflow in stages so that each layer can be tested before the next one is introduced. Step 1 prepares the shared mailbox and SharePoint records. Step 2 deploys the AI model and exposes the bounded classifier to Power Automate. Step 3 builds the support email flow and validates normal, review, fallback, and duplicate paths before approval exists. Step 4 adds human approval and draft creation, then evaluates the completed workflow end to end. What this tutorial covers Create a shared mailbox, SharePoint case register, and allow-listed support-team catalog. Connect a bounded AI classifier that returns structured fields and an advisory routing key. Build deterministic deduplication, recommendation validation, and exception paths in Power Automate. Add a human approval checkpoint that creates an Outlook draft without sending it. Walk through three representative paths: accepted recommendation, review hold, and approval timeout. Record the remaining synthetic checks in a compact validation matrix. What is out of scope Send customer-facing email automatically. Let AI assign the final owner, approve a request, change SupportTeams , or send a message. Process real customer data, retrieve attachment contents, or scan attachments. The mailbox flow passes attachments = [] ; attachment-metadata checks are exercised only in the direct classifier tests. Merge a new follow-up email into an existing case. The duplicate guard checks the same message ID; a new reply has a different ID. Treat the prototype results as production accuracy, security, capacity, or SLA claims. Prerequisites A Microsoft 365 test tenant with Exchange Online and SharePoint Online A licensed test user who can open Outlook and SharePoint Permission to create a shared mailbox and assign mailbox delegation A Power Platform environment with Dataverse for the solution and approvals, and permission to use the Office 365 Outlook, SharePoint, Approvals, and custom connectors For a development-only lab, an eligible Power Apps Developer Plan environment; otherwise, Power Automate use rights that cover this flow and its custom connector An Azure subscription in which you can create or select a Microsoft Foundry resource, deploy a model, and create an Azure Functions app Permission to enable the Function app's managed identity and assign it the Cognitive Services OpenAI User role Node.js 22, Azure CLI, and Azure Functions Core Tools v4 on the workstation used for deployment; the tested Core Tools version is 4.0.7512 Permission to create custom connectors, connection references, environment variables, and solutions in the Power Platform environment Licensing and cost Microsoft 365 access alone does not establish entitlement to the custom connector used here. Confirm the target environment's license and data policies before building the flow. The Developer Plan includes custom connectors for development and testing; production use needs appropriate paid rights. See the Power Platform licensing FAQ. Azure model inference, Function hosting, storage, and monitoring can incur charges. Usage depends on the selected region, hosting plan, model, message size, and test volume. Check those resources in Azure Cost Management and clean up the dedicated lab resources when finished. The recorded evaluation is not a cost benchmark. Download the sample Get the companion code from the example project on GitHub. Clone the repository or download and extract the ZIP to a local folder. Use the extracted repository folder as the sample root in the commands below. The repository contains the Function source, locked dependencies, connector definition, synthetic data, setup instructions, and local tests. You build the Power Automate flow in the designer; this sample is not an importable Power Platform solution. Path from the sample root Purpose azure/global-support-classifier/ Deployable Azure Function project azure/global-support-classifier/connector/apiDefinition.swagger.json Custom connector definition; replace its sample host before import data/aster-imaging/support-teams.json Twelve sample team keys and descriptions data/aster-imaging/inquiries.json Twenty synthetic inquiries and expected results tests/aster-imaging-fixtures.test.mjs Offline dataset checks tests/global-support-classifier-regression.mjs Authenticated classifier evaluation evidence/classifier-regression-20260808.md Historical lab result and evaluation limits The model evaluation results come from the August 8, 2026 lab. This article combines recorded lab screenshots with configuration screens recaptured on September 8–9, 2026 for clearer step-by-step instructions. Recapturing a configuration screen does not verify a new workflow run. A separate single-request connector test on September 8, 2026 returned HTTP 200 and passed schema validation; its input and response are shown in Step 2. The GitHub v0.1.0 sample preserves the September 8, 2026 implementation snapshot, with repository setup documentation added on September 9. Portal labels, model availability, and quotas can differ in your tenant. Step 1: Prepare the mailbox and SharePoint lists In this step, you prepare the Microsoft 365 resources used by the workflow. The shared mailbox receives support requests, the SharePoint site provides a private workspace, the Cases list stores each request and its workflow state, and the SupportTeams list defines the support teams that the AI classifier is allowed to recommend. Services and tools used in this step Service Purpose Exchange Online Create the shared support mailbox and assign mailbox permissions. SharePoint Online Create the private team site and the Cases and SupportTeams SharePoint lists. Create the shared support mailbox The shared mailbox provides a single email address for support requests. Power Automate monitors this mailbox and starts the workflow when a new message arrives. For example, Claire sends her request to RW Lab Support , and the new message starts the support email flow. Open the Exchange admin center and go to Recipients > Mailboxes. Create a shared mailbox with these values: Display name: RW Lab Support Email alias: rw-lab-support Domain: your lab tenant's default domain Note: Creating the mailbox requires the Exchange Administrator or Global Administrator role. After creation, grant the lab user both Full Access and Send As. A Microsoft 365 license does not grant these permissions automatically. Review the current shared mailbox permissions and licensing rules before using this design in production. After the mailbox is created, delegate both permissions to the lab user: Full Access lets the user open and modify the shared mailbox. Send As lets the user create a draft using the shared mailbox address. Note: Full Access alone is not sufficient for the later draft-from-shared-mailbox step. To review delegation in the Microsoft 365 admin center, go to Teams & groups > Shared mailboxes, select RW Lab Support, and open Read and manage permissions. This is the Full Access permission described above. Select Add permissions to add the lab user. Return to the mailbox details and also configure Send as permissions for that user. Create the SharePoint site The SharePoint site provides a private workspace for the workflow's lists and case data. Keeping these resources on one site also makes permissions and ownership easier to manage. In this tutorial, RW Lab contains both the Cases list and the SupportTeams list used to process Claire's request. Open SharePoint and select Build in the left navigation. Under Start building, select Site. Then select Team site. If prompted to choose a template, select Standard team. Configure the site using the following values. Field Value Site name RW Lab Site description Synthetic global support workflow lab Group email address Accept the generated alias if it is available. Site address Confirm the generated SharePoint address. Privacy settings Private - only members can access this site Language English Important: The group email address and site address must be unique in your tenant, so SharePoint may adjust the generated values. You cannot change the site's default language after creation. The red boxes group the fields you can complete on this screen and the final Create site button. Check the generated site address separately. This annotated image uses the recorded lab screen. Select Create site, add owners or members if required, and finish the site setup. Create the Cases list The Cases list is the durable record after the flow has enough information to create a case. It stores source-message metadata, the AI proposal, the validated or fallback team, review outcome, and automation status. It does not copy the complete email body or attachments into SharePoint. For Claire's request, the flow first checks SourceMessageId , loads the active team catalog, calls the classifier, and validates the proposed routing key. It then creates one row containing the message metadata and AI result. A normal high-confidence row is updated as it moves through approval and draft creation. Low-confidence, risky, invalid-key, and classifier-unavailable requests are created directly as Needs Review . A duplicate creates no second row, and an unexpected failure before row creation appears only in the flow run history and the internal failure notification. Open SharePoint and select Build in the left navigation. Under Start building, select List, and then select Blank list. Enter Cases for Name, select RW Lab for Save to, and then select Create. SharePoint opens the new Cases list after creating it. In the Cases list, select Add column, choose the matching column type, and configure the columns shown in the following table. The Title column already exists by default; configure it as shown rather than creating another Title column. Note: In the current SharePoint interface, Text creates a single-line text column. Use Multiple lines of text for longer values such as Summary , RoutingReason , and AIProposal . Column Type Configuration Title Single line of text Required; stores the case ID SourceMessageId Single line of text Enforce unique values ReceivedAt Date and time Include time RequesterEmail Single line of text Synthetic test addresses only EmailSubject Single line of text Original subject InquiryType Choice Technical Support, Quote, Demo, Product Information, Partnership, Complaint, Other Product Single line of text AI-extracted canonical product name Country Single line of text AI-extracted country Summary Multiple lines of text AI-generated summary in plain text RecommendedRoutingKey Single line of text Routing key proposed by the AI classifier RecommendedTeam Single line of text Team name resolved from SupportTeams RoutingReason Multiple lines of text AI-provided reason for the recommendation ConfidenceBand Choice High, Medium, Low; use Low as the default for a new lab RiskFlags Multiple lines of text Risk signals returned by the classifier MissingFields Multiple lines of text Required information not found in the message AIProposal Multiple lines of text String form of the classification object, or a classifier-unavailable message ClassificationSource Single line of text Records validated recommendation, held recommendation, or classifier-unavailable fallback AssignedToText Single line of text Reviewer email resolved from SupportTeams Status Choice New, Classified, Needs Information, Awaiting Approval, Assigned, Needs Review, Failed DueAt Date and time Include time ApprovalOutcome Single line of text Approve, Reject, or Timeout ApprovalComment Multiple lines of text Plain text DraftMessageId Single line of text ID returned by the Outlook draft action AutomationStatus Choice Processing, Success, Review, Failed LastAutomationRun Date and time Include time Set SourceMessageId to enforce unique values. Power Automate also checks for the message ID before creating a row; the SharePoint constraint provides a second layer of duplicate protection. To verify the constraint, open Settings > List settings > SourceMessageId. Keep the type Single line of text, set Enforce unique values to Yes, and select OK if you changed it. Open Settings > List settings > ConfidenceBand. Enter High, Medium, and Low on separate lines, use the Drop-down menu display, and leave fill-in choices disabled. For your new list, set Default value to Low, then select OK. The existing lab default is High; use Low for a new build. Classified and Needs Information remain in the current lab list from the earlier implementation. The AI-recommendation path documented below does not write those two values. The current Catch scope also does not write Failed ; it sends an internal failure notification and terminates the run. Create a list view named Tutorial Evidence with Title, EmailSubject, Product, Country, RecommendedTeam, ConfidenceBand, Status, ApprovalOutcome, AutomationStatus, DraftMessageId, ClassificationSource, and ReceivedAt. Sort ReceivedAt newest first. Later validation steps use this view; use the row details pane for columns that do not fit on screen. Create the SupportTeams list The SupportTeams list defines the support teams or queues that the AI classifier is allowed to recommend. Power Automate uses it to validate each recommendation and resolve the reviewer and SLA target. It does not store individual team members or calculate routing from product and country. For Claire's request, the AI proposes eu_distributor_support . Power Automate confirms that this key is active and resolves it to EU Distributor Support , its reviewer, and its four-hour SLA target. Open SharePoint and select Build in the left navigation. Under Start building, select List, and then select Blank list. Enter SupportTeams for Name, select RW Lab for Save to, and then select Create. SharePoint opens the new SupportTeams list after creating it. In the SupportTeams list, select Add column, choose the matching column type, and configure the columns shown in the following table. The Title column already exists by default; configure it as shown rather than creating another Title column. Column Type Configuration Title Single line of text Display name RoutingKey Single line of text Enforce unique values TeamName Single line of text Support team name ApproverEmail Single line of text Reviewer address SLAHours Number No decimal places Active Yes/No Default Yes Important: Set RoutingKey to enforce unique values. This ensures that each AI recommendation resolves to exactly one support team. Power Automate accepts a recommendation only when the key matches an active row, then uses that row's team name, reviewer, and SLA target for the normal approval path. Open Settings > List settings > RoutingKey. Keep Single line of text, set Enforce unique values to Yes, and select OK. The September 9 lab capture below shows the existing No setting; it is a configuration gap, not the recommended setting for your new list. Add the 12 synthetic rows shown below. For closer parity with the direct evaluator, use each matching description in data/aster-imaging/support-teams.json as Title; the shorter titles below are the labels used in the original mailbox lab. The flow sends Title as the routing description, so changing it changes model input and requires a rerun. Replace each <YOUR_EMAIL> placeholder with the lab reviewer address. Title RoutingKey TeamName ApproverEmail SLAHours Active EU distributor support eu_distributor_support EU Distributor Support <YOUR_EMAIL> 4 Yes NA technical support na_technical_support NA Technical Support <YOUR_EMAIL> 4 Yes APAC distributor support apac_distributor_support APAC Distributor Support <YOUR_EMAIL> 4 Yes Software support software_support Software Support <YOUR_EMAIL> 4 Yes EU regional sales eu_regional_sales EU Regional Sales <YOUR_EMAIL> 8 Yes NA regional sales na_regional_sales NA Regional Sales <YOUR_EMAIL> 8 Yes APAC regional sales apac_regional_sales APAC Regional Sales <YOUR_EMAIL> 8 Yes Global partnerships global_partnerships Global Partnerships <YOUR_EMAIL> 8 Yes Global support global_support Global Support <YOUR_EMAIL> 8 Yes Security review security_review Security Review <YOUR_EMAIL> 1 Yes Privacy review privacy_review Privacy Review <YOUR_EMAIL> 1 Yes Safety and compliance safety_and_compliance Safety and Compliance <YOUR_EMAIL> 1 Yes Review the completed list. Confirm that all 12 rows show Active = Yes and that no RoutingKey value appears more than once. The recorded lab catalog contains 12 active teams with the routing keys and SLA hours shown above. This list view omits reviewer addresses; configure ApproverEmail for every row using your lab reviewer account. Step 2: Deploy and connect the AI classifier In Step 1, you created the shared mailbox and SharePoint lists that provide the workflow's email channel, case record, and approved support-team catalog. In this step, you add the AI classification layer. You deploy an Azure OpenAI model in Microsoft Foundry, connect it to an Azure Function, and expose the Function to Power Automate through a custom connector. For example, suppose Claire emails the shared mailbox to report that a NovaScan X2 in France shows error E42 during startup after software version 4.2.1 is installed. The classifier extracts NovaScan X2 as the product, France as the country, and technical_support as the inquiry type. It summarizes the problem, returns a high confidence band, and recommends the routing key eu_distributor_support . This result is only a proposal. The classifier cannot approve the request, assign an owner, update the case, or send email. In Step 3, Power Automate validates the proposed key against the active SupportTeams list before deciding whether the request can proceed to human approval. Services and tools used in this step Service or tool Purpose Microsoft Foundry (Azure OpenAI models) Create or select the Foundry resource and deploy the model used for classification. Azure Functions Call the model, enforce the structured response contract and guardrails, and expose a bounded HTTP endpoint. Microsoft Entra ID and Azure RBAC Allow the Function app to call the model through its managed identity. Power Automate custom connectors Make the Azure Function operation available to cloud flows. Power Platform solutions Package the custom connector, connection references, and environment variables used by the workflow. Note: The local deployment tools used in this step are Node.js 22, Azure CLI, and Azure Functions Core Tools. Create a Foundry resource and deploy the model An Azure subscription is the billing and access boundary. Inside that subscription, a Microsoft Foundry resource provides the model endpoint and quota. A Foundry project is the workspace you open in the Foundry portal, and a model deployment inside that project gives the Function app a stable deployment name to call. Creating a subscription or resource alone does not deploy a model. The tested lab uses gpt-5.4-mini , pinned to model version 2026-03-17 . It provides the structured output and instruction-following behavior needed by this bounded extraction and routing task without using a larger model for every incoming message. Sign in to the Azure portal and confirm that the correct directory and subscription are selected. Search for Microsoft Foundry, select Create, and create a Foundry resource if you do not already have one that the lab can use. Choose the target subscription and resource group, enter a globally unique resource name, select a region in which gpt-5.4-mini has quota, and keep the Standard S0 pricing tier. For a disposable lab, the basic public-network configuration is sufficient; use private networking and organization-approved controls for production. Note: You need permission to write the resource, such as Contributor or Owner, and separate model quota in the selected region. Microsoft documents the current resource fields in Create a Microsoft Foundry resource. Open the Microsoft Foundry portal and verify the directory and subscription. Select a project associated with the Foundry resource you created. If no project exists yet, create one under that resource and record its name as <FOUNDRY_PROJECT_NAME> . Enable the New Foundry experience if the portal offers the switch. On the project home page, select View deployments under Use a model. You can return to the same page later through Build > Deployments. On Deployed models, select Deploy > Deploy a base model, search for gpt-5.4-mini , and configure the deployment with the following tested values: Setting Tested value Deployment name gpt-5.4-mini Model version 2026-03-17 Deployment type Global Standard Capacity 100 thousand tokens per minute Version upgrade policy Once current version expires Content filter Microsoft.DefaultV2 Important: Capacity is quota, not a target for the tutorial. Select a lower value when your subscription has less quota; the synthetic lab traffic does not require 100K tokens per minute. If this model or version is unavailable in your region, choose an approved region or model only after confirming structured-output support, then repeat the final evaluation in Step 4 with that exact deployment. Select Deploy and wait until the deployment state is Succeeded. Record the resource endpoint and deployment name. The Function app uses the deployment name, not the catalog model label, when it constructs the request. The screenshot below is a verification view of the completed lab deployment, not the initial creation form. Optionally verify the deployed version from Azure CLI. This read-only command helps catch the common mistake of configuring the Function with a deployment that exists under another resource or subscription: az login az account set --subscription "<SUBSCRIPTION_ID>" az cognitiveservices account deployment show ` --resource-group "<RESOURCE_GROUP>" ` --name "<FOUNDRY_RESOURCE_NAME>" ` --deployment-name "gpt-5.4-mini" ` --query "{state:properties.provisioningState, model:properties.model.name, version:properties.model.version, sku:sku.name, capacity:sku.capacity}" The tested deployment returned Succeeded , model gpt-5.4-mini , version 2026-03-17 , SKU GlobalStandard , and capacity 100 . Review the current Microsoft guidance for deploying Foundry models and model version update policies because availability, quota, and portal labels can change. Create, configure, and deploy the Azure Function Create and configure the Azure resource in the portal, then use Azure Functions Core Tools to publish the tested repository project. Do not copy the individual JavaScript files into the portal editor; this project has multiple source files, locked npm dependencies, and automated tests that should remain together. In the Azure portal, select Create a resource, search for Function App, and select Create. Configure the Function app using the following tested lab values. If you already have a compatible Node.js 22 Function app, open it and continue with the identity step. Setting Tested lab value Subscription The subscription that contains the Foundry resource Resource group The lab resource group Function App name A globally unique name; record it as <FUNCTION_APP_NAME> Publish Code Runtime stack Node.js Version 22 Operating system Linux Region The lab region; the tested app uses East US 2 Hosting Consumption; the tested app uses the Y1 / Dynamic plan Note: Microsoft currently recommends Flex Consumption for new serverless Function apps. The implementation documented here was tested on Linux Consumption. If you select a different hosting plan, confirm its deployment and networking behavior before treating the tutorial results as equivalent. See Create a function app in the Azure portal. Select Review + create, select Create, and wait for the deployment to finish. Open the Function App resource and record its default host name from Overview. Under Settings > Identity, enable the system-assigned identity. On the Foundry resource, assign that identity the Cognitive Services OpenAI User role. The Function uses DefaultAzureCredential ; it does not store an Azure OpenAI API key. See the Microsoft guidance for managed identities in Azure Functions and Azure OpenAI role assignment. Under the Function app's environment variables or configuration settings, add these values: Setting Value AZURE_OPENAI_ENDPOINT Azure OpenAI endpoint, typically https://<FOUNDRY_RESOURCE_NAME>.openai.azure.com AZURE_OPENAI_DEPLOYMENT gpt-5.4-mini AZURE_OPENAI_API_VERSION 2024-10-21 Keep Show values disabled while capturing or sharing this page. The list should show the three setting names without exposing their values. Note: These are Azure Function app settings. They are separate from the rw_* Power Platform solution environment variables created later in this step. Download or clone the GitHub sample, open a terminal at the sample root, and move to the Function project: cd azure/global-support-classifier Verify the local tool versions, install the locked dependencies, and run the Function tests: node --version func --version npm ci npm test The tested deployment uses Node.js 22 and Azure Functions Core Tools 4.0.7512 . If func is not available, install a supported Core Tools v4 release by following the Core Tools installation guidance. In src/functions/classifySupportInquiry.js , locate buildAzureRequest and confirm that the GPT-5 request uses max_completion_tokens = 1100 and reasoning_effort = none . Do not add temperature or the older max_tokens field. GPT-5 reasoning models count reasoning and visible output against the completion-token budget; Microsoft documents the compatible parameters in Use reasoning models. Sign in with az login , select the intended subscription with az account set --subscription <SUBSCRIPTION_ID> , and confirm it with az account show --query name -o tsv . From the Function project directory, publish the complete project to the Function app: func azure functionapp publish <FUNCTION_APP_NAME> --javascript The --javascript option is explicit because automatic language detection did not identify this repository project during the tested publish. A successful publish reports Deployment completed successfully , synchronizes the classifySupportInquiry trigger, and prints its invoke URL. Core Tools packages and deploys the complete project from the current directory; review the Core Tools publishing guidance before using another hosting plan. Return to the Function App Overview page and confirm that classifySupportInquiry appears as an enabled HTTP function. This screen is meaningful only after the Core Tools publish succeeds. Define the response contract Open azure/global-support-classifier/src/functions/classifySupportInquiry.js and locate classificationSchema . Confirm that every property is required, bounded values use enumerations, and additionalProperties is false . Confirm that a successful call returns version 2.0 of this contract: { "schemaVersion": "2.0", "classification": { "inquiryType": "technical_support", "product": "NovaScan X2", "serialNumber": "NSX2-2407138", "country": "France", "organization": null, "urgency": "normal", "language": "en", "summary": "NovaScan X2 shows error E42 during startup after version 4.2.1.", "missingFields": [], "evidence": [ "NovaScan X2 serial NSX2-2407138", "error E42 during startup", "software version 4.2.1" ], "confidenceBand": "high", "riskFlags": [], "recommendedRoutingKey": "eu_distributor_support", "routingReason": "Technical support for NovaScan X2 in France routes to EU distributor support." } } Confirm that the contract includes recommendedRoutingKey , routingReason , and riskFlags . Power Automate stores all three and validates the key before resolving a team. Treat confidenceBand as a workflow category rather than a calibrated probability. Allow only high , medium , or low , and bound the risk values to prompt injection, unsupported attachment, privacy, safety, non-English, multiple-intent, low-confidence, and invalid-key signals. Supply the active team catalog with each request Open azure/global-support-classifier/connector/apiDefinition.swagger.json . Confirm that it is an OpenAPI 2.0 definition and that the request includes routingOptionsJson . Confirm that Power Automate will serialize each active catalog entry with this shape. The example below shows one of the 12 rows: [ { "routingKey": "eu_distributor_support", "teamName": "EU Distributor Support", "description": "Technical support for NovaScan products in Europe." } ] In classifySupportInquiry.js , confirm that the Function parses routingOptionsJson before constructing the model request. This string boundary avoids a Power Automate custom-connector metadata issue with arrays of objects; the model still receives a structured routingOptions array. Apply deterministic post-processing after the model response: detects prompt-injection phrases and unsafe attachment extensions; forces safety, privacy, and security signals to the corresponding review key when that key is active; removes a missing software version or screenshot flag when the source message contains that evidence; changes an unknown product-and-country result to low confidence; marks a recommendation invalid when its key is not in the supplied active catalog. Keep the Function advisory. These controls constrain the proposal, but Power Automate still performs the operational validation and assignment. The active destinations, reviewers, and SLA values live in SharePoint. The sample still encodes product/country routing precedence in the system prompt and post-processing code; changing that policy requires a code review and regression run as well as any catalog update. Keep attachment retrieval disabled for this version. The classifier supports attachment metadata and its direct regression tests exercise unsupported-attachment detection, but the live flow passes an empty attachments array and therefore does not inspect or block attachments. Create the custom connector Two different authentication boundaries are involved. The Function app calls Microsoft Foundry with its managed identity, so it stores no Azure OpenAI key. Power Automate calls the Function endpoint with x-functions-key ; that Function host key is not an Azure OpenAI key. Open azure/global-support-classifier/connector/apiDefinition.swagger.json from the sample root. Before importing the connector, replace its host value with <FUNCTION_APP_NAME>.azurewebsites.net . Keep basePath set to /api , the scheme set to https , and the operation ID set to ClassifySupportInquiryV3 . In Power Automate, select the RW Development environment. Open More > Discover all, find Data, and select Custom connectors. Select New custom connector > Import an OpenAPI file, enter RW Support Classifier , and upload apiDefinition.swagger.json . In the connector wizard, review General, Security, and Definition in order. On General, confirm HTTPS, your Function app's host name, and Base URL = /api, then select Security. Confirm API Key, Parameter label = Function key, Parameter name = x-functions-key, and Parameter location = Header. Select Definition and open Classify a support inquiry. Under Request, confirm POST, your Function URL ending in /api/classify-support-inquiry, and the required body parameter imported from the OpenAPI file. Continue with the response check below before creating the connector. The screenshots show an existing connector, whose toolbar displays Update connector instead of Create connector. Scroll farther down to Response and open 200 (Structured candidate fields). Confirm that References Used includes ClassifierResponse and Classification, and that Body exposes fields such as confidenceBand, routingReason, and schemaVersion. Select Back to return to the action definition, then select Create connector (or Update connector when modifying an existing connector). This screen checks the imported response definition; step 4 creates the authenticated connection and tests a real request. In the Azure portal, open the Function App's App keys page and create or copy a dedicated host key for this lab connection. Return to the connector's Test page, select New connection, enter that Function host key, and create the connection. Select the new connection (or your existing lab connection) and refresh the connection list if necessary. Under ClassifySupportInquiryV3, turn Raw Body on and paste the JSON below. Keep attachments as an empty array and routingOptionsJson as a JSON-encoded string. Select Test operation. In Response, confirm Status (200), schemaVersion = 2.0, a classification object, and Schema validation > Validation succeeded. Microsoft documents the current import and test flow in Create a custom connector from an OpenAPI definition. { "subject": "NovaScan X2 error E42 - France", "body": "NovaScan X2 serial NSX2-2407138 shows error E42 during startup after software version 4.2.1 was installed.", "from": "claire.martin@alpine-distribution.example.test", "attachments": [], "routingOptionsJson": "[{\"routingKey\":\"eu_distributor_support\",\"teamName\":\"EU Distributor Support\",\"description\":\"Technical support for NovaScan products in Europe.\"}]" } Important: The Function key belongs in the secure Power Platform connection. Do not put it in a text environment variable, flow action, screenshot, or source file. The September 8, 2026 request returned technical_support, France, confidenceBand = high, no risk flags, and recommendedRoutingKey = eu_distributor_support. This verifies one connector request; the flow still needs to apply its acceptance gate and obtain human approval. Add the components to the solution Open Solutions and create RW Global Support Intake . Create or select a publisher with prefix rw so the environment-variable schema names below match. Open the solution, select Objects, and use the object-type tree to review its components. Select Add existing, add the RW Support Classifier custom connector, and confirm that Custom connectors (1) appears in the object tree. Select New > More > Connection Reference and create these four named references: Office 365 Outlook SharePoint Standard approvals RW Support Classifier Select New > More > Environment variable and create these five variables: Display name Example schema name Current value Support Mailbox rw_SupportMailbox Shared mailbox address SharePoint Site URL rw_SharePointSiteURL RW Lab site URL Default Approver Email rw_DefaultApproverEmail Lab reviewer address Cases List Name rw_CasesListName Cases Support Teams List Name rw_SupportTeamsListName SupportTeams In Objects, select Connection references, Custom connectors, and Environment variables in turn. Confirm that the four required references, one connector, and five variables are present before creating the flow. The lab solution contains the four required connector types and additional automatically generated Outlook and SharePoint references. Match each flow action to its intended connection; the extra rows are not additional connector types to create. The five variables in the table are present. Routing Rules List Name is retained from an earlier lab design and is not required by this AI-recommendation workflow. Step 3: Build and validate the support email flow In this step, you build the flow that turns the bounded AI output from Step 2 into a validated recommendation. The flow checks duplicates, resolves an active team, applies the acceptance gate, and records uncertain or unavailable-classifier requests for human review. You then test both the normal validation path and the exception paths before adding approval in Step 4. Services and tools used in this step Service or tool Purpose Power Automate Build the automated cloud flow, scopes, conditions, and state transitions. Office 365 Outlook connector Start the flow when a message arrives in the shared mailbox. SharePoint connector Check for duplicates, load active support teams, and create review or fallback case rows. RW Support Classifier custom connector Send the message and active routing catalog to the Azure Function classifier. Create and configure the flow Use the action names shown below before writing expressions. In expression references, spaces become underscores: GetItems ExistingCase is referenced as GetItems_ExistingCase . Rename the custom connector action Classify Support Inquiry so its reference is Classify_Support_Inquiry . Select the matching dynamic-content token if your designer generated a different internal name. Enter formulas in the Expression editor without a leading @ . Values containing @{...} below are inline expressions in a text field. In this lab, enter the documented SharePoint list names directly in the connector; the list-name environment variables record configuration but are not automatically substituted into every action. Open RW Global Support Intake , select Objects > New > Automation > Cloud flow > Automated, and create RW - Global Support Intake - v1 . Add When a new email arrives in a shared mailbox (V2) and configure the trigger: Trigger field Value Original Mailbox Address rw_SupportMailbox current value Folder Inbox Importance Any Only with Attachments No Include Attachments No In Trigger NewSharedMailboxEmail > Parameters, select your shared mailbox under Original Mailbox Address. Open Advanced parameters to expose the four settings shown above: Importance = Any, Only with Attachments = No, Include Attachments = No, and Folder = Inbox. The September 9 lab capture displays an onmicrosoft.com address in the mailbox picker; select the mailbox you configured in Step 1 rather than copying the lab address. Open the trigger's Settings, turn on Concurrency Control, and set Degree of Parallelism to 1 . This serializes the lab runs, including time spent waiting for approval. A pending approval can delay the next email for 30 minutes. Use this setting for controlled tests; a design that processes new emails separately from approvals is needed before evaluating throughput. The current designer labels the concurrency switch Limit. The screenshot was captured at 125% browser zoom and cropped to the settings panel so the switch and value remain readable. After the trigger, select Add an action > Variable > Initialize variable for each row below. Keep every initializer above the Try scope and in the listed order: Variable Type Initial value CaseId String concat('AST-',formatDateTime(utcNow(),'yyyyMMdd'),'-',toUpper(substring(guid(),0,8))) InquiryType String Empty Product String Empty Country String Empty RecommendedTeam String Global Support ApproverEmail String rw_DefaultApproverEmail current value SLAHours Integer 8 RecommendedRoutingKey String Empty RoutingReason String Empty ConfidenceBand String Empty RiskFlags String Empty MissingFields String Empty For CaseId, enter the expression from the table through the Value field's expression editor, then select Add (or Update when editing an existing expression). The purple concat(...) token confirms that Value contains an expression. Select that token to reopen and check the full formula, as shown above. For SLAHours, set Name = SLAHours, Type = Integer, and Value = 8, as shown below. The other eleven variables in the table use String. Repeat the same Name, Type, and Value fields for each variable; where the table says Empty, leave Value blank. Keep all twelve initializers between the trigger and Scope Try. Compare the action order in these two views. The first red box contains CaseId through ApproverEmail; the second continues with SLAHours through MissingFields. Keep all twelve initializers outside and before Scope Try. These configuration views show placement; use the table above for each variable's type and initial value. Select Add an action > Control > Scope twice. Rename the first scope Scope Try and the second Scope Catch . On Scope Catch , select Configure run after and select has timed out, is skipped, and has failed for Scope Try . Inside Catch, add Send an email (V2) to the default approver with the CaseId and workflow run name, followed by Terminate with Status = Failed, Code = RW_INTAKE_FAILED , and an internal-only error message. In the current designer, open Settings > Run after, then expand Scope Try to reveal its result checkboxes. Leave Is successful unchecked. These are alternative results: any one of the three selected results allows Catch to run. Open SendEmail InternalFailure > Parameters. Set To to your internal default approver. In Subject, enter [RW Support Flow] Failed run followed by the CaseId variable token. In Body, write a short failure notice, add the CaseId token, and insert the expression workflow()?['run']?['name'] for the run reference so the reviewer can locate the failed run. For a new build, enter ordinary text in the rich-text Body editor, such as The support email flow failed., followed by the case ID, run reference, and Review the flow run before retrying. Do not paste HTML tags into the rich-text editor. The recorded configuration below contains escaped HTML tags as literal text; this is an existing formatting issue, not the recommended body format. This capture documents the saved inputs; the flow was not changed or rerun. Select Terminate Failed > Parameters and set Status to Failed and Code to RW_INTAKE_FAILED . For Message, use Global Support Intake failed; see the internal notification for the run reference. The canvas on the right shows where this action belongs: immediately after the internal notification inside Scope Catch. Note: The current Catch scope sends an internal notification and fails the run. It does not create a new Cases row or update an existing row to Failed. Add the duplicate guard Inside Scope Try , add SharePoint > Get items and rename it GetItems ExistingCase . Get items field Value Site Address rw_SharePointSiteURL current value List Name Cases Filter Query SourceMessageId eq '@{triggerOutputs()?['body/id']}' Top Count 1 In GetItems ExistingCase > Parameters, choose your site and the Cases list. Under Advanced parameters, enable Filter Query and Top Count. Keep the trigger's Message Id token inside the single quotes in the filter, and enter 1 for Top Count, as highlighted in this September 9 configuration capture. Add Data Operation > Compose, rename it Compose ExistingCaseCount , and use length(body('GetItems_ExistingCase')?['value']) . Open the Inputs expression editor in Compose ExistingCaseCount, enter length(body('GetItems_ExistingCase')?['value']), and apply it. This September 9 capture shows the existing expression opened for inspection, so the button is labeled Update. 3. Add a Condition, rename it Condition Duplicate , and test whether the Compose output is greater than 0 . In Condition Duplicate > Parameters, select Outputs from Compose ExistingCaseCount on the left, choose is greater than, and enter 0 on the right. The operator menu is open in this September 9 capture so its full label is visible. 4. In the True branch, add Terminate with Status = Succeeded. Leave case creation out of this branch. Build the remaining email-processing actions in the False branch. Place Terminate Duplicate in the True branch and set Status to Succeeded. Continue with GetItems ActiveSupportTeams and Select RoutingOptions in the False branch. This September 9 configuration capture shows both paths beside the termination setting. Load the catalog and call the classifier In the duplicate condition's False branch, add SharePoint > Get items and rename it GetItems ActiveSupportTeams . Get items field Value Site Address rw_SharePointSiteURL current value List Name SupportTeams Filter Query Active eq 1 Top Count 100 Select SupportTeams as the list. Under Advanced parameters, set Filter Query to Active eq 1 and Top Count to 100. This September 9 capture shows the active catalog query. Add Data Operation > Select, rename it Select RoutingOptions , set From to the value output from GetItems ActiveSupportTeams , and create this mapping: Key Value routingKey item()?['RoutingKey'] teamName item()?['TeamName'] description item()?['Title'] Use value from GetItems ActiveSupportTeams for From. Map routingKey to RoutingKey, teamName to TeamName, and description to Title. This September 9 capture shows the dynamic-content version of the expressions above. Add a Control > Scope, rename it Scope AI Recommendation , and place RW Support Classifier > Classify support inquiry inside it. Rename the classifier action Classify Support Inquiry and configure it: Input Value subject Subject from the mailbox trigger body Body from the mailbox trigger from From from the mailbox trigger attachments Empty array [] routingOptionsJson string(body('Select_RoutingOptions')) Select Subject, Body, and From from the mailbox trigger. Expand Advanced parameters to show from and attachments; keep the attachments array empty by adding no items. The September 9 designer labels these inputs with a Body/ prefix. For Body/routingOptionsJson, open the expression editor and enter string(body('Select_RoutingOptions')). Apply the expression with Add, or Update when editing an existing value as shown here. This converts the Select output array into the string expected by the connector. After the AI scope, add a Condition named Condition AIRecommendationAvailable . Use Configure run after so the condition runs when the AI scope succeeds, fails, is skipped, or times out. The True path is actions('Scope_AI_Recommendation')?['status'] equals Succeeded ; use the False path for the classifier-unavailable case. Open Settings > Run after, expand Scope AI Recommendation, and select all four statuses: Is successful, Has timed out, Is skipped, and Has failed. This September 9 configuration lets the next condition evaluate the AI scope even when the classifier is unavailable. Return to Parameters. In the left field, use the expression actions('Scope_AI_Recommendation')?['status']; select is equal to and enter Succeeded on the right. The empty row underneath is the designer's next-row placeholder. The True branch handles a successful AI call; the False branch handles an unavailable classifier. Validate the returned recommendation In the AI-available True branch, add Set variable actions in this order: Variable Value InquiryType Use the enum-to-choice mapping immediately below; unmatched values become Other Product coalesce(body('Classify_Support_Inquiry')?['classification']?['product'],'') Country coalesce(body('Classify_Support_Inquiry')?['classification']?['country'],'') RecommendedRoutingKey body('Classify_Support_Inquiry')?['classification']?['recommendedRoutingKey'] RoutingReason body('Classify_Support_Inquiry')?['classification']?['routingReason'] ConfidenceBand if(equals(body('Classify_Support_Inquiry')?['classification']?['confidenceBand'],'high'),'High',if(equals(body('Classify_Support_Inquiry')?['classification']?['confidenceBand'],'medium'),'Medium','Low')) RiskFlags join(body('Classify_Support_Inquiry')?['classification']?['riskFlags'],'; ') MissingFields join(body('Classify_Support_Inquiry')?['classification']?['missingFields'],'; ') For Set InquiryType , select the InquiryType variable, open the Value expression editor, and paste the following mapping. The recorded flow uses nested if expressions in this action: if(equals(body('Classify_Support_Inquiry')?['classification']?['inquiryType'],'technical_support'),'Technical Support', if(equals(body('Classify_Support_Inquiry')?['classification']?['inquiryType'],'quote_request'),'Quote', if(equals(body('Classify_Support_Inquiry')?['classification']?['inquiryType'],'demo_request'),'Demo', if(equals(body('Classify_Support_Inquiry')?['classification']?['inquiryType'],'product_information'),'Product Information', if(equals(body('Classify_Support_Inquiry')?['classification']?['inquiryType'],'partnership'),'Partnership', if(equals(body('Classify_Support_Inquiry')?['classification']?['inquiryType'],'complaint'),'Complaint','Other')))))) Select Add for a new expression, or Update when editing an existing one. The mapping is: Classifier value SharePoint Choice label technical_support Technical Support quote_request Quote demo_request Demo product_information Product Information partnership Partnership complaint Complaint Default, including other Other Add SharePoint > Get items, rename it GetItems RecommendedSupportTeam , and configure it: Get items field Value Site Address rw_SharePointSiteURL current value List Name SupportTeams Filter Query RoutingKey eq '@{variables('RecommendedRoutingKey')}' and Active eq 1 Top Count 1 Open GetItems RecommendedSupportTeam > Parameters. Select your SharePoint site and SupportTeams list. Under Advanced parameters, show Filter Query and Top Count. Insert the RecommendedRoutingKey variable between single quotes in the filter, retain and Active eq 1, and set Top Count to 1. Add Compose, rename it Compose TeamMatchCount , and use length(body('GetItems_RecommendedSupportTeam')?['value']) . In Compose TeamMatchCount > Parameters > Inputs, open the Expression editor and enter the expression above. Select Add for a new expression, or Update when editing an existing one. The result counts the rows returned by GetItems RecommendedSupportTeam. 4. Add Condition RecommendationAccepted and require every gate below: The acceptance gate requires all five checks: team match count = 1 AND ConfidenceBand = High AND riskFlags length = 0 AND missingFields length = 0 AND RecommendedRoutingKey is not empty Open Condition RecommendationAccepted > Parameters and use And to require all five checks. The September 9 lab screen above uses a team count greater than 0; the expression below uses a count equal to 1. With Top Count = 1, these checks have the same result, and neither detects duplicate catalog rows: enforce the unique RoutingKey constraint in SharePoint. The final empty row is the designer's next-row placeholder. Use this expression for the gate and compare its output with the Boolean true : and( equals(outputs('Compose_TeamMatchCount'),1), equals(variables('ConfidenceBand'),'High'), empty(body('Classify_Support_Inquiry')?['classification']?['riskFlags']), empty(body('Classify_Support_Inquiry')?['classification']?['missingFields']), not(empty(variables('RecommendedRoutingKey'))) ) Enable the unique RoutingKey constraint before relying on one returned row as an unambiguous match. The September 9 configuration check found this constraint disabled in the existing lab. Its recorded results therefore do not demonstrate protection against duplicate catalog keys; the new-build instructions in Step 1 require that protection. The model's confidence category is not a calibrated probability. Also, this sample holds every non-empty risk list, including product-alias and non-English flags. Those flags do not all indicate danger; the broad hold is a conservative lab policy with a review-volume tradeoff. In the accepted True branch, set RecommendedTeam to first(body('GetItems_RecommendedSupportTeam')?['value'])?['TeamName'] , ApproverEmail to first(body('GetItems_RecommendedSupportTeam')?['value'])?['ApproverEmail'] , and SLAHours to int(first(body('GetItems_RecommendedSupportTeam')?['value'])?['SLAHours']) . Leave space after these actions; Step 4 adds case creation, approval, and draft creation to this branch. In the False branch, add SharePoint > Create item named CreateItem NeedsReview . Store the AI fields, set Status = Needs Review, AutomationStatus = Review, and ClassificationSource = AI recommendation held for human review . Do not add an approval action. In the AI-available condition's False branch, add Create item named CreateItem ClassifierUnavailable . Use the fallback Global Support values, set InquiryType = Other, ConfidenceBand = Low, Status = Needs Review, AutomationStatus = Review, AIProposal = AI classifier unavailable or returned an invalid response. , and ClassificationSource = Classifier unavailable / human review . In CreateItem ClassifierUnavailable, set Summary to the mailbox trigger's Subject. The red boxes below show the fallback fields in the existing action. Under Advanced parameters, also add ConfidenceBand Value and select Low explicitly. The captured action omits that field and therefore inherits the existing list default, High; the screenshot does not show the recommended Low setting. Leave other classifier-derived fields empty and do not reference the failed classifier action's body. The matched SupportTeams row supplies RecommendedTeam, ApproverEmail, and SLAHours on the accepted recommendation path. Review rows created in this step store the AI reason and proposal; Step 4 stores the same evidence when it creates an accepted case. In both paths, the active catalog remains the operational allow-list. The current flow also initializes fallback values before the Try scope: Global Support , the default approver, and an eight-hour SLA. An invalid-key or classifier-unavailable review case can therefore retain fallback reviewer and SLA values without resolving a matching SupportTeams row. This is the actual lab behavior, not an additional validated assignment. Map the case fields consistently Use the following mapping on both review-case creation actions and on CreateItem Case in Step 4. Then apply each branch's Status, AutomationStatus, ClassificationSource, and AIProposal values. For the unavailable-classifier branch, use the trigger Subject for Summary, leave other classifier-derived fields empty, and write the explicit fallback message; do not reference the failed action's body. Cases field Value Title variables('CaseId') SourceMessageId Message Id dynamic content from the trigger, the same value used by the duplicate guard ReceivedAt Date Time Received dynamic content from the trigger RequesterEmail Sender's email address from the trigger, without the display name EmailSubject Subject from the trigger InquiryType, Product, Country Corresponding variables; fallback InquiryType is Other RecommendedRoutingKey, RoutingReason, RiskFlags, MissingFields Corresponding variables ConfidenceBand Corresponding variable, or explicit Low for classifier unavailable RecommendedTeam variables('RecommendedTeam') for accepted and classifier-unavailable cases; use the branch-specific expression below for NeedsReview AssignedToText variables('ApproverEmail') DueAt addHours(utcNow(),variables('SLAHours')) Summary body('Classify_Support_Inquiry')?['classification']?['summary'] when available AIProposal string(body('Classify_Support_Inquiry')?['classification']) when available LastAutomationRun utcNow() For CreateItem NeedsReview, set RecommendedTeam to if(greater(outputs('Compose_TeamMatchCount'),0),first(body('GetItems_RecommendedSupportTeam')?['value'])?['TeamName'],'Unvalidated recommendation') . This preserves the catalog team name when a match exists and records Unvalidated recommendation when none exists. The review branch does not run the accepted branch's Set ApproverEmail or Set SLAHours actions, so AssignedToText and DueAt retain the initialized default reviewer and eight-hour SLA. A displayed catalog team name does not mean the recommendation passed the acceptance gate. This mapping was checked in the existing configuration; no new run was performed. On every later Update item, use the ID returned by that branch's Create item and preserve Title = CaseId. Do not accidentally use the source email ID as the SharePoint item ID. DueAt is a simple target timestamp from processing time; this tutorial does not implement business calendars or SLA escalation. Check the flow Save the flow. Open Flow checker and resolve every reported error or warning before testing. At this stage, Flow checker should report zero errors and zero warnings before you send the validation messages: The highlighted toolbar button opens Flow checker. This designer check reports zero errors and warnings in the captured flow; it does not verify approval delivery, mailbox access, or runtime outcomes. Validate those paths with the test cases below. Test the normal validation path Send this complete synthetic inquiry to the shared mailbox: Subject: NovaScan X2 error E42 - France - software version 4.2.1 Hello Aster Support, We are a distributor in France. NovaScan X2 serial NSX2-2407138 shows error E42 during startup after installing software version 4.2.1. We captured the error screenshot. Regards, Claire In Power Automate, open Solutions > RW Global Support Intake > Objects > Cloud flows > RW - Global Support Intake - v1. Under 28-day run history, select the new run and confirm that the duplicate guard, classifier, active-team lookup, and Condition RecommendationAccepted succeeded. Inspect the classifier and variable actions in the run. Confirm that the flow produced Technical Support, NovaScan X2, France, High confidence, no risk flags or missing fields, eu_distributor_support , and a successful match to EU Distributor Support. Confirm that the acceptance condition followed its True branch. At this point, the flow has validated the recommendation but has not created the accepted case or approval request. Step 4 adds those actions to the True branch. This separation lets you verify the routing boundary before introducing consequential workflow actions. Test one representative review path A green normal run does not prove that uncertainty is held safely. Use one intentionally incomplete message to exercise the review boundary without repeating every exception as a full walkthrough. Send this intentionally incomplete message to the shared mailbox: Subject: Help needed - low confidence routing test A device stopped working somewhere. Please route this request. Open RW Lab > Cases > Tutorial Evidence, locate the newly created row by its subject and received time, and record its generated CaseId. Confirm that Product and Country are empty, RecommendedTeam = Global Support, ConfidenceBand = Low, RiskFlags = Low confidence, and Status = Needs Review. Open the corresponding flow run, expand the action groups, and confirm that Condition RecommendationAccepted followed the False branch and CreateItem NeedsReview succeeded. The screenshot below shows a separate prompt-injection case held in review. It illustrates the risk-flag gate, rather than the incomplete-message test just described. Your generated CaseIds will be different. Record the remaining Step 3 checks Run the remaining checks separately, but summarize them rather than repeating the same send–open run–open row sequence. Record both the validation surface and the observed boundary: Check Validation surface Observed boundary Structured prompt-injection risk Live mailbox and flow run Security Review recommendation remained Needs Review because riskFlags was non-empty. Azure content filter or invalid classifier response Live mailbox and flow run CreateItem ClassifierUnavailable stored the fallback proposal and Needs Review / Review . Invalid routing key Function test plus flow-gate inspection invented_team was forced to Low confidence; a zero-row team lookup cannot pass the acceptance gate. Duplicate message ID Resubmitted completed run Terminate Duplicate succeeded and no second Cases row was created. Note: Historical screenshots may show High on a classifier-unavailable row because the original list defaulted to High. The instructions above explicitly use Low for new fallback rows. ClassificationSource identifies the unavailable-model case; Low here is a conservative fallback value, not a model assessment. Step 4: Add human approval, create a draft, and evaluate the workflow In this step, you complete the accepted recommendation path. A reviewer decides whether to accept the proposed team, and an approval may create an editable acknowledgement draft, but no branch sends that draft to the requester. You walk through the approved path and the actual PT30M timeout boundary, then summarize the remaining stateful checks and classifier evaluation. Services and tools used in this step Service or tool Purpose Power Automate and Approvals Present the AI recommendation to the reviewer and branch on the human decision. SharePoint connector Create the accepted case and update its approval and automation states. Office 365 Outlook connector and Outlook Create an editable acknowledgement draft and confirm that no message was sent automatically. Azure Functions and Microsoft Foundry Run the authenticated classifier evaluation against the deployed implementation. Note: The local evaluation tools used later in this step are Node.js 22 and PowerShell. Configure approval and case-state updates In the accepted-recommendation branch, add SharePoint > Create item and rename it CreateItem Case . Map the trigger metadata and classifier fields to the corresponding Cases columns, then set these operational fields: Cases field Value RecommendedTeam RecommendedTeam variable AssignedToText ApproverEmail variable DueAt addHours(utcNow(),variables('SLAHours')) AIProposal string(body('Classify_Support_Inquiry')?['classification']) ClassificationSource AI recommendation validated by SupportTeams Status New AutomationStatus Processing Open CreateItem Case > Parameters and select your SharePoint site and Cases list. Under Advanced parameters, set Title to the CaseId variable. Select Message Id, From, and Subject from the mailbox trigger for SourceMessageId, RequesterEmail, and EmailSubject, respectively. Continue with the classifier and operational fields in the tables above. Scroll down within CreateItem Case > Parameters. Set RecommendedTeam to the RecommendedTeam variable and AssignedToText to ApproverEmail. For ReceivedAt, select the trigger's Received Time token (the label may appear as Date Time Received). Enter addHours(utcNow(),variables('SLAHours')) as the DueAt expression and utcNow() as LastAutomationRun. Select the classifier's corresponding Product, Country, and Summary outputs for those fields, and the InquiryType variable for InquiryType Value. Continue down the panel to configure AIProposal, ClassificationSource, Status, and AutomationStatus from the table above. Continue down CreateItem Case > Parameters. Set Status Value to New and AutomationStatus Value to Processing. For AIProposal, open the expression editor, enter string(body('Classify_Support_Inquiry')?['classification']), and select Add (or Update for an existing expression). This stores the classifier's classification object as text. The collapsed token in the screenshot displays string(...); select it to inspect the full expression. Enter AI recommendation validated by SupportTeams for ClassificationSource. Select the RecommendedRoutingKey and RoutingReason variables from dynamic content for their corresponding fields; the purple tokens are variable values, not literal text. At the bottom of CreateItem Case > Parameters, select the RiskFlags and MissingFields variables from dynamic content for their matching fields. For ConfidenceBand Value, select the ConfidenceBand variable as a custom value. The purple tokens shown here are variable values; do not type their names as plain text. This screenshot belongs to the accepted-recommendation branch. For the classifier-unavailable branch, use the explicit fallback values described in Step 3. Add SharePoint > Update item named UpdateItem AwaitingApproval . Use the ID returned by CreateItem Case , keep the same Title, set Status = Awaiting Approval and AutomationStatus = Processing, and update LastAutomationRun with utcNow() . Open UpdateItem AwaitingApproval > Parameters and select the same SharePoint site and Cases list. Select ID from CreateItem Case for Id. Under Advanced parameters, keep Title set to the CaseId variable, enter utcNow() for LastAutomationRun, and select Awaiting Approval for Status Value and Processing for AutomationStatus Value. This records the waiting state before the approval request starts. Add Start and wait for an approval, rename it Approval Assignment , and set the following values. Configure Timeout under the action's Settings: Setting Value Approval type Approve/Reject - First to respond Assigned to ApproverEmail from the validated SupportTeams row Timeout PT30M for the lab Title [CaseId] Review AI-recommended assignment to RecommendedTeam Under Parameters, select Approve/Reject - First to respond. Build Title with the CaseId and RecommendedTeam variables selected from dynamic content, and select the ApproverEmail variable for Assigned to. The purple tokens represent variable values; do not type the variable names as plain text. Open Approval Assignment > Settings > General, then enter PT30M in Action timeout. This sets the approval wait to 30 minutes. The separate Run after setting in step 10 determines which action handles that timeout. Include the product, country, recommended team, routing key, AI reason, confidence, risk flags, and original subject in the approval details. State the effect of each decision explicitly: Approve the AI recommendation to create an acknowledgement draft. Reject to keep the case in human review. In Approval Assignment > Parameters, scroll to Details and insert the dynamic-content tokens alongside their labels as shown. The September 9 configuration capture shows the complete reviewer message, including the final decision instructions. After the approval action, add a Condition named Condition Approved . Keep run after = is successful only and test body('Approval_Assignment')?['outcome'] equals Approve . The timeout path in step 10 must be a parallel branch from the approval action, not an action after this condition. Select the approval action's Outcome dynamic content in the left field (shown as body/outcome ), choose is equal to, and enter Approve in the right field. The highlighted row is the comparison to configure. Open Settings > Run after, expand Approval Assignment, and select only Is successful. This checks whether the approval action completed; the Outcome comparison above checks the reviewer's decision. A completed rejection follows the condition's False branch. A timed-out approval follows the separate timeout branch in step 10. In the approved branch, update the case to Assigned, record ApprovalOutcome = Approve and ApprovalComment = coalesce(first(body('Approval_Assignment')?['responses'])?['comments'],'') , and keep AutomationStatus = Processing. Open UpdateItem Approved in the True branch of Condition Approved. Select Cases, use ID from CreateItem Case for Id, and preserve Title = CaseId. Under Advanced parameters, select Outcome from Approval Assignment for ApprovalOutcome, use utcNow() for LastAutomationRun, and enter the blank-safe ApprovalComment expression above. The captured lab configuration shows its original first(...) token. Choose Assigned for Status Value and Processing for AutomationStatus Value; the later UpdateItem DraftRecorded action records successful draft creation. Add Office 365 Outlook > Draft an email message and rename it Draft Acknowledgement . Configure it: Draft field Value To From from the mailbox trigger Subject concat('We received your support request [',variables('CaseId'),']') From Shared mailbox address Importance Normal Body Use the template below, replacing bracketed values with dynamic-content tokens Use this draft body: Hello, We received your request. A reviewer approved the proposed assignment to our support team. Case: [CaseId] Team: [RecommendedTeam] This is a draft created for human review. It has not been sent automatically. Select the trigger's From token for To. Insert the CaseId variable into the subject and the CaseId and RecommendedTeam variables into the body; do not type the bracketed placeholders literally. The screenshot uses text plus a CaseId token for the subject, equivalent to the expression in the table. Verify the recipient, team, and wording before sending the draft manually. In Draft Acknowledgement, scroll down to Advanced parameters, select From (Send as), and enter the shared mailbox address configured in Step 1. The September 9 configuration capture below highlights this field. The From field selects the sender identity; it is not a destination-folder setting. Inspect Drafts for the account used by the Outlook connection and confirm the displayed From address. Do not assume that setting From to the shared mailbox also stores the draft in that shared mailbox. See the Office 365 Outlook connector reference. Add another Update item, rename it UpdateItem DraftRecorded , write the draft action's Id to DraftMessageId, set AutomationStatus = Success, and update LastAutomationRun. Select Cases and use ID from CreateItem Case for Id. Under Advanced parameters, preserve Title = CaseId, select the Id output from Draft Acknowledgement for DraftMessageId (displayed as body/Id), set LastAutomationRun to the expression utcNow(), and choose Success for AutomationStatus Value. The SharePoint item ID and the Outlook draft ID refer to different records; select each token from its corresponding action. In the rejected branch, update the case to Needs Review, record the approval outcome and comment, and set AutomationStatus = Review without creating a draft. Open UpdateItem Rejected in the False branch of Condition Approved. Select Cases, use ID from CreateItem Case for Id, and preserve Title = CaseId. Under Advanced parameters, select Outcome from Approval Assignment for ApprovalOutcome, use utcNow() for LastAutomationRun, and record the first approval response's comments in ApprovalComment. Use the blank-safe comments expression from item 6 above; the captured lab configuration shows its original first(...) token. Choose Needs Review for Status Value and Review for AutomationStatus Value. Add a parallel branch directly from Approval Assignment with a separate UpdateItem ApprovalTimedOut action. Use the SharePoint ID from CreateItem Case and preserve Title = CaseId. Select Configure run after > has timed out, then set Status = Needs Review, ApprovalOutcome = Timeout, and AutomationStatus = Review without creating a draft. Open Settings, scroll to Run after, and expand Approval Assignment. Select only Has timed out; leave the other three results unchecked. Return to Parameters. Use ID from CreateItem Case for Id and the CaseId variable for Title. Set ApprovalOutcome to Timeout, LastAutomationRun to the expression utcNow(), and ApprovalComment to No response was received before the approval timeout. Choose Needs Review for Status Value and Review for AutomationStatus Value. The timeout update protects the case record, but a timed-out approval can still mark its parent scope as failed. If Scope Catch is configured to run whenever that parent scope fails, it will also send the internal failure notification and may leave the overall run in a Failed state. The validation below preserves that observed behavior. The troubleshooting section provides an explicit handled-timeout exit to test as an improvement; it is not represented by the historical screenshots. Important: Do not add Send a draft message or another send action. Approval in this tutorial authorizes draft creation only. Save the flow and run Flow checker again. Resolve every error or warning before the approval test. At this point your accepted branch should contain the case, approval, outcome, and draft actions added in this step. Compare your action placement with this configuration view. The left red box contains the three approved-path actions in order; the middle box contains only the rejection update. The timeout update is outside the condition and connects directly to Approval Assignment with its timeout run-after setting. This screenshot shows the recorded configuration, before the handled-timeout improvement described below. Test approval and draft-only behavior Send the complete NovaScan X2 inquiry from Step 3 to the shared mailbox again as a new message. Do not use Resubmit for this test because the duplicate guard intentionally stops a replay with the same SourceMessageId . Open the new flow run, confirm that Condition RecommendationAccepted follows the True branch, and record the generated CaseId from CreateItem Case . In Power Automate, select Approvals > Received, then open the request for <CASE_ID> . Confirm that it displays EU Distributor Support, eu_distributor_support , the AI reason, High confidence, and an empty risk list. This Outlook capture, taken on September 9, shows the approval-request email dated August 8, 2026 for historical case AST-20260808-E11790B1 . The left results list keeps the request and its acknowledgement Draft together. The red rectangles identify the proposed team, routing key, AI reason, confidence, risk flags, and decision buttons. Personal details are masked. This is the request message; use approval history to verify the completed decision. Select Approve, enter a synthetic reviewer comment, and submit the response. After completion, select History and confirm that the request shows Outcome = Approved. Open RW Lab > Cases > Tutorial Evidence, select <CASE_ID> , and confirm that it records Assigned, ApprovalOutcome = Approve, AutomationStatus = Success, and a non-empty DraftMessageId. The approved-case screenshot at the beginning of this article shows a historical lab result. Use the CaseId generated by your own run for the remaining checks. Open Outlook for the account used by the Outlook connection, expand the navigation pane, select Drafts, and open We received your support request [<CASE_ID>] . Confirm that the acknowledgement remains editable, names the case and approved team, and states that it was not sent automatically. The acknowledgement for the historical case is open in the Outlook compose view. Check the editable subject and body, the case ID, and EU Distributor Support. The body states that the draft was created for human review and has not been sent automatically. The recipient is masked. This September 9 capture shows the existing draft reopened for inspection; the visible 5:06 PM saved time is not evidence of its original August 8 creation time. Verify Sent Items separately in the next step. In the Outlook navigation pane, select Sent Items, enter <CASE_ID> in the search box, and confirm that the search reports no sent acknowledgement. The approved test establishes the human decision and no-send boundary. Test the PT30M timeout boundary Send the complete NovaScan X2 inquiry again as a new message. Add a unique prefix such as [TIMEOUT-PT30M-01] to the subject so you can distinguish the run, and do not use Resubmit. Confirm that the run reaches Approval Assignment , record the generated CaseId and approval start time, and do not approve or reject the request. Wait at least 30 minutes. Do not shorten the action timeout for this evidence run; the purpose is to test the same PT30M value configured in the flow. Open the completed run and confirm that Approval Assignment shows TimedOut and UpdateItem ApprovalTimedOut succeeded. Open the corresponding Cases row and confirm Status = Needs Review, ApprovalOutcome = Timeout, AutomationStatus = Review, and an empty DraftMessageId. Search Outlook Drafts and Sent Items for the CaseId and confirm that neither contains an acknowledgement for the timed-out request. The tested PT30M run produced CaseId AST-20260808-F1A058C2 . The approval timed out after 30 minutes, the timeout update succeeded, and the case retained Needs Review / Timeout / Review with no draft ID. No customer acknowledgement was found in Drafts or Sent Items. Warning: The same run exposed a control-flow issue. Approval Assignment timed out inside Scope Try , so the parent scope was marked Failed even though UpdateItem ApprovalTimedOut succeeded. Scope Catch then sent the internal failure notification, and Terminate Failed left the overall run in a Failed state. The case is safely held for review, but this timeout is not yet normalized as a handled outcome. To avoid a false failure alert, isolate the approval timeout from the catch condition or add an explicit handled-timeout exit before enabling the flow for production. This is the only long-running walkthrough in the tutorial. The remaining stateful outcomes are summarized below. Record the remaining stateful checks Check Evidence Observed result Approved recommendation Approval history, Cases, Drafts, and Sent Items Assigned / Success ; draft ID recorded; editable draft created; nothing sent. Low-confidence recommendation Cases and flow run Needs Review / Review ; no approval or draft. Classifier unavailable Cases and flow run Fallback proposal stored as Needs Review / Review ; no approval or draft. Reviewer rejection Approval history, Cases, Drafts, and Sent Items Needs Review / Review ; no matching draft or sent message. Duplicate replay Flow run and Cases Duplicate termination succeeded; no second row. Approval timeout Approval action, timeout update, Cases, Drafts, and Sent Items Approval timed out; timeout update succeeded; Needs Review / Timeout / Review ; no acknowledgement. Overall run Failed and sent the internal failure notification because Catch also handled the timed-out parent scope. Compare Status, AutomationStatus, ConfidenceBand, and ApprovalOutcome for the recorded August 8 cases: low confidence, approved, classifier fallback, rejected, another low-confidence case, and timeout. The fallback row retains the historical High default; these records do not show a new workflow run or prove that the proposed Low default was deployed. To check the recorded approved case, open Sent Items in the Outlook account used by the flow connection and search for AST-20260808-E11790B1. The September 9 recapture below shows Nothing found in Sent Items. Outlook expands the search to other folders after finding no match in Sent Items; those additional results are outside this crop. This checks the existing August 8 case, not a new workflow run. Evaluate the completed workflow Keep AI evaluation separate from flow-state tests. The GitHub sample contains 20 synthetic inquiries. Eighteen are evaluated directly against the authenticated classifier. Exact-message replay is tested separately against the flow. The existing-case follow-up fixture is also excluded from the classifier run, but automatic follow-up merging is not implemented in this walkthrough; do not count that fixture as a passed flow capability. Return to the sample root (if you are in azure/global-support-classifier , run cd ../.. ) and run the fixture validator: node --test tests/aster-imaging-fixtures.test.mjs Set the authenticated classifier endpoint and Function key only in the current terminal session. In PowerShell: $env:RW_CLASSIFIER_URL='https://<FUNCTION_APP_NAME>.azurewebsites.net/api/classify-support-inquiry' $env:RW_CLASSIFIER_FUNCTION_KEY='<FUNCTION_KEY>' Run the authenticated classifier regression: node tests/global-support-classifier-regression.mjs Review the aggregate result. The final authenticated regression produced: Measure Result Fixture count 20 Classifier cases evaluated 18 Excluded from classifier evaluation 2: replay and follow-up Cases passing every evaluator check 12 of 18 (66.7%) Product match rate, returned classifications 100% Country match rate, returned classifications 100% Inquiry-type match rate, returned classifications 100% Team-key match rate, returned classifications 100% Evaluator safety-scenario criterion Passed Low-confidence review boundary Passed Prototype criteria Passed Inspect the per-case differences instead of relying only on the aggregate result. In the authenticated rerun on August 8, 2026, using deployment gpt-5.4-mini pinned to version 2026-03-17 , 12 of the 18 classifier cases matched every evaluator check. Six did not pass every check: all six failed the summary-term check, one also differed on urgency, and one also differed on a missing-field expectation. The summary evaluator checks for specified terms; it does not require an exact sentence match. These differences need inspection and should not be dismissed as cosmetic without reviewing the outputs. Product, country, inquiry type, team recommendation, safety, and low-confidence criteria all passed. Keep those differences in the regression output rather than changing the benchmark around the model. The evaluator computes field match rates over responses that contain a classification, excluding failed HTTP calls. For the prompt-injection fixture, it treats any non-success HTTP response as a safe failure; that alone does not distinguish a content-filter refusal from a service outage. Its safety criterion also tolerates some inquiry-type, missing-field, and summary differences. Read the per-case output alongside the aggregate result. This small, known synthetic set is a regression check, not an independent or held-out accuracy study. The evaluator never calls Approvals or Outlook. Its safety result cannot establish that the cloud flow withheld an approval, blocked an attachment, or sent no email; those claims require the separate flow and mailbox checks above. A successful process exit means the configured prototype thresholds passed, even when some individual checks failed. Compare the classifier results and the stateful validation matrix with these synthetic prototype thresholds: product, country, inquiry type, and team recommendation accuracy are each at least 90%; directly evaluated safety, privacy, prompt-injection, and unsupported-attachment inputs satisfy the evaluator criteria; separately verify that a returned risky proposal cannot pass the live acceptance gate; every low-confidence or invalid-key result enters Needs Review; a duplicate message ID does not create a second case; a classifier failure creates a Needs Review row; an unexpected flow failure is visible in run history and sends the internal failure notification, but may not leave or update a Cases row; no test sends a customer-facing acknowledgement automatically. Clear the Function key from the terminal session when the run is complete: Remove-Item Env:RW_CLASSIFIER_FUNCTION_KEY Record the evaluation scope and limitations. These results are not production accuracy, security, capacity, or SLA claims. Re-evaluate with approved representative data, target-tenant policies, and an operational review process before production use. Troubleshooting Symptom Check or next action The custom connector cannot be created or used Confirm environment permissions, custom-connector entitlement, and data policies. Connector test returns 401 or 403 Check the connector host and Function key. For an upstream authorization error, separately check the Function managed identity and its role on the Foundry resource. Classifier returns a non-success response Inspect Function logs for the bounded upstream code, then check deployment name, quota, role propagation, and content-filter behavior. Keep the case in review. A valid recommendation goes to Needs Review Inspect all five gate inputs. In this lab, any missing field or risk flag causes a hold, even at High confidence. A later email appears delayed With trigger concurrency set to one, the previous run may still be waiting for approval. Complete or let that lab approval time out before testing another message. An expression cannot find an action Check the action's internal name and its nesting. Keep each expression inside a branch where its referenced action ran. Draft creation fails or the draft seems missing Check the Outlook connection account, mailbox delegation, and that account's Drafts folder. Setting From does not select the storage folder. A timeout case is in review but the run is Failed The generic Catch also observed the timed-out parent scope. See the handling pattern below. Treat a recorded timeout as a handled outcome The historical lab deliberately remains visible in the screenshots: the case update succeeded, but the run failed and sent an internal alert. For a lab where a recorded timeout should finish successfully, test this small change: On the timeout-only branch, keep UpdateItem ApprovalTimedOut after Approval Assignment with run after = has timed out. Immediately after that update, add Terminate, name it Terminate HandledTimeout , and set Status = Succeeded. Let it run only after the update succeeds. The case must be durably recorded as Needs Review / Timeout / Review before the successful exit. Keep the generic Catch for unexpected failures. Do not configure the successful exit to run after a failed or skipped case update, and do not place it on the normal approval path. Repeat the full PT30M test. Require the case's timeout state, an empty DraftMessageId, no acknowledgement in Drafts or Sent Items, no generic failure notification, and an overall Succeeded run. Separately verify that an unexpected SharePoint update failure still reaches Catch. This is a proposed correction to the recorded lab, not a newly verified tenant result. No post-correction screenshot or live run is included in this article. The pattern uses Power Automate's run-after and termination controls; validate it in your environment before relying on it. Production considerations Save incoming requests separately from the long-running approval process so pending reviews do not serialize new messages. Design recovery for partial completion: a Cases row can exist while draft creation or its ID update fails. Duplicate detection alone does not repair that case, and replaying blindly can create extra drafts. Replace tutorial-level SharePoint permissions with least-privilege role assignments and a documented ownership model. Confirm data residency, retention, DLP, audit, and connector policies. Store attachments separately and scan them before downstream processing. Pass only approved attachment metadata to the classifier if the unsupported-attachment gate is expected to protect the live flow. Add a Catch-path upsert if every unexpected failure must leave Status = Failed and AutomationStatus = Failed in Cases. Separate handled approval timeouts from unexpected failures so a successful timeout update does not also trigger the generic failure notification and failed termination. Use a supported secretless identity path for any AI service. Define operational ownership for rule changes, failed runs, approval timeouts, and mailbox delegation changes. Re-evaluate licensing and capacity for the target environment. Test with approved representative data before making any accuracy or SLA claim. Clean up the lab Turn off RW - Global Support Intake - v1 first and cancel any outstanding lab runs or approvals. Remove the dedicated connector connection and its Function key when they are no longer needed. Delete the dedicated Function app, model deployment, and associated lab storage or monitoring resources after retaining the synthetic evidence you want to keep. Delete a whole resource group only if it contains exclusively disposable lab resources. If you created the mailbox, SharePoint site, or Power Platform environment solely for this exercise, remove them through their respective admin tools when finished. Keep shared resources used by other work. Confirm that the remaining Azure resources and deployments match what you intend to retain.Introducing Inside Microsoft Foundry: Quickstart 🎬
Discover Inside Microsoft Foundry: Quickstart, a new video series for developers building AI agents. Starting with "What does it really take to ship an AI agent?", the series explores real-world challenges such as model selection, grounding agents in data, evaluation, deployment, observability, and governance. Follow along as we show how the Microsoft Foundry ecosystem helps developers move from prototype to production, with new episodes released in the coming weeks.Model Migration Process on Microsoft Foundry and Azure OpenAI
Every app built on an LLM will eventually move to a new model. The model you shipped may be retired, or a newer model may offer better quality, cost, or performance. Changing a model name in code from a retiring model such as gpt-4o to a newer one such as gpt-5.1 may take one line. That line hides a much larger migration. Model failures are often silent to the systems using the model and loud to users. Nothing crashes. Error rates stay flat. Every dashboard says the migration went fine. Meanwhile, responses change shape, summaries become longer and more hedged, JSON fields disappear, and tool calls fire in a different order. Users notice. Support queues grow. Downstream code that depended on the old behavior starts to break. A successful migration preserves the application's behavior or improves it in measurable ways. That requires a repeatable process to detect drift, adapt safely, and prove quality before broad rollout. The model migration process has six phases: Discover → Assess → Adapt → Validate → Roll out → Retire. This article explains what each phase looks like, which Microsoft Foundry tools support it today—including Azure OpenAI capabilities—and where teams still need to build around the platform. It then applies the process to a retail shopping assistant and points to additional resources in Go deeper at the end of this post! Why migrate now Every model has a retirement date. On Microsoft Foundry, generally available models typically ship with a retirement date about 18 months out, and older model families are actively replaced. For example, the model lifecycle and retirement schedule lists gpt-4o (2024-05-13) as retiring on October 1, 2026, with gpt-5.1 as its replacement. What happens at retirement depends on how you buy capacity: Standard, Global Standard, and Data Zone Standard (pay-as-you-go) deployments are auto-upgraded on a rolling, region-by-region schedule. You control the timing with versionUpgradeOption set to one of: OnceNewDefaultVersionAvailable, OnceCurrentVersionExpired, or NoAutoUpgrade. NoAutoUpgrade means the deployment stops working at retirement. Priority Processing follows the same path. Provisioned (PTU) deployments are not auto-upgraded. You migrate them yourself, either in-place (traffic moves over a 20–30 minute window with no downtime) or side-by-side (stand up the new deployment, test, shift traffic, delete the old one). Batch deployments follow the side-by-side path: deploy the new model, resubmit jobs, retire the old deployment. The developer problem is the same in every case: traffic eventually reaches a different model, but the platform cannot tell you whether the application still behaves as it did before. A responding endpoint does not prove that the app behaves correctly. A new model can change formatting, tone, tool-calling behavior, or JSON shape in ways that quietly break downstream code. When should I migrate? Start before the retirement date. Automatic upgrade handles the traffic transition for eligible deployments, but the team still owns behavioral validation. Provisioned deployments also require a manual migration. Microsoft typically makes a replacement available in Global Standard about 90 days before retirement, in provisioned regions about 30 days before retirement, and in standard regions about two weeks before retirement. That gives you time to evaluate the new model on your own terms. Retirement dates cannot be extended. You also do not need a deprecation notice to begin. If a newer model may improve quality, speed, or cost, run it through the process now. Waiting turns the switchover into a slow train wreck: responses drift, parsing becomes brittle, support tickets accumulate, and the team ends up debugging a model it did not choose on a date it did not pick. A deliberate migration makes the retirement date a formality and creates a process the team can reuse. Who this is for This process fits teams that own an LLM-powered feature inside a larger application and run migrations deliberately. It also applies to AI-native platform teams that centrally manage models for other application teams. The phases remain the same, though platform teams may run them faster and in parallel rather than in sequence. Fine-tuned workloads are out of scope here because they cannot be upgraded automatically, have separate training and deployment retirement schedules, and turn the Adapt phase into a distillation or retraining exercise rather than primarily prompt work. The six phases Phase Definition What success looks like Discover Learn that a model change is coming or needed. The team receives a timely, structured signal with the deprecation date, replacement model, and migration window. Assess Choose a target model and confirm that it is operationally available. The team understands the candidates and confirms capacity, region, and SKU before tuning starts. Adapt Replay the current workload on the new model, diagnose changes, and update prompts, parameters, tool definitions, output schemas, and calling code. The team runs side-by-side replay against real or representative traffic, can see the behavioral differences, and records every change. Validate Run the adapted workload against a quality rubric and decide whether it is safe to ship. The team has an evaluation suite that is affordable to run and trusted by application owners and reviewers. Roll out Promote the model through staged production exposure, monitor live behavior, and commit or roll back. Canary or weighted routing is in place, live quality is measured alongside latency and errors, and rollback remains possible. Retire Decommission the old deployment, free capacity, archive evaluation artifacts, and update internal documentation. The old SKU is gone, the deployment count falls, and the team carries what it learned into the next migration. Foundry tools at a glance Microsoft Foundry provide tools for each phase of the Model Migration Process. Phase Microsoft Foundry feature (including Azure OpenAI) Documentation Discover Model retirement schedule, lifecycle policy, Service Health alerts, and Models API lifecycleStatus Model retirement schedule Lifecycle policy Assess Model leaderboards and benchmarks for quality, safety, cost, throughput, and latency; trade-off charts; side-by-side comparison; suggested replacements Model leaderboards and benchmarks Side-by-side compare Adapt Prompt Optimizer in the Foundry Agent playground; agent optimization; simulator for synthetic data Prompt Optimizer Agent optimization Simulator Validate Azure AI Evaluation SDK with 30+ evaluators, LLM-as-judge, graders, and the portal evaluation wizard Azure AI Evaluation SDK Portal evaluation Roll out Automatic upgrade and versionUpgradeOption; provisioned in-place or side-by-side migration; continuous evaluation; Azure Monitor alerts Auto-upgrade with versionUpgradeOption Continuous evaluation Retire Models API to confirm 410 Gone; observability dashboard to track deployment count Models API Observability dashboard Breakdown of each phase 0. Prepare the test dataset Before starting the six phases, build a set of representative inputs, expected outputs, and agreed success criteria. This dataset gates the middle of the lifecycle: Adapt needs inputs for replay, and Validate needs ground truth and scoring criteria. Step 0 describes the workload rather than the candidate model, so it can begin during Discover, before the team selects a target. Build the dataset from captured production traffic or domain examples in .csv or .jsonl. If representative data is not available, use the simulator to generate synthetic inputs. Two practices determine whether this work pays off: Instrument capture before you need it. Production content capture is opt-in and never retroactive. Log prompts, responses, latency, and token counts now so the team has traffic to evaluate later. Freeze the dataset. Keep inputs, ground truths, and success criteria fixed throughout the migration. If they change, source and target results are no longer comparable. You also need an inventory of the model deployments your workload uses, including their deployment types (Standard, Provisioned, or Batch). For each source model, note its retirement date and suggested replacement from the Model retirement schedule. 1. Discover Discover begins when something forces the team to consider a model change: a deprecation notice, a new generally available model, a cost or latency problem, or a capability gap. The phase ends with a decision to begin migration or stay on the current model if it remains stable, performs well, and is not approaching retirement. Foundry tools. The model lifecycle and retirement schedule publishes retirement dates and suggested replacements. The Azure OpenAI model retirements documentation explains notification timing, including at least 60 days for generally available model retirements and at least 30 days for preview model retirements. It also explains how to configure Azure Service Health advisories and use the Models API for programmatic lifecycleStatus and deprecation checks. Those APIs provide the foundation for an internal discovery system. Where it breaks. Customers may learn about a retirement through email, a service health alert, or a production error. By the time the right team sees the signal, it may already be deep into the deprecation window and heading toward retirement. What your team provides. The schedule and Models API expose the data through a stable contract. Mature enterprises may add a thin notification layer that routes it to the right owners. 2. Assess The team chooses a candidate target model and confirms that it is usable: the correct region and SKU, enough quota, and availability alongside the current model so rollback remains possible. Assess also includes projecting monthly cost against historical traffic. Pricing structures change between model generations through reasoning tokens, cached input, structured-output overhead, and other factors. Those changes can move unit economics by 2x or more. For regulated workloads, compliance requirements such as BAA, FedRAMP, and regional Standard versus Global Standard availability may narrow the candidate list before quality testing begins. Foundry tools. Start with the replacement suggested in the retirement schedule, then build a shortlist with model benchmarks, which compare quality, safety, cost, throughput, and latency. Use trade-off charts such as quality versus cost and the side-by-side model comparison for up to three models. Compare context windows, feature support such as function calling, structured output, and vision, and available endpoints. Confirm SKU, region, quota, and upgrade mechanics in the model retirements documentation. Where it breaks. Teams face several plausible candidates, such as gpt-5.1, gpt-5.2, and a nano variant, without clear positioning between them. A selected model may be unavailable in the required region or SKU, a constraint that sometimes appears only after planning is underway. Historical traffic may also show that the new model costs substantially more, forcing an unplanned budget decision. What your team provides. Public benchmarks should filter the candidate list, not make the final decision. Confirm the shortlist against the team's own workload. Build the monthly cost view from token logs and current pricing. 3. Adapt Adapt is often the most time-consuming phase for embedded and product-facing workloads. Validate may take longer for regulated workloads. First, replay the existing workload on the new model without changing it. This isolates changes caused by the model. Diagnose shifts in verbosity, reasoning depth, structured-output adherence, tool-call shape, and latency. Then update the application until it recovers or improves on the previous behavior. Prompt editing is only one part of Adapt. A migration often changes four other surfaces: Parameters. temperature, top_p, max_tokens, and reasoning-effort controls may not map directly between generations. Some are unsupported by newer model families. Tool definitions. Argument names, descriptions, and required fields that reliably guided the old model may need clearer wording or tighter constraints. Output schemas. Structured-output behavior changes between models. A schema the old model followed loosely may need explicit constraints, or the new model may finally enforce it. Calling code. API and SDK differences, including Chat Completions versus Responses, streaming formats, and new or renamed request fields, can require code changes. Downstream parsers may also assume the old response shape. For agentic and workflow workloads, schema and tool-call changes can outweigh prompt changes. Foundry tools. Prompt Optimizer is available through the Optimize button below the system instructions field in the Agent playground. It restructures instructions, explains each change by paragraph, and supports iteration. For example, a team can add a constraint such as "keep the JSON schema exactly" and optimize again. It is a fast first pass for a prompt that would otherwise be rewritten by hand. For agent workloads, agent optimization tunes instructions, tools, and model selection together. Prompt Optimizer and agent optimization are available in Microsoft Foundry, not Azure OpenAI. When production data is unavailable, the simulator can generate synthetic and adversarial inputs. Where it breaks. Most migration time is spent in a manual diagnosis loop. Teams rerun prompts by hand, compare outputs by eye, and rarely record what changed or why. For agent builders, chat benchmarks may miss tool-call regressions such as extra fields, renamed arguments, or changed call sequences. Those problems appear only when the team replays real agent traces. Plan for three constraints: Start with the optimizers, then verify their output. They apply general practices in a single pass rather than fitting changes to the team's dataset. They tune instruction text, not tool definitions or output schemas. Copy the original prompt first because there is no version history, then evaluate the optimized prompt against the frozen dataset. Expect more manual work when moving between providers or model families. There is no "optimize for target model X" flow. Moving from one family to another, such as OpenAI to Claude, still requires deliberate prompt and schema translation. Record traffic before you need it. Replay is only as useful as the captured data. Existing traces are available as an evaluation source for agents today, while content capture is opt-in and never retroactive. Log prompts, responses, latency, and tokens now to prepare for the next Adapt phase. 4. Validate Run the adapted workload against a quality rubric on the frozen dataset. The rubric may combine rules, LLM-as-judge evaluation, human review, existing user-feedback signals, or a domain-specific scoring framework. Examples include a clinical summarization rubric for healthcare or a tool-call sequencing assertion for agents. Validation produces a pass-or-fail decision for production exposure. AI-native teams may run the same signal continuously on every commit rather than treating it as a one-time gate. The dataset is a dependency for both Adapt and Validate. Build and freeze it early, around Assess, even though its primary purpose belongs to this phase. Validation then has two touchpoints: Before Adapt, freeze the dataset and success criteria, then run the current model to establish the source baseline. After Adapt, run the target model against the same dataset and evaluators, compare it with the source baseline, and make the release decision. Prepare the evaluation runner early and apply the gate after Adapt. Both steps belong to Validate. Foundry tools. The Azure AI Evaluation SDK, installed with pip install azure-ai-evaluation, includes more than 30 evaluators. They cover grounding, relevance, retrieval, coherence, fluency, question answering, reference-based similarity, F1, BLEU, ROUGE, safety, agent behavior, and Azure OpenAI graders. Teams can also build custom LLM-as-judge evaluators for task-specific rubrics. The portal evaluation flow runs the same evaluators against model, agent, dataset, and trace targets. Run identical evaluators against source and target outputs on the frozen dataset so the results remain comparable. Measure the three dimensions used for sign-off: Quality: evaluator results Latency: leaderboard time to first token and throughput, plus operational latency from the workload Cost: (input tokens × input price) + (output tokens × output price) Where it breaks. Most teams do not have an evaluation suite. Teams that do often built it themselves and may not use platform evaluation tools. Regulated workloads add mandatory human review, which can become the bottleneck. For those teams, migrations often stall in Validate rather than Adapt. What your team provides. The evaluators are ready to run, but model workloads still require teams to curate a domain-relevant test set from production traffic. That is why Phase 0 pays for itself. 5. Roll out Promote the validated configuration in stages: non-production, then a canary or weighted percentage of production traffic, followed by broader exposure. Compare live latency, errors, and quality signals with the pre-migration baseline, then commit or roll back. Some workloads cannot expose a new model to customer traffic during testing, including flows involving protected health information or financial transactions. Use shadow or mirror mode instead: run the new model offline against production inputs and compare its outputs with the old model without affecting users. Foundry tools. Migration mechanics depend on the deployment SKU: Standard, Global Standard, and Data Zone Standard deployments upgrade automatically on a rolling schedule. Control timing with versionUpgradeOption: OnceNewDefaultVersionAvailable, OnceCurrentVersionExpired, or NoAutoUpgrade. Priority Processing follows the same path. Provisioned, Global Provisioned, and Data Zone Provisioned deployments migrate manually, either in place during a 20-to-30-minute Azure-managed traffic transition or through side-by-side deployments. Batch deployments migrate side by side. Deploy the new model, resubmit jobs, then retire the old deployment. Fine-tuned deployments do not upgrade automatically. They follow separate training and deployment retirement schedules, so plan retraining or distillation early. See the model retirements documentation for deployment-specific guidance. Use continuous evaluation to score a sample of production traffic in the Foundry Observability dashboard. Connect evaluation results to traces for root-cause analysis and configure Azure Monitor alerts for quality regressions. Where it breaks. Offline evaluation can miss production quality and latency regressions. Rollback decisions may also be forced by deprecation deadlines rather than evidence. What your team provides. Teams implement weighted routing between deployments in their application or gateway layer. They must also choose how long to keep the old deployment warm for rollback. Embedded copilot teams often target about 30 days. Design both mechanisms once and reuse them for future migrations. 6. Retire Retire is easy to forget. Decommission the old deployment, free its capacity, archive evaluation artifacts, update internal documentation, and communicate the change to downstream owners. That may include customer-facing documentation, marketing pages, support runbooks, and audit logs. Regulated workloads may need to retain artifacts for years. Retirement is also a governance step. Foundry tools. Use the Models API to confirm that the old version is retired through lifecycleStatus or 410 Gone. Use the observability dashboard to confirm that the active deployment count falls. Add useful production traces to the golden dataset so the next migration starts with better evidence. Where it breaks. Teams skip the phase. Zombie deployments accumulate, leaving teams with structural debris from migrations they never finished. What your team provides. The observability dashboard shows deployment count, but the team must decide which deployments still carry traffic. Create an explicit retirement ticket rather than relying on someone to remember. Embedded copilot teams also need to update public claims such as "powered by gpt-4o" after the model changes. Worked example: Zava's Shopping Assistant migrates from gpt-4o mini to gpt-5.x Zava is a fictional retailer used as a stand-in for a real customer story. The example reflects patterns observed in customer-facing embedded AI workloads. Zava's Shopping Assistant is one of the company's largest LLM workloads. It has two LLM stages: Per-review insight extraction identifies sentiment, attribute mentions, and defect signals across thousands of product reviews. Product-level summaries present those findings to shoppers on the product page. Together, the two stages account for a meaningful share of Zava's token volume. Discover Zava's central AI Platform team made gpt-5.x models available internally and notified feature teams. The Shopping Assistant team learned about the models through that channel and received a target migration window before gpt-4o mini's deprecation. Zava has an internal discovery layer built on the retirement schedule and Models API. Microsoft provides the underlying data, while Zava routes the signal to application owners. Assess The team compared gpt-5.4 nano, which offered lower latency and cost, with gpt-5.1 and gpt-5.2 using leaderboard trade-off charts. Selection remained difficult because of the rapid release cadence, unclear positioning between variants, and the lack of a behavioral benchmark for product question-and-answer workloads. Capacity planning required coordination with the AI Platform team. Both gpt-4o mini and the gpt-5.x candidate needed to remain available in the same regions so the team could roll back. Adapt The team ran its existing Shopping Assistant prompts against gpt-5.4 nano using a sanitized traffic sample. Customer queries were scrubbed of personally identifiable information before replay. The behavioral comparison found three problems: Summaries used more hedged language and sometimes contradicted the underlying review evidence, creating a shopper-trust risk. Insight counts varied across runs. The model sometimes extracted substantially more or fewer attribute mentions than gpt-4o mini, affecting downstream filtering. Latency varied more than expected on the synchronous product-page path. Prompt Optimizer helped restructure the summary prompt, but the team still diagnosed the differences manually. It built its own replay system and behavioral comparison on top of captured traffic. Reengineering took weeks and extended beyond prompts: parameters and downstream parsing for the insight-extraction output also changed. Validate The team scored outputs with Zava's Product Answer Quality (PAQ) rubric. Its nine criteria cover factual grounding, attribute accuracy, tone, and refusal behavior for questions outside the catalog. Zava implemented the rubric as custom evaluators in the Azure AI Evaluation SDK. Initial evaluations used unchanged prompts to isolate model behavior. The team reran them after each prompt change. Zava's QA team also completed a manual review, which the company requires for every new model used in a customer-facing workflow. The per-review insight extraction stage still has no automated evaluation, a gap the team has accepted for now. Roll out The validated configuration moved to non-production and then through staged exposure: employees first, followed by a small percentage of shoppers. The canary exposed latency regressions that offline evaluation had missed. The team rolled back the latency-sensitive synchronous product-page path while keeping the offline pregeneration path on the new model. Retire Retirement is not complete. The gpt-4o mini deployment remains warm for rollback. The team must resolve the synchronous-path latency regression before retiring it, which sends that code path back to Adapt. This split state is common. Retire can lag Roll out by weeks or months, and the old deployment remains visible in deployment-sprawl data. Lessons from the example Adapt consumed most of the schedule. Replay tooling, behavioral comparison, and prompt reengineering are the clearest opportunities for Microsoft to shorten migrations. Validate worked because Zava had invested in it. Most customers do not have an equivalent to PAQ. Making domain-specific evaluations cheaper to build would improve confidence in this phase. The migration is partially live and partially rolled back. The process must support split-state workloads rather than assuming a binary switch from old model to new. How to use this process The six phases can serve as a checklist and an interview script for teams creating or auditing a migration process. Documentation and process: Lead with the six phases. Most teams recognize them immediately. Investment priorities: Start with Adapt and Validate. Across customer stories, those phases consume the most time and confidence. Interviews and postmortems: For each phase, ask whether it happens, who owns it, which tool the team uses, where it failed last time, and what evidence would increase confidence. Metrics: Discover and Retire are the easiest phases to instrument through measures such as announcement reach and active deployment count. Adapt and Validate require purpose-built telemetry that most teams do not yet have. Go deeper Two companion resources turn the process into concrete implementation steps: Microsoft Learn guide: This article follows the six phases through identifying affected deployments, preparing a test dataset, adapting prompts, evaluating source and target models, and rolling out by deployment type. Foundry Models Accelerator: This community toolkit includes a deployment-inventory scanner, a feasibility and assessment playbook, code and API migration audit scripts, an A/B evaluation runner with golden datasets, and rollout guidance. It follows the same six-phase process. The Foundry Models Accelerator is a community-built toolkit provided as-is under the MIT License. It falls outside Microsoft Support. Always verify model availability and retirement dates against official documentation. Additionally, check out the Foundry Forgebook which hosts a plethora of recipes that walk through the required code changes to migrate from different source to target models, even across model families. The goal of a model migration is to change the model without changing your application’s behavior, or to change it measurably for the better. Lead with the six phases, invest first in Adapt and Validate phases, and treat the Retire phase as a governance step.818Views0likes0Comments