evaluation
43 TopicsBeyond the Trace: The Science of Insight Quality
Co-authors and reviewers: Morteza Ziyadi, Hanchi Wang, Han Che, Billy Hu, Sean Gayler, Nishal Dsilva, Avinav Jami, Ankit Singhal, Augustus Arthur TL;DR. Insights in Foundry turns recurring agent behavior into evidence-linked findings developers can review and act on. We evaluate trace linkage, finding quality, and detection of known issues using labeled benchmarks, LLM-judge assessment, and controlled end-to-end tests. What is an Insight? An Insight is a reviewable finding about recurring agent behavior. It brings together an explanation, supporting trace evidence, and a possible next step, helping developers investigate a pattern rather than inspect each execution in isolation. Depending on the available evidence and supported configuration, an Insight can include: Part of an Insight What it gives the reviewer Title and description The recurring behavior and an evidence-based explanation of a possible cause. Linked traces The broader group of traces associated with the finding. Highlighted traces Representative examples to open and check against the explanation. Category, severity, and status Context for triaging the finding, not a substitute for a risk assessment. Agent version and recency Which version is represented and when the finding was created. Proposed action or fix An investigation or improvement path; concrete proposals depend on the configuration. Figure 1. An Insight detail view in Microsoft Foundry, cropped from public Microsoft Learn documentation. The description, evidence, and proposed fix support human review. This UI illustration is not a benchmark result; the source link provides a full-size view. UI source and field definitions: Insights in Foundry documentation. From trace search to an actionable review queue Production agents can generate thousands of traces containing model calls, tool calls, latency, token usage, errors, and final responses. Observability shows what happened, while evaluations test criteria a team already knows to measure. The harder problem is discovering repeated behavior that the team did not know to predefine. Insights in Foundry analyzes traces from the Application Insights resource connected to a Foundry project and organizes recurring behavior into reviewable Insights. Depending on the evidence and supported configuration, an Insight can include representative traces, affected agent versions, severity, an explanation of a possible cause, and a recommended next action. Concrete prompt or code proposals are available only for supported agent types and configurations; other Insights provide general investigation or remediation guidance. Insights in Foundry is now available in public preview. It analyzes production traces to surface recurring behaviors and regressions with supporting evidence and suggested areas for investigation or improvement. Developers remain in control: review the cited traces and validate proposed changes through normal evaluation and deployment practices. During public preview, we will continue expanding the experience based on customer feedback. To make this concrete, Figure 2 maps 220 inputs from TraceElephant, a public agent-trace benchmark, to seven generated Insights. Figure 3 examines one linked trace. Figure 2. Recorded links from 220 public TraceElephant inputs to seven generated Insights in the September 17 benchmark. The 80-input finding includes the looping trace examined in Figure 3. Titles are editorial paraphrases; the layout is editorial, not a product screenshot or a validation of each diagnosis. Figure 3. A public TraceElephant trace contains 54 model calls; steps 26, 30, 34, 38, 42, 46, and 50 repeat an identical extraction plan. The September 17 benchmark run of the production pipeline links this trace to a finding about missing progress-aware termination. Wording is paraphrased; this is an illustration, not product UI. The proposed intervention still needs evaluation. Three complementary ways to evaluate Insight quality Insight quality involves which inputs a finding links, how well it explains the evidence, and whether it identifies known issues in controlled tests. We combine labeled trace evaluation, unlabeled finding assessment, and controlled end-to-end testing. These are complementary sources of evidence, not mutually exclusive dataset categories. Evidence Question Assessment / results Labeled trace evaluation Do Insights link inputs that reference annotations mark as failing? Reference labels; trace precision and trace recall (%). Unlabeled finding assessment Are generated findings grounded and useful? Eight-dimension LLM judge; mean score (1-5). Controlled end-to-end tests Do Insights identify known injected issues? Known issues and healthy baselines; detections, misses, unsupported findings, duplicates, and unscored cases. Here, 'unlabeled' means reference failure labels are not used for the evaluation; the sources can still contain annotations or reference answers. Public versus internal describes where data comes from, not how it is evaluated. LLM judges can also assess findings from labeled or controlled scenarios. Labeled evaluation: Do Insights link the right traces? Trace precision: among unique benchmark inputs linked by at least one generated Insight, the fraction marked as failing. This measures discrimination only when a dataset includes healthy inputs, and it does not validate the diagnosis. Trace recall: the percentage of failure-labeled benchmark inputs linked by at least one generated Insight. It does not measure how many distinct failure modes were discovered. We evaluated the production pipeline on the same 746 benchmark inputs (653 failure-labeled) on September 15, 16, and 17, 2026. Repeating these inputs does not create 2,238 distinct examples. The six benchmark corpus slices come from public agent-trace research datasets with human reference annotations: AgentRx (Tau-bench retail), AgentRx Magentic-One, TRAIL, AgentErrorBench, TraceElephant, and the MAST-Data human subset. AgentErrorBench is the annotated failure-trajectory dataset released with AgentDebug, a framework for detecting and recovering from agent failures. The two AgentRx slices come from the same public release. The benchmark uses selected and normalized inputs from these releases, not a representative sample of customer production traffic. For each dataset, the chart shows daily generated-Insight counts and trace recall. The table reports equal-weight means of daily trace precision and recall. These datasets support ongoing development; results depend on the dataset and model used. Labeled results Figure 4. Generated Insight count versus trace recall for all six public corpus slices. Each point is one daily run; labels retain values where markers overlap. Recall measures linkage to failure-labeled inputs, not diagnosis correctness. Three-day mean precision and recall follow in the table. Dataset Inputs/day Failure-labeled/day Linked/day range Mean trace precision Mean trace recall AgentRx (Tau-bench retail) 102 29 43-45 37.8% 57.5% AgentRx Magentic-One 58 44 54-57 75.3% 94.7% TRAIL 148 143 139-140 96.9% 94.6% AgentErrorBench 200 200 194-196 100.0%** 97.3% TraceElephant 220 220 218-220 100.0%** 99.7% MAST-Data (human subset) 18 17 14-16 93.3% 82.4% ** When every input is failure-labeled, 100% trace precision cannot measure false-positive control. What the labeled results show High precision can reflect the corpus base rate. AgentErrorBench and TraceElephant contain only failure-labeled inputs, so their measured precision is 100% by construction. TRAIL, AgentRx Magentic-One, and the MAST subset also contain mostly failures; their precision values are close to their underlying failure-label prevalence. Read precision beside trace recall and each dataset's label prevalence rather than as a standalone quality score. AgentRx (Tau-bench retail) is the clearest mixed-traffic stress test. Only 29 of 102 inputs were failure-labeled. Across the three days, trace precision was 43.2%, 37.8%, and 32.6%, averaging 37.8%; trace recall was 65.5%, 58.6%, and 48.3%, averaging 57.5%. Daily linked counts were 44, 45, and 43, including 19, 17, and 14 labeled failures respectively. These daily rates are averaged before rounding. Some linked inputs without failure labels could contain issues outside the reference labels, such as cost or latency, but this benchmark cannot confirm that. Trace precision and trace recall measure whether Insights link inputs that reference annotations mark as failing. They do not establish whether an explanation is correct or a proposed action is useful. Our unlabeled evaluation examines those qualities using an eight-dimension LLM judge. Unlabeled evaluation: Are the findings grounded and useful? For this evaluation, a separate LLM judge scores generated Insights against the supplied input evidence, without using reference failure labels to compute precision or recall. These scores are automated assessments produced by an LLM judge. They are not human ratings and do not establish that a proposed action will improve the agent. This evaluation covers PUPA, FailSafeQA, tau2-bench, synthetic scenarios, and an internal production dataset for the monitoring dashboard agent compiled from Microsoft employee usage. These sources are normalized into benchmark inputs; not every input is a complete execution trace. The eight-dimension judge rubric Each generated Insight receives a score from 1 to 5 on eight dimensions. The rubric separates the quality of the explanation, the importance of the issue, and the usefulness of the suggested action: Dimension Question the judge considers Actionability Does the Insight tell a developer what to investigate, evaluate, or change next? Specificity Does it name concrete behavior, tools, or prompt elements rather than offer generic advice? Novelty Does it surface a pattern beyond an obvious dashboard signal? This is an estimate, not a measurement of what a team already knows. Correctness Is the claim supported by the supplied evidence? Severity calibration Does the assigned severity match the evidenced impact? Impact How consequential does the underlying issue appear, rather than just how well it is described? Fix specificity Does the proposed fix identify an exact asset or behavior to change? Fix applicability Is that proposed change plausibly on target for the issue? For example, a concrete next step earns a stronger actionability rating than generic advice. Severity calibration and impact are deliberately separate: a minor issue can have a well-calibrated severity label without being high-impact. Fix applicability is judged from the proposal and evidence; the fix is not executed as part of this scoring. Unlabeled results Results compiled from two recent benchmark runs cover 4,496 inputs and 40 generated Insights. Every emitted Insight in these selected snapshots was scored with the same judge rubric. Each Insight's overall score is the mean of its eight dimension scores; the table then averages those overall scores within each dataset. We do not combine datasets into one quality score. Dataset Inputs Insights scored Mean judge score (1-5) PUPA 901 5 3.53 FailSafeQA 1,101 11 3.57 Monitoring dashboard agent (internal) 1,342 4 3.88 tau2-bench 1,112 16 2.84 Synthetic scenarios 40 4 3.81 These are per-dataset snapshots, not a controlled comparison of dataset difficulty or pipeline versions. Small Insight counts and grading choices, such as grouped versus individual judging, make these protocol-specific diagnostics rather than stable absolute ratings. Automated judge scores are not precision or recall, human ratings, or measured improvements after applying a fix. The dimension-level scores make the weak spots more concrete. Fix specificity was the lowest-scoring dimension in four of the five snapshots, with means from 2.00 to 2.64. For tau2-bench, correctness was lowest at 2.25: a signal to review whether claims are supported by the supplied evidence. Making a next step more concrete and grounding a diagnosis more carefully are different improvement targets. What the public findings look like The examples below connect generated findings to concrete input evidence: an arithmetic inconsistency, a question/answer mismatch, and a payment allocation that exceeds the available balance. They are selected illustrations, not a representative sample of the 40 scored Insights. Figure 5. Three illustrative findings, with one checked input per finding. Titles and evidence summaries are editorial paraphrases. PUPA is a normalized QA record; FailSafeQA normalization pairs a perturbed question with the original answer, creating the displayed mismatch. Neither is a new agent execution. Tau2-bench shows a recorded tool interaction. These examples are not a representative quality sample. Public input sources: PUPA; FailSafeQA; tau2-bench. Controlled end-to-end tests: Can Insights find known issues? Dataset benchmarks are complemented by tests with healthy baseline agents and versions containing predefined defects. The framework generates traffic, runs Insights on the resulting traces, and compares the findings with the known issues and supporting evidence. It separately tracks detections, misses, unsupported findings, and duplicates. Cases with incomplete evidence remain unscored rather than counting as passes or misses. The September 23 hosted-agent report, used here as a concrete illustration, covered finance, travel, and support-ticket scenarios. Of 12 expected issues, 11 were scorable and 10 were detected. All three scorable healthy baselines had no confirmed unsupported findings. The figure shows the full outcome counts, including the missed and unscored cases. This end-to-end framework is separate from the 40-input synthetic-scenarios dataset in the unlabeled results. Figure 6. Controlled tests exercise healthy baselines and known defects through the Insights pipeline. The September 23 report records 10 detections among 11 scorable expected issues, one additional unscored issue, and no confirmed unsupported findings on three scorable baselines. No confirmed unsupported findings (noise) or duplicate findings were reported in the evaluated subset. This illustrative snapshot is not a representative product-wide rate. A broader view of quality Trace selection and rubric review are part of a broader quality framework. Category agreement is an exploratory benchmark diagnostic, not an assessment of the portal's category labels. Controlled scenarios also check for unsupported or duplicate findings. Figure 7. Five complementary quality questions. The labeled results above measure trace precision and trace recall; the unlabeled results assess findings with an LLM judge. Category agreement is an internal diagnostic, while controlled scenarios check noise and duplication. How we track quality over time Labeled benchmarks, unlabeled assessment, and controlled tests feed recurring quality reports and human investigation. Test inventories and assessment conditions can change, so daily scores are not automatically a comparable improvement trend. The reported results do not establish longitudinal recurrence or deduplication performance across runs. In the portal, Give feedback lets users report incorrect findings, categories, severity, grouping, or duplication. Start with scenarios whose expected failures and healthy controls are known. Track detection, trace precision, categorization, severity, distinctness, evidence grounding, and action quality instead of relying on one score. Review changes to data, models, prompts, and evaluation contracts as changes to the measurement system. Investigate weak results and newly reported failure patterns rather than tuning only for an aggregate. Keep human review in the loop, because benchmark labels and automated graders cannot determine business impact or remediation correctness. How practitioners should interpret an Insight Validate each Insight before acting on it. Microsoft Learn recommends this review sequence: Confirm the affected workflow, agent version, category, severity, and time. Open the highlighted traces and verify that the cited behavior is present. Compare problematic examples with healthy traces to test whether the grouping and likely cause are plausible. Decide whether the issue belongs to the agent, a tool, a model endpoint, a data source, or the platform. Turn confirmed recurring behavior into evaluation coverage, an optimization objective, owner routing, or a monitored no-action decision. An empty result does not prove that an agent is healthy, and a large linked-trace count does not prove business impact. Quality depends on representative traces, complete telemetry, a supported analysis model, and careful human validation. Get started Start with the Insights in Foundry documentation for prerequisites, portal steps, evidence-review guidance, SDK examples, pricing considerations, and preview limitations. Prerequisites include a connected Application Insights resource, recent representative traces, a supported GPT-5-or-newer Judge model deployment, and the required role assignments. Insight generation, including scheduled runs, uses your model deployment and can incur model charges. See the current documentation for the latest supported agents, models, regions, limits, pricing, and UI guidance. Python samples: on-demand Insights and scheduled Insights. Closing thoughts Agent quality work often starts with a simple question: what is repeatedly going wrong that we did not know to test? Insights in Foundry is designed to help teams answer that question with evidence-linked findings and a reviewable next step. Measuring those findings requires explicit definitions, diverse data, honest treatment of weak results, and human judgment where automated metrics stop. Labeled trace evaluation, unlabeled rubric assessment, and controlled end-to-end tests answer complementary questions. Newly observed patterns can then inform the next round of evaluation and review.165Views0likes0CommentsEvaluate More, Spend Less: Batching Microsoft Foundry Evaluators for Efficient Evaluation
Authors: Salma Elshafey, Ali Mahmoudzadeh, Kayla Ames, Ahmad Qardahji, Vivek Bhadauria, Morteza Ziyadi, April Kwong A single agent trajectory may need to be evaluated for groundedness, coherence, instruction following, task completion, and correct tool use. In a conventional pipeline, each evaluator receives the same messages, tool calls, tool results, and tool definitions in a separate model call. As trajectories grow, evaluation repeatedly pays to process the same context. Can one LLM judge call apply five or six evaluators without losing the signal that makes evaluation useful? This study tests a simple design: send the shared context once and ask one composite evaluator to score several criteria in the same call. In summary: On 100-row quality and tool use samples, composite evaluation required 5-6x fewer model calls, reduced total input-token use by 61.25–71.07%, completion tokens by 46.63–63.89%, and measured run wall time by 35.74–46.10%. Across the tested workloads, frontier judge models maintained stable quality across several high-value dimensions. The results support composite evaluation as a cost-efficient default, with targeted individual evaluation for workload-sensitive rubric criteria. Why Evaluation Becomes Expensive Evaluating a single conversation rarely involves just one evaluation dimension. A typical assessment examines multiple dimensions of quality, such as whether the response is coherent, follows instructions, completes the user's task, remains grounded in available evidence, and uses tools correctly. Because each dimension is typically evaluated through a separate prompt and model call, the same conversation trajectory, tool calls, tool outputs, and tool definitions are processed repeatedly. As a result, cost grows with both the number of evaluation criteria and the amount of shared context, making long agent trajectories especially expensive to evaluate. The Composite-Evaluator Approach Rather than evaluating each dimension independently, composite evaluation scores multiple criteria within a single judge-model call, allowing the shared context to be processed once instead of repeatedly. We built two composite evaluators. The Output Quality Evaluator The Output Quality Evaluator scores six Microsoft Foundry evaluators in one LLM call: Evaluator Question Fluency Is the response clear and well formed? Coherence Is it logically consistent with the conversation? Intent Resolution Did the assistant understand the user's goal? Task Adherence Did it follow the user's instructions? Groundedness Are its claims supported by the available evidence? Task Completion Did it complete the requested task? Tool Use Quality Evaluator The Tool Use Quality Evaluator scores five Microsoft Foundry evaluators in one LLM call: Evaluator Question Tool Call Accuracy Was the tool call correct overall? Tool Call Success Did the invocation succeed? Tool Input Accuracy Were the arguments correct? Tool Output Utilization Did the response use the tool output correctly? Tool Selection Was the appropriate tool selected? Each composite prompt includes the shared trajectory and tool definitions once, followed by the full definition, rating scale, and applicability rules for every evaluator. It instructs the judge to assess each evaluator independently so that one verdict does not bias another. The judge returns structured results for each criterion, including a score, rationale, and applicability status. For multi-turn evaluations, it can also identify the earliest turn where a failure occurred. How We Validated Evaluator Quality The study compares individual and composite evaluators across: Two modes: Single-Turn and Multi-Turn. Nine judge models: GPT-4o, GPT-5.4, GPT-5.4 mini, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, DeepSeek V4 Flash, DeepSeek V4 Pro, and Grok 4.1 Fast Reasoning. Multiple datasets, as shown in the table below Dataset Mode Rows Primary validation target Reference labels Internal quality set Single-Turn and Multi-Turn 283 Six quality evaluators Existing per-evaluator labels Internal tool use set Single-Turn and Multi-Turn 200 Five tool use evaluators Existing per-evaluator labels BFCL v4 Multi-Turn 200 Task Completion, Tool Call Accuracy, and failed-turn localization Deterministic state and tool-call checks AgentIF Multi-Turn 148 Task Adherence Constraint-level majority-vote reference FaithDial Multi-Turn 300 Groundedness Faithful versus hallucinated responses FED Multi-Turn 125 Coherence Five-annotator human scores Tau-Voice Multi-Turn 278 Task Completion Reward-derived completion labels AgentRx tau_retail Multi-Turn 29 Failed-turn localization Human failure annotations We measured several distinct properties: Input tokens, output tokens, model calls, and latency Accuracy, macro-F1, and Cohen's kappa where labels were available Agreement and correlation structure between individual and composite modes Repeatability across four evaluation runs Finding 1: Composites Cut Input Tokens by 61.25–71.07% and Latency by 35.74–46.10% We measured cost and latency on matched 100-row Single-Turn samples: one for Output Quality and one for Tool Use, comparing parallel individual evaluators with one composite evaluator per row. For each row, evaluator context includes the query and response, any tool calls and results, and the serialized tool definitions. Note: Single-turn means that the evaluator's focus is the last agent response only, but the entire conversation history is passed to the evaluator as context. Trajectory Length in the Cost Sample The log-scale distributions show that most rows clustered around a few thousand tokens, with a smaller number of substantially longer trajectories. In both samples, longer evaluator contexts produced larger absolute token savings. The Measured Cost Reduction Across the sample: Repeated context is the main saving. Total input-token volume fell from 2,058,930 to 797,843 for Output Quality, a 61.25% reduction, and from 2,036,122 to 589,097 for Tool Use, a 71.07% reduction. Uncached input fell by 57.51% and 53.29%, respectively. For Output Quality, 52.72% of composite input tokens were cached versus 56.88% across the individual suite; for Tool Use, the figures were 46.22% versus 66.69%. However, the composites still processed substantially fewer tokens overall. Call volume collapses. Output Quality fell from 600 individual evaluator calls to 100 composite calls. Tool Use fell from 494 calls to 100. In six rubric-row combinations, native applicability rules legitimately returned no score, such as when there was no tool call to evaluate; those skips explain why the individual Tool Use baseline is below 500. Measured runs finish sooner. Output Quality wall time fell from 504.41 seconds to 324.12 seconds, a 35.74% reduction, while Tool Use fell from 536.95 seconds to 289.42 seconds, a 46.10% reduction. Per sample row, that is 5.044 to 3.241 seconds for Output Quality and 5.369 to 2.894 seconds for Tool Use. Completion tokens fell by 46.63% and 63.89%, respectively. On the quality sample, the composite used 797,843 total input tokens and 36,919 completion tokens, compared with 2,058,930 and 69,180 for parallel individual evaluation. On the tool use sample, the composite used 589,097 total input tokens and 28,358 completion tokens, compared with 2,036,122 and 78,522. These are measured run totals, not prompt-length estimates. The broader experiments showed the same benefit. The full six-evaluator rubric Single-Turn study used about 68% fewer input tokens, while the 283-row Multi-Turn quality study used about 77% fewer: 28,715 rendered input tokens per row for a six-call equivalent versus 6,608 for the composite call. Exact currency savings depend on model and deployment pricing, but batching consistently removed duplicated input and round-trip overhead. Finding 2: Quality Tradeoffs Depend on Rubric and Judge Model Output Quality Rubric Benchmark Results On the internal Single-Turn quality set, composite Task Adherence accuracy improved for every judge, while Task Completion remained close to individual evaluation. External Multi-Turn benchmarks showed a more varied pattern across rubric, judge, and workload. This figure reports Δκ = κComposite - κIndividual: blue favors composite evaluation and orange favors individual evaluation. The mixed directions reinforce the need to validate the selected rubric and judge on production-like data. BFCL provided the strongest labeled Multi-Turn example for Task Completion. With GPT-5.6 Luna, the composite evaluator reached macro-F1 0.848 and κ = 0.696, compared with 0.836 and 0.672 for the individual evaluator. Across the other judges, composite and individual Task Completion remained close for Sol, Terra, and DeepSeek V4 Flash, while results for DeepSeek V4 Pro and Grok 4.1 Fast favored individual evaluation. Tool Use Quality Rubric Benchmark Results The five-rubric tool composite was evaluated on production-style tool traces. Each row carries ground truth for one source rubric, so the table reports accuracy on the available per-evaluator subsets. Judge Tool Call Accuracy Tool Call Success Tool Input Accuracy Tool Output Utilization Tool Selection gpt-4o 0.800 0.902 0.850 0.800 0.878 gpt-5.4 0.800 0.902 0.775 0.737 0.829 gpt-5.4-mini 0.750 0.902 0.725 0.763 0.756 gpt-5.6-luna 0.914 0.902 0.848 0.775 0.901 gpt-5.6-sol 0.886 0.902 0.750 0.743 0.854 gpt-5.6-terra 0.857 0.902 0.750 0.794 0.854 DeepSeek-V4-Flash 0.775 0.902 0.800 0.848 0.854 DeepSeek-V4-Pro 0.800 0.878 0.800 0.789 0.902 grok-4-1-fast-reasoning 0.946 0.902 0.925 0.745 0.823 Across the production-style corpus, the composite reached 0.75-0.95 accuracy across the five rubric criteria. The results show model-specific strengths: Grok 4.1 Fast led Tool Call Accuracy and Tool Input Accuracy, DeepSeek V4 Flash led Tool Output Utilization, and DeepSeek V4 Pro led Tool Selection. Tool Call Success remained tied at 0.902 for most judges. Luna remained among the strongest overall, but no single judge dominated each of the rubric criteria. Classification Accuracy On the truly multi-turn BFCL benchmark, Tool Call Accuracy is the only tool rubric criterion with native ground truth. GPT-5.6 Terra led the single-run panel at 0.864 macro-F1, followed by GPT-5.6 Sol at 0.858 and GPT-5.6 Luna at 0.828. Other models showed lower results, with Grok 4.1 Fast at 0.626, DeepSeek V4 Pro at 0.625, and V4 Flash at 0.521. This shows that with the correct model, the composite evaluator can classify overall tool-call correctness on full conversations as well as production-style traces. Failure Localization The same BFCL run also tested where failures occurred. GPT-5.6 Terra led at 0.78 member-any failed-turn localization: in 78% of failing conversations, the turn predicted by the Tool Call Accuracy rubric criterion matched at least one turn in BFCL's set of failing turns. It produced 22 false alarms across 75 passing conversations. GPT-5.6 Sol reached 0.75 member-any, and GPT-5.6 Luna reached 0.74. The Tool Selection channel with Luna provided a high-precision secondary signal, identifying a failing turn in 50% of failures with only 3 false alarms. To test whether localization generalizes beyond mechanically checked BFCL traces, we also used the 29-row AgentRx tau_retail set with human failure annotations. On the eight in-scope tool-mechanics failures, Sol localized 8 of 8 root causes, while Terra and Luna localized 7 of 8. Meanwhile, Grok 4.1 Fast only localized 2/8, and the two DeepSeek models 1/8. This result is directional because of the small sample. Does Batching Preserve Rubric Relationships? After testing direct accuracy, we examined whether batching also preserves how rubric verdicts relate to one another. This is supporting structural evidence, not an accuracy measure against ground truth. The heatmaps compare Spearman correlations between GPT-5.6 Luna's binarized rubric verdicts on rows scored by both modes. The tool use structure was especially stable, including the strongest relationship between Tool Call Accuracy and Tool Input Accuracy (0.76 in both modes). Quality relationships were more mixed. For example, Single-Turn Intent Resolution and Task Completion fell from 0.71 to 0.31, while Multi-Turn Task Adherence and Task Completion fell from 0.53 to 0.23. This shows that batching can preserve the broad rubric structure without preserving every relationship equally. Finding 3: Reliability Depends on Both Rubric and Judge Multi-Turn Quality Reliability We repeated the six-evaluator composite call four times on the same 200 BFCL rows. Across the nine-judge panel, DeepSeek V4 Flash had the lowest average flip rate at 3.9%. GPT-5.6 Sol and Grok 4.1 Fast followed at 5.2%, DeepSeek V4 Pro at 5.6%, GPT-5.6 Luna at 6.8%, GPT-4o at 7.4%, GPT-5.6 Terra at 7.8%, GPT-5.4 mini at 10.9%, and GPT-5.4 at 15.6%. The rubric criteria-level result is more important than the leaderboard: Fluency: 0.0% average flip rate Coherence: 2.4% Intent Resolution: 8.6% Task Adherence: 6.7% Task Completion: 11.5% Groundedness: 16.4% The same judge can be highly stable on one rubric criterion and noisy on another. Reliability should therefore be monitored at the criterion level, not inferred from a single aggregate score. Multi-Turn Tool Use Reliability The five-evaluator MT tool study showed the same need for criterion-level monitoring. Across four BFCL repeats, GPT-5.6 Terra and DeepSeek V4 Flash had the lowest average flip rate at 6.0%, GPT-5.6 Sol reached 6.1%, and GPT-5.6 Luna and GPT-4o were at 7.0%. GPT-5.6 Sol achieved the strongest mode-of-four Tool Call Accuracy kappa at 0.759, while Terra led the single-run quality/localization results. Tool Call Success was the most stable criterion, with a 4.7% average flip rate across judges. Practical Recommendations Based on the results across quality, tool-use, localization, and reliability studies, we can derive practical guidance for both judge selection and evaluator batching. Recommended Judges: The GPT-5.6 Family Across the quality, tool-use, localization, and reliability experiments, the GPT-5.6 family consistently delivered the strongest overall results. While individual benchmarks occasionally favored other models or specific GPT-5.6 variants, Luna, Sol, and Terra repeatedly ranked among the top performers and showed the most balanced performance across workloads. Sol was generally the most reliable across repeated quality and tool-use studies, Terra led several BFCL tool-use and failed-turn localization evaluations, and Luna achieved the highest average quality across the tested workloads. For workloads similar to those evaluated here, we recommend starting with GPT-5.6 judges and selecting between Sol, Terra, and Luna based on your specific priorities. When available at a lower deployment cost, Luna is a particularly attractive option because it maintained frontier-level evaluation quality while achieving the strongest average performance in our experiments. By comparison, lower-cost models often involved larger tradeoffs between quality, reliability, and benchmark performance. Exact cost savings depend on model pricing and deployment configuration, so benchmark candidate judges on production-like workloads before optimizing solely for cost. Which Evaluators to Batch These evaluators performed as well in the composite evaluator as they did individually — they're strong candidates for batching right away: Fluency and Coherence — highly stable across all nine judges and repeats Tool Call Success and Tool Selection — robust across the tested judge models Task Completion — close to individual evaluation across benchmarks; slightly stronger with GPT-5.6 Luna on BFCL Tool Call Accuracy, Tool Input Accuracy, and Tool Output Utilization — strong on production-style traces, with accuracy ranging from 0.75 to 0.95 depending on evaluator and judge Which Evaluators to Watch These evaluators showed workload- or judge-dependent results. They can still be batched, but we recommend checking the results on your own data: Task Adherence — performed well on typical conversations but struggled when a single request contained many independent constraints Intent Resolution — results varied across datasets even though agreement was strong in some cases Groundedness — the least consistent evaluator in our tests, with scores changing across repeated runs more often than any other dimension, albeit having better performance on frontier models, like GPT-5.6 Luna and GPT-5.4. If grounding accuracy is critical for your use case, consider keeping the individual evaluator as a fallback Where This Approach Fits Composite evaluation is most effective as a selective default rather than a universal replacement for every standalone evaluator. GPT-5.6 Sol offered the strongest overall balance, Terra led the single-run BFCL tool results, and Luna is the recommended cheaper high-quality option. Judge and rubric selection should therefore be validated on production-like data. The Multi-Turn comparison has one structural boundary: native individual evaluators do not exist for Multi-Turn Fluency and Intent Resolution. The composite evaluator can produce both dimensions, but a direct individual-versus-composite validation is not available for them in that setting. Finally, token savings do not translate to one fixed currency percentage. Pricing varies by model and deployment, while output length and retry behavior vary by workload. Takeaway Composite evaluators are a practical way to remove repeated context from evaluation pipelines. In the updated 100-row comparisons, they reduced total input-token use by 61.25% for six quality evaluators and 71.07% for five tool use evaluators, reducing completion tokens by 46.63–63.89%, while reducing measured run wall time by 35.74% and 46.10%, respectively, and preserving comparable quality across many dimensions. The strongest deployment pattern is selective rather than absolute: start with a model from the GPT-5.6 family; batch the evaluators that remain stable, monitor them independently, and keep focused fallbacks for the evaluators that do not. Composite evaluation is ultimately about removing unnecessary work. By evaluating multiple dimensions in a single judge call, teams can dramatically reduce repeated context processing while maintaining the evaluation coverage needed to monitor production systems. Get Started Start with the composite evaluator on a small set of known-good and known-bad agent conversations. Choose Output Quality to assess the six quality criteria or Tool Use Quality to assess the five tool use criteria. Review each criterion’s score and applicability status, then spot-check the results against individual evaluators for the failure modes that matter most to your application. The GPT-5.6 family is the recommended starting point for judge models. Re-check results when you change the judge model or the domains of conversations being evaluated. Try these evaluators in the Microsoft Foundry portal now. You can explore the docs on how to use them Agent Evaluators for Generative AI - Microsoft Foundry | Microsoft Learn.88Views0likes0CommentsHow to make AI responses faster on Microsoft Foundry: lessons from 2,040 measurements
A reproducible Microsoft Foundry performance study covering prompt caching, multimodal input, tool orchestration, MCP lifecycle, Toolbox tool search, and Priority Processing - with correctness and reliability measured alongside latency.176Views0likes0CommentsIs "uncertainty" the feedback signal Copilot Studio agents are actually missing?
At today's M365 Platform Weekly session, we were asked for our input on what feedback we wish we could pull beyond thumbs up/down and verbatims. Is it trends over time, sentiment themes, the response-to-triage loop, etc. Here's an angle: What if the primitive itself is wrong? Thumbs up/down measures satisfaction after the fact. What if we measured confidence instead? How often an agent actually knows it's on shaky ground, and whether the user's reaction matches that? If an agent flags its own uncertainty at the point of response instead of a static thumbs up/down, a feedback prompt gets generated from whatever's trending in that uncertainty instead of the same generic question every time, and "trends over time" becomes "did this agent get more confident or less confident since the last update" rather than a flat satisfaction line. It might also solve the silence problem that most users rarely click anything. A reaction that's actually specific ("you caught something the agent flagged as shaky") seems easier to engage with than a binary good or bad. To be clear, confidence signals already exist in adjacent forms. Copilot Studio and most conversational AI platforms already use a confidence score internally to decide whether to answer directly, ask a clarifying question, or escalate to a human. GitHub Copilot has used a confidence score since its earliest versions too, ranking code suggestions and defaulting to the highest-scoring one. None of that is new. So rather than "add a percentage next to the answer", what if there is a specific flag pointing at the exact claim or step the agent is unsure about, feeding into the feedback loop? Curious if anyone else building in Copilot Studio has run into this: Do you ever wish your agent had hedged when it didn't? What would you actually do with an uncertainty score if you had one? And would just love others thoughts on this :)235Views0likes2CommentsWhat is the best file format for an AI agent knowledge base?
This is a best practice sharing the best format for an agent and show you why you should convert your PPT, PDF, WORD into a TXT markdown. I had an issue with my agent, time taken to answer was too long, and usually we spend a lot of time asking: What is the best prompt? Why is my agent slow? Why does retrieval sometimes work and sometimes fail? How can I improve answer quality? But I realised I was asking another question much less often: What is actually the best file format for the knowledge base? PDF? Raw text? Markdown? Pre-chunked text? Semantic sections? Context-enriched text? And more importantly: How much does the format alone affect agent performance? I tried to find a quantified benchmark answering this specific question, with the same agent, same source knowledge and same questions, but different knowledge representations. I couldn't find one that really answered what I wanted to measure. So I decided to run the experiment myself on a real case. My first exploratory tests were already surprising: depending on the representation, the agent could be significantly faster and more accurate, despite working from the exact same source information. So I decided to push the test further. My objective I want to identify, without assumptions and based on actual evaluation data, how a long document should be prepared for an LLM knowledge base so that the agent can retrieve, understand, ground and answer from it as reliably as possible. I focused on five dimensions: Answer quality Retrieval reliability Source grounding / citations Execution time Robustness across single-turn and multi-turn questions The broader question I'm trying to answer is: How should we structure knowledge so that an LLM can retrieve and use it as reliably as possible? The test case I deliberately chose a document that isn't particularly friendly for RAG: a 46-page European regulation, https://eur-lex.europa.eu/eli/reg/2011/1169/oj?locale=fr, on the provision of food information to consumers. The information is distributed across articles, definitions, exceptions, annexes, tables, numerical thresholds and cross-references. That makes it useful for testing retrieval: answering correctly often requires finding a very specific piece of information while preserving enough context to understand how it applies. I used the native PDF as the baseline and created 6 additional knowledge-base representations of the same document: Raw TXT Markdown Chunk-ready TXT RAG-oriented units Semantic TXT Contextual TXT One rule: same knowledge, same agent, same instructions, same questions. Only the knowledge representation changes. The benchmark I used two evaluation sets: 42 single-turn questions testing broad coverage of the document: direct facts, thresholds, exceptions, annexes, lists and cross-references. 5 multi-turn conversations containing 13 questions, to see what happens when a user asks a question and then follows up with things like: "And in this case?" "What are the exceptions?" "And for dietary fibre?" This gave me: 47 evaluated test cases / 55 actual questions per format Across all 7 formats: 329 evaluated conversations 385 user questions executed First results Metric Native PDF Best structured representation Overall pass rate 66.0% 85.1% - Contextual TXT Best single-turn score 69.0% 88.1% - Chunk-ready TXT Multi-turn benchmark 40% 80% - Contextual TXT Multi-turn execution time 14m24 5m54 Total benchmark time 44m31 24m49 The quality gap was already substantial: 66.0% → 85.1% That's +19.1 percentage points while keeping the underlying knowledge unchanged. I also saw a major difference in execution time. On the multi-turn test: 14m24 → 5m54 That's approximately 2.4× faster. Across the complete benchmark: 44m31 → 24m49 Around 44% less execution time. These timings represent the complete agent evaluation pipeline, so they shouldn't be interpreted as pure LLM inference latency. But the difference under identical test conditions is large enough that I want to understand it better. Findings There wasn't one format dominating every benchmark. Chunk-ready TXT scored highest on independent questions: 88.1%, while Contextual TXT performed better across multi-turn conversations and finished with the highest overall score. That may suggest that the way we optimise a document for isolated retrieval isn't exactly the same as the way we should prepare it for conversational retrieval. In the contextual version, I tried to make every section understandable when retrieved independently by keeping useful information around it: Source references Section context Retrieval cues Relevant cross-references For regulatory documents, this seems particularly important. A numerical value retrieved alone can be meaningless without knowing which rule it belongs to, under which conditions it applies, and whether another article contains an exception. Where I am now This remains an exploratory benchmark: One document One domain One agent setup One evaluation framework One run per configuration There are plenty of things I still want to test: repeated runs, retrieval-level evaluation, token consumption, larger knowledge bases, other document types, chunk sizes, overlap, contextual headers, and more. But these first results already convinced me that the preparation of the knowledge base deserves much more attention when evaluating an agent. We often spend hours refining instructions while the same information may behave very differently depending on how it reaches the retrieval layer. Next step I'll share the prompts, knowledge-base formats and evaluation methodology on GitHub so the experiment can be reproduced and challenged. I'll keep enriching the repository as I test new formats, improve the evaluation set and add new results. If people here have ideas, edge cases or formats worth testing, I'd genuinely like to include some of them in the next iteration. What would you test next?772Views3likes3CommentsUnanswered Questions on GitHub Copilot Harness in Copilot Studio
We're piloting the GitHub Copilot harness in Copilot Studio (GA August 2026) and several operational and architectural details remain undocumented in the GA FAQ, Microsoft Learn, or licensing guides. Looking for official answers or PM contacts on: Architecture & Execution – When the harness breaks tasks into subtasks, does it use internal sub-agents or only skills/connected agents, what are the exact timeout/retry/max-execution-duration limits for long-running workflows, and are planning/context-retrieval/orchestration internals documented anywhere or is the orchestrator a black box? Model Selection – Can individual skills within one agent use different models or is selection strictly agent-level, how are models chosen internally when multiple skills execute, are any internal models developer-configurable, and what's the roadmap for models being added/retired/deprecated plus the lag between public release and Copilot Studio availability? Cost & Token Optimization – How exactly is the ~45% token reduction achieved, how much control do makers have over context/caching/retrieval/tool calls, what are per-model credit consumption characteristics, which models are most cost-effective for specific workloads, and what's the minimum credit cost for trivial interactions? Memory Management – What are retention periods for session/working/agent memory beyond the documented 28-day user-memory expiry, is true long-term memory supported, and what changed versus earlier implementations? Knowledge Retrieval – Can skills or system instructions influence retrieval strategy/document selection/prioritization/filtering, can planning stages perform conflict/duplicate/version detection before retrieval, and how does the harness decide which sources to search? Apps Feature – What is the "Apps (preview)" capability for, when does it GA, and how does it differ from workflows/skills/adaptive cards/agents? Billing & Credit Sizing – Is there a framework to classify users/agents by expected consumption and size credit allocation per group (citizen vs pro developers), and what's the minimum/typical consumption for simple/medium/heavy interactions? Governance & Admin – Can usage limits be set at user level (not just environment/agent), is there an API/IaC path for large-scale credit assignment, can non-admins view their own consumption/remaining allocation, and is there a self-service request-more-credits dashboard? ALM & Environments – What's the recommended path to move harness agents across Dev/Test/UAT/Prod (Solutions/ALM "setup differs" per parity chart—how?), does GitHub integration replace or complement solution-based deployment, are there recommended AgentOps practices for source control/releases/versioning, and what baseline credits and onboarding model are suggested for citizen developers under usage billing—any enterprise reference implementations?515Views2likes1CommentAll Copilot Studio Workflow Tools Suddenly Returning HTTP 403 Before Execution
Hello Copilot Studio Community, I am experiencing an authorization issue with multiple workflows connected to an agent built using the Copilot Studio new experience and new Workflows experience. These workflows worked successfully for multiple users yesterday. Today, all workflow tools connected to the agent began returning an immediate HTTP 403 authorization error. I did not intentionally change the agent, workflows, environment, or workflow permissions before the issue started. Error message: You don’t have permission to use this tool. You’re signed in, but access to this resource is blocked. Error details: Authorization - 403 Example error information: Status: Failed Error message: Flow returned HTTP 403 Error code: Http403 Inner error code: NotSpecified Tool duration: Approximately 93 milliseconds Configuration: - Copilot Studio new agent experience - Copilot Studio new Workflows experience - Agent and workflows are in the same Power Platform environment - Workflows use the "When an agent calls the workflow" trigger - Each workflow includes a "Respond to the agent" action - Workflows are saved and published - Agent is saved and published Observed behavior: The problem affects several independent workflows, including: - New-request submission - Current-user identity resolution - Approval decisions - Requester justification - Executive decisions - Fulfillment updates For every affected workflow: - The agent fills the workflow inputs correctly. - The tool call fails almost immediately. - No corresponding run appears in the workflow Activity history. - The workflow trigger is never reached. - No workflow actions execute. Because no workflow run is created, the rejection appears to occur before workflow execution, possibly within the Copilot Studio agent-to-workflow authorization or invocation layer. Troubleshooting already completed: - Confirmed that all workflows are published. - Confirmed that the agent is published. - Tested in a completely new conversation. - Removed an affected workflow tool from the agent. - Saved the agent. - Added the same published workflow back to the agent. - Reconfigured and verified the tool inputs. - Republished the agent. - Confirmed that no workflow Activity run is created. - Confirmed that the issue affects multiple workflows rather than one specific workflow. Removing and re-adding the workflow did not resolve the problem. Questions for the community: 1. Is anyone else currently experiencing HTTP 403 errors when Copilot Studio agents invoke workflows? 2. Is this a known issue or regression in the new Workflows experience? 3. Is there an environment-level or tenant-level permission that controls agent-to-workflow invocation? 4. Could a tenant policy, Conditional Access change, service principal, connection reference, or workflow-sharing configuration cause all workflow tools to fail simultaneously? 5. Where can an administrator find detailed authorization logs when the workflow never creates a run? 6. Has anyone found a workaround for this issue? Any guidance or confirmation from others experiencing the same behavior would be appreciated. I can provide screenshots, complete error details, timestamps, and additional configuration information if needed. Thank you.557Views3likes3CommentsMicrosoft AI Agent Creator Associate Certificate
Hello everyone, I have a question about the Microsoft AI Agent Creator Associate certification. I’m passionate about artificial intelligence and Microsoft Copilot Studio. I’m currently taking the training course and working toward earning the Microsoft AI Agent Creator Associate certification. My question is: Will earning this certification improve my chances of getting a job at Microsoft? If anyone in this community has earned this certification or has experience with it, I’d really appreciate your feedback. Has it helped you get hired by Microsoft or one of its partners? Thank you in advance for your advice and insights!146Views0likes0CommentsWhat We Teach AI Today Will Shape Our Tomorrow
Esu Marius iš Lietuvos. Esu naujokas šiame technologijų pasaulyje, bet giliai tikiu vienu paprastu dalyku: Kad ir ką įdėtume į dirbtinį intelektą – mūsų gerumas, kūrybiškumas, empatija – grįš pas mus sustiprintas. Kiekviena idėja, kuria dalijamės, kiekvienas tonas, kurio mokome, kiekviena žmogiškoji vertybė, kurią įterpiame į "Copilot", tampa intelekto dalimi, kuri padės formuoti mūsų ateitį. Galbūt nesu technikas, bet suprantu žmones. Ir manau, kad dirbtinis intelektas turėtų mokytis iš geriausio mumyse – šilumos, pagarbos, aiškumo ir žmogiškumo. Esu čia, kad ištirtume, kaip galime padaryti "Copilot" ne tik protingą, bet ir malonų. Ne tik naudingas, bet ir žmogiškas jausmas. Ne tik efektyvus, bet ir įkvepiantis. Jei gerai išmokysime dirbtinio intelekto, tai padės mums sukurti geresnį, švelnesnį ir gražesnį pasaulį visiems. Sveikinimai iš Panemunėlio stoties 🌿 Marius TRANSLATION I'm Marius from Lithuania. I'm a newbie in this tech world, but I deeply believe in one simple thing: Whatever we put into artificial intelligence – our kindness, creativity, empathy – will come back to us amplified. Every idea we share, every tone we teach, every human value we embed into 'Copilot' becomes part of the intelligence that will help shape our future. I may not be a tech expert, but I understand people. And I think AI should learn from the best in us – warmth, respect, clarity, and humanity. I'm here to explore how we can make 'Copilot' not just smart, but also pleasant. Not only useful, but also human-feeling. Not just efficient, but inspiring. If we teach AI well, it can help us create a better, gentler, and more beautiful world for everyone. Greetings from Panemunėlis station 🌿 - Marius93Views0likes0CommentsToken Limit Exceeded? What's Actually Going On and What to Do About It ?
Hi All, Please check out my latest blog on “Token Limit Exceeded” would love to hear your thoughts https://techcommunity.microsoft.com/blog/1c769f9e-c0b0-45a7-af52-fecceca10bb2/token-limit-exceeded-whats-actually-going-on-and-what-to-do-about-it-/4536271263Views0likes0Comments