discovery engine
1 TopicAdaptive by Design: How Microsoft Discovery Explores Science
In Microsoft’s ALE evaluation, Discovery Engine with CLIO enabled leads publicly benchmarked agentic harnesses across three scientific domains. Microsoft Discovery is incorporating CLIO into Discovery Engine. Discovery Engine can effectively test competing hypotheses, learn from unsuccessful approaches, and revise reasoning toward the strongest path for success. Microsoft Discovery is incorporating CLIO (Cognitive Loop via In-Situ Optimization) into the latest releases of Discovery Engine, now available in the Discovery app and coming soon to the enterprise platform. Discovery Engine using CLIO mode achieved higher scores than all other publicly benchmarked agentic harnesses across the three scientific domains evaluated to date. With CLIO mode enabled, Discovery Engine can effectively explore competing hypotheses using tools and models, evaluate evidence, learn from unsuccessful approaches, and adjust its reasoning as new information emerges. To test the value of this update to the Discovery Engine, we ran the Agent’s Last Exam (ALE) benchmark with GPT-5.6 Sol as the default model while providing dynamic access to a broader mixture of models. Discovery Engine scored 61.6% in health and medicine versus 57.2% (+4.4%) for Codex with GPT-5.6 Sol, 75.2% in physical sciences versus 66.6% (+8.6%) Claude Code with Opus 5, and 64.6% in life sciences versus 60.8% (+3.8%) Codex with GPT-5.6 Sol. Built for the Scientific Mindset It is well established that agentic harnesses are important for accomplishing useful work. GitHub Copilot, Claude Code, Codex, and many others have made significant contributions on top of LLMs. However, these harnesses have largely been focused on workflows as they relate to software development, rather than scientific research. Discovery provides a harness for the rest of science and engineering, bringing both the scientific mindset of ideation and evidence collection, and the engineering rigor of problem decomposition and structured work to a collaborative environment. Importantly, CLIO reflects on regrets, treating unsuccessful paths as valuable information that informs its next course of action: it identifies why an approach failed, updates downstream reasoning, and redirects exploration without discarding what was learned. With Discovery Engine’s CLIO mode, we are rapidly increasing the ability to explore ideas that do not have a clear or well-researched solution. Figure 1 – Discovery Engine with CLIO enabled increased performance across all three scientific domains and provides consistency compared to individual models. Ultimately, this test of Discovery Engine's reasoning process and ability to use multiple models shows the benefit of how to increase the outcomes above any of the available models. The benefits increase further as you move from bounded direct workflow execution through exploration of new ideas and designs, but these results show that the increase even holds across a broad swath of scientific questions. Adaptive by Design Discovery Engine’s CLIO mode expands ideation and reflection to identify better options in addition to using multiple fallback models when the base model underperforms. Built using GitHub Copilot as the base harness, CLIO adds additional depth and exploration that is fit for purpose for the technical problems encountered in science and engineering. As the model ecosystem becomes increasingly diverse across performance and efficiency, having the flexibility to change is crucial to always adopting model strengths. One example of this benefit is in pursuit of Agent’s Last Exam questions where Discovery Engine switches models dynamically when facing false rejections or is otherwise unable to answer questions sufficiently using the base model. If we compare Discovery Engine to using a single specific model and harness combination, as you would experience when using a model-provider specific solution, the performance gaps expand dramatically (e.g., to +18.3% in physical sciences versus Codex with GPT-5.6 Sol, and +8.6% from Claude Code with Opus 5). The progress of ALE science tasks represented as performance average across life sciences, physical sciences, and health & medicine questions. Discovery Engine's harness and use of multi-models is crucial to improving performance and consistency. About Agent's Last Exam ALE is a large-scale, multi-disciplinary rolling benchmark released in June 2026, with 152 publicly available tasks spanning 13 professional domains and 55 subfields. Unlike benchmarks centered on question answering, ALE focuses on measuring agentic performance on tasks ranging from minutes to hours, completing deliverables using real data and tools. The questions in ALE are sufficiently difficult to stress the benefits of this update to Discovery Engine and importantly test the ability to react to scientific tools which are gathering new information in a long running agentic process. Consistent with Microsoft Discovery’s scientific focus, our evaluation focuses on three of ALE’s thirteen top-level domains: Physical Sciences, Health & Medicine, and Life Sciences. Example tasks include clinical variant annotation, molecular simulation, epidemiological forecasting, and medical-image analysis. For consistency with public benchmarking, we mirror ALE’s Strengths by Domain analysis, which reports domain level scores only for domains containing at least 10 tasks. We therefore exclude domains below that threshold as the published analysis does not provide corresponding comparison scores. Furthermore, as specified in the analysis, we benchmark comparably by following the five-hour time limit evaluation protocol rather than the two-hour limit. All attempts were immediately terminated after five hours and graded on the work produced, allowing partial credit for completed portions of the task. With CLIO mode, Discovery Engine can explore each ALE task through multiple independent approaches, within the same five-hour limit. By comparing these approaches, Discovery Engine identifies gaps, shares learnings across belief states, and successfully resolves the strongest path for each task into a single, evidence-backed answer. Accordingly, all reported domain scores are presented as Pass@1, demonstrating not only the performance gain but also the increase in reliability. Get Started Today While benchmarks like Agent’s Last Exam are useful in demonstrating Discovery’s utility for execution of scientific workflows, it has also generated real-world outcomes for high impact scenarios. Discovery Engine with CLIO mode was recently used to discover a novel organic redox flow battery negolyte. In the coming weeks, we will continue to demonstrate additional scientific use cases from chip design to radio frequency (RF) engineering to wastewater epidemiology, showing the depth and breadth of scientific applications which Discovery can enable. We are excited to see how Microsoft Discovery and CLIO can enable your hardest scientific or engineering problems. The Microsoft Discovery app is available in preview as a localized experience that gives researchers, students, academic labs, and scientific teams a simpler way to begin using Microsoft Discovery capabilities without starting with a full enterprise deployment. It is available for download on the Microsoft Discovery GitHub and users can get started with a GitHub Copilot account. Co-authored by Christine Caggiano, Joshua Bradley, Steven Truitt, and William Chappell395Views3likes0Comments