Introducing auto: a new retrieval reasoning effort mode that recovers 97.6% of the evidence recall of the highest effort level at 46.9% fewer query planning tokens.
Enterprise retrieval workloads mix simple lookup queries with complex questions that require searching across many documents. Today, customers of Foundry IQ (Azure AI Search) Knowledge Bases must choose a fixed retrieval reasoning effort level: minimal, low or medium. That forces customers to make a trade-off across their workload: pay for deeper retrieval even on simple questions, or reduce costs and risk missing the evidence needed to answer more complex ones.
The new auto retrieval reasoning effort, available in the August 2026 preview, reduces the need to choose one fixed effort level for mixed workloads. auto assesses each query and adjusts the query planning and iterative search effort it receives. Straightforward questions take a faster, lower-cost path, while complex questions get the depth of multi-step retrieval, without manual tuning.
Customers can still use minimal, low or medium when their workloads are predictable. For mixed workloads, auto lets each query determine the effort it needs.
Near-identical retrieval quality, lower cost
Figure 1 shows the central trade-off across retrieval reasoning effort tiers. Higher fixed effort generally retrieves more complete evidence, but it also consumes more tokens. auto assigns lower effort to straightforward queries and reserves medium (the highest retrieval reasoning effort) for the queries that need it, retaining 97.6% of its evidence recall while reducing average query planning token usage by 46.9%.
Figure 2 translates the token savings into customer cost. With GPT-5.4 mini, auto reduces the average cost of 1,000 queries from $13.82 to $9.03 (a saving of $4.79, or 34.7%) compared with running every query at medium effort. Because search costs remain fixed, the dollar savings increase when a more expensive model is used for query planning.
Effort that scales with the question
Not every question needs the same retrieval depth. The auto retrieval reasoning effort matches retrieval depth to each question: minimal or low for straightforward requests, and medium when the question requires additional planning and iterative search. Figure 3 shows how those routing decisions vary across datasets. In the evaluated enterprise workloads, most queries follow the faster minimal or low paths, while datasets containing more complex questions send a larger share of queries to medium.
auto across the evaluated enterprise queries. KS = knowledge source.This routing behavior translates directly into the token savings shown in Figure 4. auto remains close to medium in evidence recall while using fewer tokens across every evaluated dataset. The datasets that route a larger share of queries to minimal achieve the greatest savings, because more queries avoid reasoning tokens when additional retrieval depth is unlikely to help. For a direct lookup query, this can mean using no reasoning tokens instead of the approximately 14,000 consumed by the average medium run.
auto and medium by dataset. The same adaptivity works in reverse. On BrowseComp, a public benchmark of deliberately difficult, multi-step research questions, auto escalates 78% of queries to medium. Its resulting evidence recall remains within a few points of running every BrowseComp query at medium, showing that the router increases effort when simpler retrieval paths are unlikely to be sufficient.
It is worth knowing the edges. Choose medium when most queries are complex, multi-step questions and retrieval quality matters more than cost. Choose minimal when the workload consists almost entirely of direct lookups. For mixed workloads, auto is the recommended starting point because it preserves close to medium-level evidence recall while reducing average cost without per-query configuration.
Figure 5 shows how auto applies per-query routing across three queries of increasing complexity: a direct lookup, a simple multi-hop question, and a complex multi-hop question. auto selects the lowest retrieval effort expected to preserve evidence quality, escalating from minimal to low or medium as the query requires more planning and retrieval steps.
Get started
auto retrieval reasoning effort is now available in preview for Foundry IQ (Azure AI Search) Knowledge Bases. For workloads that mix straightforward lookups with complex questions, auto provides a starting point without requiring you to choose a fixed effort level for every query.
If you already have a knowledge base, set retrievalReasoningEffort to auto as its default, or specify auto on an individual retrieve request to override that default for the request. See Set the retrieval reasoning effort for configuration details and supported versions.
If you are new to knowledge bases, follow the agentic retrieval quickstart to create one and run your first queries. Then try auto on a representative set of your own questions and compare the retrieved evidence and query-planning token usage with a fixed effort level.
Appendix
Similarly to our previous blog posts (Foundry IQ: Improve recall by up to 54% with knowledge bases | Microsoft Community Hub, Up to 40% better relevance for complex queries with new agentic retrieval engine) we evaluated auto on several benchmark query sets.
- Customer datasets: customer-provided corporate and member-document collections in English, covering domains such as oil and gas corporate reports and health-insurance member documents. We evaluate them as single knowledge source (KS) workloads over chunked PDF indexes.
- SEC: SEC filings of US public companies in English, evaluated in two variants: a single knowledge source setup spanning all sectors, and a routing setup with one knowledge source per Global Industry Classification Standard (GICS) sector.
- MIML: a multi-industry, multi-language corporate document benchmark in English, French, and Simplified Chinese. We evaluate both single knowledge source variants restricted to one language-industry slice and a routing variant spanning multiple languages and industries.
- BrowseComp: We use the 830 human-verified BrowseComp-Plus queries and indexed the full corpus, including distractor documents, into 512-token chunks with OpenAI text-embedding-3-large. To support continuous evidence-recall measurement, we decompose each question into multiple atomic factoid evidence nuggets. Evaluations use a two-agent neutral user plus search agent configuration with a standard system prompt and per question effort capped at 40 tool calls.