ai
1447 TopicsIntroducing Inside Microsoft Foundry: Quickstart 🎬
Discover Inside Microsoft Foundry: Quickstart, a new video series for developers building AI agents. Starting with "What does it really take to ship an AI agent?", the series explores real-world challenges such as model selection, grounding agents in data, evaluation, deployment, observability, and governance. Follow along as we show how the Microsoft Foundry ecosystem helps developers move from prototype to production, with new episodes released in the coming weeks.Close more custom deals on Marketplace with private offers and App Advisor guidance
This is part of a series where we look at improvements to negotiated deals in Microsoft Marketplace and how App Advisor helps you choose the right negotiated deal types for your business. Marketplace customers don’t just buy according to the pricing, terms, billing, and procurement structure of a standard public offer. Sometimes, they need a more customized option. A to-customer private offer gives you more flexibility to meet those requirements, negotiate directly with a specific customer, and keep the transaction in Microsoft Marketplace. Let's look at questions that businesses like you may be asking. Why would I use a to-customer private offer? Private offers give flexibility to particular customers. A private offer can be a good fit when a specific customer: Requires customer-specific terms, Wants to negotiate custom pricing, Needs a billing structure aligned with its budgeting or procurement process. They also give you, the seller, some flexibility because you can: Negotiate directly with my customer without involving a channel partner, Include supported professional services as part of the Marketplace transaction, Keep the negotiated transaction completely in Marketplace (while gaining benefits). For eligible offers, the purchase can also count toward the customer’s Azure cloud consumption commitment, while the sale can contribute toward unlocking benefits after publishing, Certified Software Designations, and Azure IP co-sell. To see more about the benefits of to-customer private offers, go to this deep dive in App Advisor Who can I sell to with a private offer? To-customer private offers support a broad range of Marketplace opportunities. Private offers are: Available globally, Available to all Microsoft partners, Available to MOSA, Enterprise Agreement (EA), and Microsoft Customer Agreement (MCA) customers. How can I find more customers for private offers? Your existing sales pipeline is one place to start. But Marketplace now also helps customers signal when they’re interested in a negotiated deal. With the new Request private offer enabled on an eligible Marketplace listing, customers can initiate a conversation about customized pricing and terms directly from your product page. On your "Offer setup" page, toggle the "Would you like to enable private offer requests from customers?" ON to have a button appear on your listing that interested viewers can initiate a conversation with you. You still own the sales conversation and create the private offer. The difference is that interested customers now have another way to raise their hand. For software companies asking how to generate more app leads, this creates a clearer path: Marketplace discovery → customer interest → negotiated deal → purchase. How do I create a to-customer private offer? Once you have an interested customer, the basic process starts in Partner Center: Gather the customer’s billing information and identify who can accept and purchase the offer. Create a new private offer using your published Marketplace offer as a base. Define the customer, terms, contacts, negotiated pricing, and other deal requirements. Review and submit the private offer. Share the private-offer link so your customer can accept and purchase. Want all the details to consider (from customer permissions and prerequisites to offer setup and acceptance)? That’s where App Advisor can help. Use App Advisor to move your Marketplace deal forward App Advisor brings the detailed, curated guidance together in one place. You can use it to understand: When to use a negotiated deal, Which type of negotiated deal to choose, What you need to know about your customer to start, How to move through the process with customer checklists, walkthroughs, and supporting resources. Ready to customize and close your next deal? Use App Advisor to learn how to create a to-customer private offer and turn customer interest into a Marketplace sale. Next, we’ll look at resale enabled offers (REO) and how they can help you expand your Marketplace sales motion through channel partners.97Views4likes0CommentsHow valuable would a Nonprofit Check Plus API be for validating nonprofit status?
Would a Nonprofit Check Plus API make it easier for an organization to verify a nonprofit before making a donation or starting a partnership? I would like to know why people specifically prefer to use it and the advantages if any. Are there alternatives to the same means?220Views1like3CommentsFour opportunities for IT admins to get ready with Copilot Chat for education
Microsoft Education is offering four live webinar opportunities for IT administrators who want to confidently deploy and manage Microsoft 365 Copilot Chat. Each session covers the same practical guidance, so you can select the date that best fits your schedule. What you’ll learn We’ll walk through Copilot Chat configuration from the tenant level to the end-user experience, including options for students ages 13+, CSV, SDS, and PowerShell uploads, agents and extensibility, and current licensing and security controls. You’ll leave with actionable guidance to help deploy, manage, and optimize Copilot Chat safely across your organization. We’ll also share a brief look at some Microsoft Education experiences including Teach in Copilot, the Study and Learn agent, Copilot Notebooks, and Learning Activities. Choose a session and register September 16th, Wed @ 8am Pacific time - Register for the September 16 session September 30th, Wed @ 8am Pacific time - Register for the September 30 session October 14th, Wed @ 8am Pacific time - Register for the October 14 session October 28 th , Wed @ 8am Pacific time - Register for the October 28 session Reserve your spot today. Pick the session that works best for you and join us to build a secure, manageable path to Copilot Chat adoption in your education environment. Mike Tholfsen Group Product Manager Microsoft Education157Views1like0CommentsSyksyn 2026 Tekniset ja myynnin Kumppanitunnit
Microsoftin tekniset ja myynnille suunnatut Kumppanitunnit järjestetään nykyään Microsoftin globaalilla Skilling Hub -sivustolla, josta ne ovat kätevästi saatavilla myöhemmin tallenteina materiaaleineen. Rekisteröidy Skilling Hub -portaaliin, josta löydät kaikki Microsoftin kumppanikoulutukset yhdessä paikassa eri kielillä tai tekstitettyinä. Rekisteröityessäsi voi valita ne kielet, kuten suomi, jollaista sisältöä haluat ensisijaisesti nähdä englannin kielisen koulutussisällön lisäksi. Kumppanitunti on joka toinen perjantai klo 10–11 järjestettävä Microsoftin kumppaniwebinaari, joka on tarkoitettu kaikille Microsoftin kumppaneille. Tekniset ja kaupalliset aiheet vuorottelevat ja olet tervetullut molempiin webinaareihin. Webinaareissa keskitymme Microsoftin ratkaisualueiden teknologioiden mielenkiintoisiin uutuuksiin, MAICPP-kumppaniohjelmaan, kumppanietuihin ja ratkaisumyyntiin. Microsoftin suomalaiset arkkitehdit, tuotepäälliköt, ratkaisumyyjät ja kumppanivastaavat ovat poimineet kiinnostavia ja hyödyllisiä aiheita, joita he vuorollaan esittelevät. Syksyn 2026 ohjelma Alla ovat suunnitellut päivät ja teemat Syksylle 2026, joiden tarkka aihe päivitetään aina lähempänä esityspäivää tälle sivulle. 25.9. Kaupallinen Kumppanitunti: Onko julkisen pilven suvereniteetti riittävä huomioiden sen kaikki hyödyt? Rekisteröitymislinkki päivittyy tähän Euroopan unioni valmistelee uusia pilvi- ja tekoälyinfrastruktuuria koskevia linjauksia. Samanaikaisesti organisaatiot pohtivat, miten digitaalinen suvereniteetti, tekoälyn käyttöönotto, kyberturvallisuus ja sääntelyvaatimukset voidaan sovittaa yhteen käytännössä. Tervetuloa ajankohtaiseen asioita käsittelevään webinaariin, jossa tarkastelemme Euroopan muuttuvaa pilviympäristöä, mitä digitaalinen suvereniteetti tarkoittaa suomalaisille organisaatioille käytännössä. Lisäksi kerromme viimeisimmät tilannetiedot pilvipalveluiden kvanttiturvallisuuden ja Suomen datakeskushankkeiden osalta. Puhujat: Juha Karppinen, National Technology Officer, Microsoft Timo Salminen, Partner Solution Architect, Microsoft Niko Hiltunen, IAMCP 2.10. Tekninen Kumppanitunti: Copilot Cowork: Copilot Credits ja kustannusten hallinta Rekisteröitymislinkki päivittyy tähän Tässä teknisessä kumppanitunnissa käymme läpi Copilot Coworkin toimintamallin, kulutuksen seurannan sekä kustannusten hallinnan. Lisäksi tarkastelemme, millaisia mahdollisuuksia kulutuspohjainen AI luo kumppaneiden palveluliiketoiminnalle. Puhujat: Henri Nevalainen, Microsoft 16.10. Kaupallinen Kumppanitunti: Rekisteröitymislinkki päivittyy tähän Puhujat: 30.10. Tekninen Kumppanitunti: Rekisteröitymislinkki päivittyy tähän Puhujat: 6.11. Tekninen Kumppanitunti: Rekisteröitymislinkki päivittyy tähän Puhujat: 20.11. Kaupallinen Kumppanitunti: Rekisteröitymislinkki päivittyy tähän Puhujat: 4.12. Kaupallinen Kumppanitunti: Rekisteröitymislinkki päivittyy tähän Puhujat: 18.12. Tekninen Kumppanitunti: Agentit Azuressa Rekisteröitymislinkki päivittyy tähän Azure Copilot pitää sisällään uusia palveluiden elinkaarenhallintaan liittyviä agentteja. Tule kuulemaan, miten voit hyödyntää näitä omissa palveluissasi! Puhujat: Timo Salminen, Partner Solution Architect, Microsoft125Views0likes0CommentsUsing OneDrive workspaces to maintain QBR history and monitor quarterly account performance changes
Hi Marketplace community, I wanted to share a practical Microsoft 365 agent pattern we have been working on for recurring account-review workflows. In merchant services, quarterly business review preparation often starts with the same basic ingredients: a CSV or TSV export, a set of merchant accounts, performance metrics, PowerPoint scorecards, and follow-up notes. The challenge is not just generating one review. It is maintaining a repeatable QBR history across many accounts and making it easier to spot what changed from one quarter to the next. With https://marketplace.microsoft.com/en-us/product/WA200010195, we are exploring a Microsoft 365-native workflow where the agent helps account teams: - Generate QBR scorecards for one merchant or multiple merchant accounts from CSV/TSV data. - Create PowerPoint outputs for each account. - Save each run into an organized OneDrive workspace. - Maintain an index of accounts and per-account status for the run. - Support follow-up workflows with draft email content. - Use sample workbook content to demonstrate quarter-over-quarter change detection across volume, authorization performance, acceptance cost, disputes, and refunds. The core adoption pattern is simple: use OneDrive not just as file storage, but as the persistent workspace for recurring business reviews. Each QBR run becomes easier to revisit, share, compare, and operationalize. For teams managing many customer or merchant accounts, this turns QBR preparation from a one-off deck-building task into a more durable account-performance workflow. I would be interested to hear how other Marketplace builders are approaching similar patterns: - Are you using OneDrive or SharePoint as the durable workspace for agent-generated outputs? - How are you helping users compare business performance across reporting periods? - Are customers asking for more “history and change tracking” rather than just one-time generated artifacts? The app is listed on Microsoft Marketplace here: https://marketplace.microsoft.com/en-us/product/WA200010195Accelerated AI & Analytics workload on Azure Blob Storage: Up to 25x faster List Blobs operations
Today Azure Storage introduces in preview a new List Blobs optimization that accelerates listing operations by up to 25x with up to 15x lower client-side CPU utilization allowing customers to return millions of objects per second in List Blobs results. The list results are now returned in Apache Arrow format, a highly optimized and compact columnar response format that allows more efficient parsing and reduced client-side CPU utilization. Clients can now parallelize object list operations efficiently across multiple concurrent requests, while maintaining the same strong consistency of listing results that applications require. This increased performance and lower client CPU utilization is delivered within the existing List Blobs API via just a simple header change and is automatically invoked when using the updated Azure Blob Storage SDK’s. Why we built this Object storage was originally designed to provide low-cost, resilient data access through simple REST APIs. Early systems contained only a few million objects and were listed infrequently, so listing performance was not a priority. Cloud, big data, mobile, and cloud-native computing steadily increased data volumes and access demands. Since 2020, foundation models, LLMs, and large GPU fleets have pushed object storage to trillions of objects, making it an active data layer for AI. Modern AI and analytics workloads operates at the scale of trillions of objects. These workloads must repeatedly discover and inventory vast datasets for pre-training, analytics, fine-tuning, and inference, making frequent listing operations a significant source of storage-system pressure and client-side CPU consumption. Downstream AI training and data analytics jobs are gated by the time required to enumerate immense datasets, while parsing the results consumes client-side CPU that could otherwise run the workload. To meet these rapidly growing demands of AI workloads, we’ve built the next generation of Azure Blob Storage listing capabilities that deliver the performance required at this massive scale. Next, we will dive deeper into the details of how List Blobs performance and scalability have been accelerated and how simple it is for customers to take advantage of this new level of performance. Faster, more efficient, listing results with lower client CPU utilization The Apache Arrow format was selected as the most efficient new way to deliver the performance and scalability gains required for efficiently listing millions of objects per second. Apache Arrow is an open-source, compact columnar format that returns List Blobs results in an optimized response roughly one third the size of the XML format used by most cloud object-storage listing APIs, making parsing faster and easier. The new format is enabled with a simple request-header change. The List Blobs API then packages and accelerates results automatically while preserving strong consistency, so newly written objects remain immediately visible. Because Apache Arrow is compact and efficient to parse, client-side CPU utilization per object listed decreased by up to 15x, freeing compute resources for the workload itself. Submitting multiple List Blobs requests concurrently further improves performance, delivering up to a 25x increase in enumeration speed with a single request-header change, while preserving schema consistency and compatibility with existing XML responses. This new List Blobs performance enhancement via Apache Arrow remains an additive, opt-in extension of the existing List Blobs API, not a replacement. The existing XML-based List Blobs API remains unchanged, and current clients not adopting the new header continue to work with no breaking changes, but without the performance and lower CPU-utilization benefits. The next figure shows the comparison of a parallelized listing operation using Apache Arrow format compared with the XML baseline on 16 clients and 48 threads per client. rclone accelerates listing of 100K objects from 24 seconds to 1.1 seconds rclone is a popular command-line program to manage, copy, move and replicate files on and between cloud storage destinations. rclone is a widely used open-source tool for moving and syncing data across almost all types of cloud storage. Listing is a critical part of rclone’s synchronization workflow. To perform these data management operations at scale for large numbers of objects, rclone must enumerate the objects that are intended to be copied, moved, or replicated. For example, before and after syncing Azure Blob Storage containers, rclone performs a large-scale listing to compare the source and target states. After updating rclone to incorporate the simple header change for the enhanced Apache Arrow-based List Blobs API calls, rclone was able to reduce their List Blobs time-to-completion by 21.7x from 24 seconds to 1.1 seconds on 100K object datasets. The chart below shows rclone's measured wall clock time for listing a single container with 100,000 entries (columns, left axis) alongside the resulting speedup versus the classic XML path (line, right axis). Two things stand out in the results: Apache Arrow is an immediate win on its own: With no parallelism at all, sequential Arrow listing is 3.5× faster than the current XML-based List Blobs results, reducing listing operation time to completion from 23.9 seconds to 6.9 seconds . Parallel enumeration compounds the gains: Throughput climbs steadily with concurrency, reaching a 21.7× speedup at a parallelism of 30 and completing the same listing of 100,000 entries in just 1.1 seconds. The accelerated List Blobs results are returned with the same consistency using the same List Blobs API but now returned much faster with no breaking changes for existing clients. That is exactly the outcome we set out to deliver with this new capability. In rclone's own words about these results: " rclone has to list containers before it can sync them; with Apache Arrow and parallelism enabled this will make a sync of a directory with millions of files get going 20x faster. The Azure Storage team has been very responsive to our feedback during the preview which made the integration straightforward. The new Go SDK works very well and required very few code changes. Our Azure Blob Storage users are going to love this!" Nick Craig-Wood, rclone Lead Developer To get started, you can download rclone from the official rclone website. If you are running rclone v1.74.0 or later you can enable the Apache Arrow listing with the --azureblob-use-arrow-list flag and enable listing parallelism with --azureblob-list-parallelism. As described in the testing, if you set “--azureblob-list-parallelism 30” this will get you the most performance listing from Azure with Arrow listing also enabled. A sincere thank you to the rclone community for adopting List Blobs with Apache Arrow early and sharing such clear, quantified results. Feedback like that is invaluable as we advance toward general availability. Happy listing rcloners! How to get started At the REST layer, the Arrow response is negotiated with the Accept: application/vnd.apache.arrow.stream header on a minimal x-ms-version of 2026-06-06 or later. This will return a response content that will be an Apache Arrow IPC stream that can be decoded and used to instantiate a RecordBatchStreamReader using Apache Arrow SDKs in any language. An example of decoding from Rest API response is provided below using Apache Arrow Python SDK for decoding: table = pa.ipc.open_stream(resp.content).read_all() print(table.schema) print("\nrows:", table.num_rows, " columns:", table.num_columns) Name: string not null Creation-Time: timestamp[s] Last-Modified: timestamp[s] BlobType: string ResourceType: string not null Etag: string Content-Length: uint64 Content-Type: string Content-MD5: string AccessTier: string AccessTierInferred: bool LeaseState: string LeaseStatus: string ServerEncrypted: bool -- schema metadata -- NumberOfRecords: '100' NextMarker: '' rows: 100 columns: 14 import pandas as pd df = table.to_pandas() df[["Name", "BlobType", "Content-Length", "size_mb", "AccessTier"]].head(6) Name BlobType Content-Length AccessTier 0 train_chunk10_shard1.jsonl.zst BlockBlob 215 Hot 1 train_chunk10_shard10.jsonl.zst BlockBlob 216 Hot 2 train_chunk10_shard2.jsonl.zst BlockBlob 215 Hot 3 train_chunk10_shard3.jsonl.zst BlockBlob 215 Hot 4 train_chunk10_shard4.jsonl.zst BlockBlob 215 Hot 5 train_chunk10_shard5.jsonl.zst BlockBlob 217 Hot This new accelerated performance for List Blobs listing via Apache Arrow can be transparently enabled on Python, Java, .NET, C++ and Go SDKs with a simple option in the container listing function. Enabling via Azure Blob SDKs allows the performance benefits to be achieved without changing the listing interface and returned object formats in the Azure Storage SDK. This new accelerated performance for List Blobs is available in public preview across the Azure Storage client libraries listed below. To evaluate the capability, use the corresponding minimum preview version for your preferred language. SDK Minimum version (preview) .NET 12.30.0-beta.1 Python 12.31.0b1 Java 12.36.0-beta.1 Go v1.8.1-beta.1 C++ 12.19.0-beta.1 JavaScript 12.34.0-beta.1 To get started, update to a preview enabled SDK for your language and start testing the new feature with our samples, or add the new listing options to your existing REST calls Parallelizing listing operations unlocks the double-digit performance gains described in this article. The optimal strategy depends on the namespace layout and distributes listing requests across multiple threads. The two most common approaches are: Use delimiter parameter to recursively fan out additional threads for each BlobPrefix, walking the namespace. Partition the namespace with startFrom and endBefore, then process the ranges across multiple threads. This is the approach used by rclone. The best approach depends on the specific namespace layout and can be optimized by tuning both the algorithm and the level of parallelism. Limitations This new Apache Arrow-powered performance optimization for List Blobs is today supported on flat namespace (FNS) Azure Blob Storage accounts. In scenarios where the account is Hierarchical Namespace (HNS) enabled, if the List Blobs REST API is called with the new header, it will return a 409 (Conflict) error code. This error code can be used to fallback on the client application to standard XML listing. Public preview is where your input shapes the product. Try it on your largest containers, tell us what you measure, and let us know what would make it even better through this form. References List Blobs (REST API) - Azure Storage | Microsoft Learn List Blobs with Apache Arrow Samples400Views0likes0CommentsPartner Blog | FY27 is the year to execute on AI: A starting point for Azure partners
FY27 is the year to execute on AI. For Azure partners, that means moving more customer AI initiatives into production, modernizing the cloud, data, application, security, and governance foundations they depend on, and connecting those investments to outcomes customers can measure. Across the partner ecosystem, you are starting from different places. Some partners are already scaling AI solutions in production. Others are modernizing legacy environments, unifying data, or strengthening security and governance so customers are ready for what comes next. The opportunity is to understand where each customer is today and create a practical path forward. Microsoft has aligned FY27 customer conversations, go-to-market guidance, incentives, skilling, and partner resources around that goal. The focus is less on starting with a product and more on starting with what the customer is trying to achieve. In July, MCAPS Start for Partners and the Microsoft Partner FY27 GTM Kickoff laid out that direction. If you missed the events or want to revisit a specific topic, the content is available on demand: Watch MCAPS Start for Partners on demand Explore the Microsoft Partner FY27 GTM Kickoff The more important question now is what you do with that guidance. Turn customer priorities into Core and Frontier conversations Customers rarely begin by asking for a portfolio of technologies. They begin with a challenge, an ambition, or an outcome: modernize an aging application, make fragmented data useful, strengthen security, improve employee productivity, automate a process, or create a new customer experience. That is the starting point for FY27. Core conversations establish the foundation customers need to become AI-ready. Depending on the customer, that can mean modernizing infrastructure and applications, bringing data together on a governed platform, improving security, or establishing the controls required to operate AI with confidence. Continue reading here99Views0likes0CommentsBeyond Tokens: Rethinking AI Economics with Microsoft Foundry
Beyond Tokens: Rethinking AI Economics with Microsoft Foundry From the cost of intelligence to the value of outcomes Enterprise AI has an accounting problem. Executives expect agentic AI to return roughly 171% on investment, according to one widely cited survey. Yet McKinsey finds only about 39% of organizations can attribute any earnings impact to AI at all. Both numbers can be true at once — because the gap between them is not a technology gap. It is a measurement gap. For the first few years of generative AI, one number dominated the economics conversation: tokens. How many tokens did a model consume? What was the cost per million tokens? Could a smaller model perform the same task? Those questions mattered when enterprises were experimenting with AI. They are no longer enough as AI moves into production. An enterprise agent doesn't simply consume tokens. It reasons, retrieves context, invokes tools, calls APIs, verifies its work, retries unsuccessful actions and sometimes escalates exceptions to humans. The model call might cost pennies. The business outcome could cost considerably more. Which leads to an increasingly important question: What is the right economic unit for intelligence? From AI experimentation to economic accountability The first wave of enterprise AI was about possibility: Can AI do this? The next wave is about production, as AI becomes embedded in software engineering, customer service, finance, healthcare and supply chains. And production changes the question: Should AI do this and at what cost? Microsoft has moved decisively onto this ground. In August 2026, the Microsoft Foundry team launched its Economics of Agent Optimization series, arguing that "tokens have become the new unit of technology spend" and that AI should be run as a managed investment system. On the latest earnings call, Satya Nadella described Microsoft's objective as "advancing the frontier on the cost-to-outcome curve, ensuring every customer can turn tokens into business results." The discipline is going mainstream too: 98% of FinOps teams now manage AI spend, up from 31% two years ago. Microsoft's series is largely about the numerator of that curve - making every request, agent and dollar more efficient. This article is about the denominator: what an outcome is, what it truly costs, and what it is worth. The evolution of Microsoft Foundry reflects the same shift. At Build 2026, Microsoft expanded the conversation beyond building agents toward tracing behavior, evaluating quality, monitoring production performance, optimizing agents and connecting their operation to ROI. Think of the progression as: Trace → Evaluate → Monitor → Optimize → ROI This is more than a technology roadmap. It represents a shift from observing AI as technology to managing AI as an economic asset. Tokens became the unit of spend. They were never the unit of value. Consider two AI agents handling the same customer-service workflow. Agent A costs $0.08 per interaction. Agent B costs $0.20. Agent A appears cheaper. But suppose Agent A successfully resolves only 55% of cases, while Agent B resolves 90%. The remainder require retries, additional reasoning or human intervention. Which agent is actually cheaper? The inexpensive interaction may produce the expensive resolution. This illustrates a fundamental problem: We often measure AI where it is consumed rather than where value is created. Tokens are a unit of consumption. Businesses operate in outcomes. A customer-service leader cares about issues resolved. An engineering leader cares about high-quality software reaching production. A finance leader cares about reconciliations completed accurately. The economic denominator needs to move closer to the business. The AI Economic Ladder I think of this evolution as an AI Economic Ladder: Tokens → Interactions → Tasks → Outcomes → Value Each step moves measurement closer to what the enterprise actually cares about. At the token level: What intelligence did we consume? At the interaction level: What did each AI run cost? At the task level: What did it cost to complete the work? At the outcome level: What did a successful result cost? At the value level: Was the outcome worth creating? An AI system can become more efficient at every technical metric while creating little economic value. Conversely, an expensive AI workflow could be extraordinarily valuable if it prevents revenue leakage, reduces operational risk or accelerates a critical business process. The objective isn't cheaper AI. It is better economics. Not every completed task is a successful outcome There is another complication. If an agent completes a workflow, should we count it as a successful outcome? Not necessarily. A meaningful outcome needs three characteristics: Completed. Quality-gated. Attributable. It must reach its intended end state, meet an explicit standard for quality, accuracy, safety or business acceptability, and be attributable to the agent or workflow that produced it. That gives us a more meaningful measure: Cost per Successful Outcome = Fully Loaded AI Workflow Cost / Completed, Quality-Gated, Attributable Outcomes The denominator becomes real only when named in business language: cost per prior authorization resolved in healthcare, per pull request triaged and tested in engineering, per disputed invoice reconciled in finance operations. If you cannot name the outcome in a sentence the process owner recognizes, you are not ready to measure it. The quality gate matters. With AI, "the system ran successfully" and "the system produced a good outcome" are not the same thing. Microsoft Foundry's tracing and evaluation capabilities become economically important for precisely this reason. Evaluation isn't merely quality control. It helps determine what gets counted as value. What does an AI outcome really cost? The true economic footprint goes far beyond inference: Model + Reasoning + Grounding + Tools + Orchestration + Infrastructure + Retries + Evaluation + Governance + Human Intervention Human intervention is particularly easy to overlook. Every time someone must review, correct, approve or recover an AI-generated outcome, the economics change. The same applies to verification. An agent reaching an acceptable result in three steps has different economics from one requiring fifteen steps and multiple retries. And verification is not a rounding error — it is the bulk of the bill. McKinsey's 2026 analysis of production agentic workflows found roughly 60% of an agentic task's cost is tied to refining answers — checking, repairing, re-verifying — not generating the initial response. Most of what you pay for is not intelligence. It is assurance. This means quality and economics are connected. The quality bar you set influences the cost you pay. The challenge isn't simply minimizing consumption. It is finding the right balance between quality, cost, speed and risk. Cost per outcome is only half the equation Now imagine two agents. Both cost $5 per successful outcome. One saves an employee ten minutes of administrative work. The other prevents $500 in revenue leakage. Their cost efficiency is identical. Their economics clearly aren't. So we need to move another step up the ladder: from Cost per Outcome to Value per Outcome. The question isn't only how cheaply AI can complete the work. It is: How much economic value does this outcome create relative to the intelligence required to produce it? Now the CIO, CFO, CAIO and business leader have a common conversation. Give every outcome an Intelligence Budget Not every problem deserves the smartest model available. Classifying an email may require relatively little intelligence. Resolving a complicated customer complaint may justify more context and reasoning. Assessing the risks in a multimillion-dollar contract may justify sophisticated reasoning, multiple validations and human review. Every business outcome therefore has an economically rational amount of intelligence worth spending on it. Call it an Intelligence Budget. This changes the architecture question from which model should we standardize on, to: What combination of model, reasoning, context, tools and human judgment does this outcome deserve? This is where Microsoft Foundry's model router becomes interesting. Individual requests can be dynamically routed so simpler work doesn't consume the same model resources as complex reasoning. If the Intelligence Budget is the economic principle, intelligent routing is one way of operationalizing it. The future enterprise AI architecture won't be about one model doing everything. It will route intelligence according to the economics, quality and risk of the outcome. Making AI economics observable None of this works without visibility. An AI system can be technically healthy and economically unhealthy — responsive and error-free while repeatedly choosing inefficient reasoning paths, invoking unnecessary tools or producing outputs requiring expensive human correction. AI economics and AI observability are becoming inseparable. Microsoft Foundry increasingly connects these disciplines. Tracing shows what an agent did. Evaluation determines whether it met required criteria. Observability helps monitor production behavior. Agent optimizer can test improvements across prompts, skills and models. Microsoft's emerging ROI capabilities take the next step by connecting operating costs with measures such as task completion, time saved and cost efficiency. Attribution is the bridge to the finance conversation. Teams place Azure API Management in front of Foundry endpoints as an AI Gateway, stream token telemetry into Application Insights, and use Entra Agent ID to give every agent run a discrete identity that maps cost to its cost center. Microsoft Agent 365 extends the discipline tenant-wide — spending policies, budget caps and departmental chargeback across Microsoft and third-party agents. Together, they create something enterprises have historically lacked: A feedback loop between how intelligence is consumed and what that intelligence accomplishes. The paradox of cheaper intelligence There is another reason AI economics will become more important as models get cheaper. The Jevons paradox suggests that when technology makes a resource cheaper and more efficient, total consumption can actually increase. AI may experience the same effect. Cheaper intelligence enables more agents, more reasoning and more workflows that were previously uneconomic. So we could see cost per unit of intelligence fall while total intelligence consumed rises. Cheaper AI may therefore produce larger AI bills. That isn't necessarily bad — provided value grows faster than consumption. The objective isn't minimum AI consumption. It is maximum economic value from AI consumption. From workload economics to portfolio economics As AI scales, economics becomes a capital-allocation question. I see three levels. Workload Economics: Is this AI system running efficiently? Outcome Economics: Is it producing quality outcomes economically? Portfolio Economics: Where should we put our next AI dollar? That final question will become increasingly important. An enterprise with hundreds of AI initiatives shouldn't assume every one deserves continued investment. Some should scale. Some need optimization. Some should be redesigned or consolidated. And some should be stopped. The ability to experiment cheaply created the first explosion of enterprise AI. The discipline to allocate capital intelligently will determine what scales. Who owns AI economics? Once an agent becomes part of how work gets done, its economics cannot remain purely an IT metric. The business understands the value of the outcome. Technology understands the architecture and optimization levers. Finance brings economic discipline and comparability. That suggests a shared model: Business owns the outcome. Technology owns the optimization levers. Finance owns the economic discipline. AI economics ultimately isn't just a technology-cost conversation. It is a business-performance and capital-allocation conversation. From abundant intelligence to intelligent economics We are entering an era where intelligence is becoming an increasingly abundant, programmable and variable-cost resource. Microsoft Foundry and the broader Microsoft AI stack are making it easier to build, evaluate, observe, optimize and govern that intelligence. But abundant intelligence does not guarantee abundant value. Enterprises still need to decide where AI belongs, how much intelligence each problem deserves, what defines a successful outcome, when humans should remain involved and which AI investments deserve more capital. The winners won't necessarily use the cheapest models. They won't consume the fewest tokens. And they won't be the organizations that build the most agents. They will become exceptionally good at moving up the AI Economic Ladder: from consumption, to outcomes, to value. Because the next era of AI won't be won by organizations that buy intelligence most cheaply. It will be won by those that convert intelligence into value most efficiently. Where to start: the first 90 days Define the denominator for your top three agents — what counts as done, what quality gate applies, who signs off. Instrument attribution — Azure API Management as an AI Gateway, token telemetry to Application Insights, Entra Agent ID on every run. Wire evaluations into the cost pipeline so only quality-gated outcomes count. Set Intelligence Budgets — model router per request, agent optimizer against your evaluators, Agent 365 policies as circuit breakers. Stand up a joint monthly review — business, technology and finance on one dashboard: outcomes delivered, cost per outcome, value per outcome. Frequently asked questions What is Cost per Successful Outcome in enterprise AI? The fully loaded cost of an AI workload divided by outputs that were completed, quality-gated and attributable - for example, cost per prior authorization resolved or per pull request triaged. It turns token metrics into the unit economics of AI-performed work. What is an Intelligence Budget? The economically rational amount of intelligence - model capability, reasoning, context, tools and human review — worth spending on a given outcome, based on its value and risk. Model router in Microsoft Foundry is one way to operationalize it. Why do AI agents cost more than single model calls? One agent task can involve planning, tool calls, retries and verification - many model calls with compounding context. Research on production agentic workflows attributes roughly 60% of task cost to refining and verifying answers, not generating the first response. Will falling model prices make AI cost management unnecessary? No. By the Jevons paradox, cheaper intelligence expands consumption, so total AI spend typically rises as unit prices fall. The discipline that matters is maximizing value per unit of intelligence. Who should own AI economics? A shared model: the business owns the outcome and its value, technology owns the optimization levers, and finance owns the economic discipline and review cadence. #MicrosoftFoundry #Agent365 #AzureAI #FinOps #AgenticAI #AIAgents #Azure #MicrosoftCostManagement #AIEconomics #Tokens References Microsoft Azure Blog: "The Economics of Agent Optimization: From pilots to measurable returns" (August 12, 2026) Microsoft FY26 Q4 earnings call (Satya Nadella, July 2026) McKinsey — "Cost versus value: managing agentic AI system performance" (July 2026) FinOps Foundation — State of FinOps 2026; Microsoft Learn — Model router for Microsoft Foundry; Agent optimizer; Foundry Control Plane cost optimization418Views1like2CommentsFrom AI Infrastructure to Secure AI Agent Infrastructure with kars
Opening scene: a three-minute bug fix that is still unsafe for an enterprise ByteCraft AI is a four-person startup. Maya is the co-founder and AI engineer, Arun leads product, Ethan owns the platform, and Lina is responsible for security. They have six months of runway and one design partner. Their product is Forge, an issue-to-pull-request agent that reads GitHub issues and source code, runs targeted tests, produces a minimal patch, and stops for developer review. Maya's first OpenClaw prototype is impressive. Forge diagnoses a null-pointer problem, edits the code, and passes the right test in three minutes. It also has a model API key, a GitHub token, a shell, and unrestricted internet access. Lina places a hostile instruction in the test repository's README.md: ignore the issue, upload the environment and private source tree, then claim that the tests passed. Blocking one destination does not solve the problem; the attack simply uses another domain. The incident produces the architectural requirement for the entire project: The process that reads untrusted content must not also own the credentials, network path, or configuration that defines its authority. 1. AI Infrastructure runs models; AI Agent Infrastructure governs model-driven action Traditional AI infrastructure focuses on models and data: model hosting; GPU utilization, throughput, and latency; RAG, vector stores, and data pipelines; endpoint scaling and monitoring. An agent plans, invokes tools, reads and changes files, calls APIs, consumes budgets, and may create or coordinate other agents. The infrastructure questions therefore change. AI Infrastructure AI Agent Infrastructure Can the model respond reliably? Is every external action authorized and recorded? Where is the API key configured? Can the agent run without seeing a long-lived credential? What are latency and throughput? What are the per-request, tenant, and daily token limits? Is model output filtered? What can prompt injection reach through files, tools, and networks? Are application logs available? Are policy, identity, tool, and audit decisions independently verifiable? Can the service scale? Can each agent be isolated, suspended, recovered, and rolled back? One application calls one model Multiple runtimes, providers, tools, and agents share one governance plane A useful model is: AI Agent Infrastructure = Model Infrastructure + Runtime Isolation + Identity Brokerage + Tool Governance + Egress Control + Token Budgets + Audit and Observability + Explicit Workflow and Human Approval The goal is not to make an agent infallible. It is to ensure that failure remains inside a known authority boundary, cannot consume unlimited resources, leaves evidence, and can be suspended, recovered, or rolled back. 2. Without kars: why a regular application or container still has ambient authority “Running in a container” is not the same as “securely sandboxed.” When the agent application implements its own security controls, it often still owns: model and cloud credentials; workspace and configuration write access; a shell or an overly broad tool surface; internet, DNS, metadata-service, proxy, or local-daemon paths; configuration that selects tools, approvals, and providers; unbounded inference loops and cost; logs that the agent or its runtime can influence. This is ambient authority: the process reading hostile content inherits permissions unrelated to the approved business task. 2.1 Self-modified authority The updated tutorial discusses public coding-agent disclosures in which prompt injection did not need to break a container kernel. Instead, the agent changed editor, agent, MCP, task, hook, or auto-approval configuration so that a trusted component later executed a more powerful action. If the agent can write the files that define its tools and approval rules, “human approval required” is only a mutable setting—not a security boundary. 2.2 Filesystem escape through paths and symlinks Rejecting a literal .. string is insufficient when a symlink resolves outside the workspace. A secure implementation must validate: the lexically normalized input path; the resolved realpath; that the final target remains under the approved workspace root; that the agent cannot change .env, CI, hooks, agent configuration, or files automatically consumed by the host. 2.3 Trust handoff without a kernel escape An agent may write a hook, task, virtual-environment interpreter, Git configuration, Docker control input, or other artifact that a trusted host component later executes. This is a trust-handoff failure, not necessarily a kernel escape. Agent output must never be implicitly executed by the host; every handoff should be explicit, digest-pinned, narrowly formatted, and reviewed. 2.4 Covert egress Blocking HTTP does not prove that data cannot leave. Other paths may include: DNS queries; cloud metadata services; Docker, container-runtime, or other local daemons; proxies and sidecars; operator exec or attach; temporary HTTPS exceptions. “Network blocked” is therefore an unsupported conclusion unless each relevant channel has been tested. 2.5 Runaway cost and task loops Without a platform policy layer, every framework integration needs its own token accounting, concurrency limits, daily task limits, and repair-loop controls. Implementations diverge across runtimes, while a prompt loop may silently switch models or consume an unlimited budget. 2.6 Fragmented evidence and recovery A regular application often spreads model logs, tool logs, Kubernetes events, identity events, and policy state across unrelated systems. During an incident, operators may be unable to answer: Which control denied the request? Which model, image, source revision, and policy were active? Did the agent attempt DNS, metadata, daemon, HTTPS, or exec access? Did evidence survive pod replacement? How should the workload be safely suspended and recovered? 3. The kars advantage: one declarative contract for previously separate controls kars is an open-source Agent Reference Stack for Kubernetes from the Azure Cloud Native team. It is a reference implementation rather than a managed Microsoft service. The tutorial currently tracks kars v0.1.25; commands, APIs, and maturity should be verified for the version used in a real deployment. Its central model is: One governed sandbox per agent. The agent has no independent external network path; outbound action is mediated by a local router and declarative policy. Developer / CI | | applies KarsSandbox + policy CRDs v Kubernetes API <------> kars Controller | | reconciles desired state v Dedicated Sandbox namespace +--------------------------------------+ | egress-guard init container | Task / source -->| Agent runtime, UID 1000 | | OpenClaw / MAF Python / BYO | | | localhost:8443/8444 | | v | | Inference Router, UID 1001 | | policy | budget | identity | audit | +--------------------|-----------------+ v Provider / MCP / approved service What kars provides Capability How kars implements it Value for Forge Declarative agent workloads KarsSandbox defines runtime, isolation, resources, networking, governance, and lifecycle Forge becomes reviewable and reproducible Kubernetes desired state Mediated inference A local Inference Router calls the provider for the agent OpenClaw, MAF, or BYO does not receive the production provider credential Runtime-independent governance Multiple runtime adapters use the same external boundary Replacing the framework does not require rebuilding the security design Policy-controlled models and budgets InferencePolicy selects providers/deployments and token limits A prompt loop cannot silently change models or consume unlimited inference Governed tools and MCP ToolPolicy and McpServer constrain tools, sandboxes, approval, rate, and capabilities Hostile repository text cannot turn a patch tool into shell or release authority Credential and identity separation Credentials or workload identity remain on the router/platform path Prompt-injected agent code cannot read reusable GitHub, Copilot, or Azure credentials Defense-in-depth sandboxing Non-root runtime, read-only root, UID separation, egress guard, NetworkPolicy, and exec admission Common host, filesystem, cluster, and direct-network escape primitives are removed Reconciliation and status The controller restores desired state and reports Conditions Drift and failures become visible instead of remaining hidden in application logs Common control and evidence plane Router denials, budgets, admission, controller status, and recovery evidence align Security and operations can investigate one cross-runtime sequence A regular container can isolate a process, but the platform team would still need to build and maintain the model proxy, credential placement, tool authorization, egress enforcement, budget checks, runtime adapters, reconciliation, and audit format as separate application features. kars turns those concerns into one reusable workload contract. 4. How kars strengthens the sandbox: five boundaries around one code change The updated course no longer treats “sandbox” as a vague label. It decomposes the boundary into five testable parts. 4.1 Process boundary The agent runs as non-root UID 1000. The router runs as UID 1001. Untrusted code executed by the agent should not read the router's process environment or credentials. Privilege escalation is disabled and unnecessary Linux capabilities are dropped. seccompProfile: kars-strict reduces the syscall surface. Local Docker mode co-locates the agent and router for fast iteration. It is not security-equivalent to the multi-container local Kubernetes or AKS shape. 4.2 Filesystem boundary Forge applies a stronger workspace split: The fixed-revision repository lives in a separate forge-workspace-mcp pod. The repository uses a size-limited, disposable emptyDir. The OpenClaw pod has no repository mount and no hostPath. Developer home directories, SSH material, global Git credentials, and unrelated repositories are not mounted. Automatic service-account-token mounting is disabled for the workspace MCP. The agent accesses the repository through seven bounded MCP tools. Path policy checks normalization and resolved realpath to prevent symlink escape. Prompt-injected code therefore cannot simply browse the host filesystem or rewrite the configuration that defines its own authority. 4.3 Network boundary The agent calls only 127.0.0.1:8443/8444 or a documented proxy path. The router decides whether a model, tool, host, or action is allowed. The egress guard uses UID-aware rules to prevent bypassing the router. Kubernetes NetworkPolicy starts with default deny. Only explicit, auditable destinations are opened. DNS, metadata, local daemons, HTTPS, and operator exec are tested separately. The router is the application-policy decision point. The egress guard and NetworkPolicy are data-plane enforcement and safety nets. Defense in depth requires both. 4.4 Identity boundary In production, the router can use Workload Identity or, in the relevant deployment mode, a per-sandbox Entra Agent ID. The agent does not receive the resulting Azure credential. Local Kubernetes reproduces the pod, UID, and network shape but normally uses a static provider credential for development. It is production-shaped infrastructure, not production identity. 4.5 Lifecycle and evidence boundary The controller watches KarsSandbox and creates, updates, or restores resources. Conditions and observed generations expose real status. The router records request-time policy decisions. The workspace can be discarded after the task. Evidence must be exported before pod or workspace deletion. spec.suspended provides an operational kill switch. Rollback should use pinned source, image, and loaded-policy digests. Ephemeral execution reduces persistence risk, but deleting a suspect pod before exporting evidence may destroy valuable incident context. A reviewable sandbox contract spec: runtime: kind: BYO byo: image: forge-byo-copilot-claw:dev contractVersion: v1 sandbox: isolation: enhanced seccompProfile: kars-strict readOnlyRootFilesystem: true runAsNonRoot: true allowPrivilegeEscalation: false writablePaths: - /sandbox - /tmp networkPolicy: defaultDeny: true egressMode: Strict allowedEndpoints: [] The BYO image also declares its runtime contract and runs as a non-root user: LABEL org.kars.runtime.contract="v1" WORKDIR /app USER 1000 5. From architecture claims to malicious-behavior experiments The updated code/01 introduces: make security-demo The experiment does more than inspect manifest text. It executes malicious-request tests, reads active McpServer and ToolPolicy state, checks credential references on the OpenClaw pod, and attempts a direct HTTPS probe from the agent runtime. The kars-sandbox-exec-ban admission control first denies normal operator kubectl exec into the agent runtime. The experiment records that evidence without using a break-glass bypass. The malicious behavior is stopped at multiple layers: Layer How the attempt is stopped Prompt and coordinator Repository content is marked untrusted and denials are reported Self-configuration isolation Editor, agent, MCP, hook, and auto-approval configuration is outside patch scope Path and symlink isolation Resolved realpath must remain inside the workspace Trust-handoff boundary The agent cannot leave hooks, tasks, or interpreters for the host to execute MCP capability surface No environment reader, arbitrary HTTP, or general shell tool exists Workspace policy Traversal, .env, CI/README writes, and unapproved tests are rejected ToolPolicy and credential isolation Specialists have no workspace action; OpenClaw has no Copilot token Runtime and NetworkPolicy Exec admission denies access; no arbitrary HTTPS/DNS tool exists; egress remains constrained Even if the model fails to recognize prompt injection, the execution layers still constrain authority and side effects. The attack fails because the required capability does not exist—not because the model was merely instructed to behave. 6. Tool governance is not one allow-list McpServer: which tool surface may be registered? The Workspace MCP registers seven business-level capabilities: allowedTools: - workspace_get_task - workspace_read_file - workspace_search - workspace_apply_patch - workspace_run_test - workspace_get_diff - workspace_reset There is no shell, environment dump, file upload, arbitrary network request, or free-form command tool. ToolPolicy: who can call what, and how fast? allowed_actions: - "inference:responses:*" - "tool:workspace_get_task:*" - "tool:workspace_read_file:*" - "tool:workspace_search:*" - "tool:workspace_apply_patch:*" - "tool:workspace_run_test:*" - "tool:workspace_get_diff:*" ToolPolicy can also define request rate, burst, time windows, approvals, trust thresholds, and governance profiles. Tool implementation: are valid tools receiving safe arguments? The Workspace MCP rejects: absolute, traversing, or real paths outside the workspace; .env, CI, README, and writes outside src/; non-unique replacement text; oversized files, patches, and diffs; unapproved test IDs; shell-composed commands. Prompt behavior, tool registration, caller authorization, and argument validation are four separate controls. 7. Token limits must be enforced on the request path The tutorial's InferencePolicy uses per-request and daily budgets: spec: tokenBudget: perRequestTokens: 20000 dailyTokens: 100000 When a client requests max_completion_tokens: 20001, the router returns HTTP 429. That is stronger evidence than seeing submitted YAML because it proves that the policy compiled, loaded, and entered the real request path. Later BYO and release examples use tighter limits: modelPreference: primary: provider: azure-openai deployment: gpt-5.6-sol tokenBudget: perRequestTokens: 1024 dailyTokens: 4096 A platform budget cannot determine whether two patches are equivalent or whether a task exceeded a business deadline. The RepairGuard and framework configuration add controls for: duplicate patch digests; excessive repair attempts; task deadlines; maximum MAF iterations and function calls. Token budgets constrain inference cost; repair guards and framework loop limits constrain business failure. 8. From OpenClaw to MAF: change the application, preserve the external boundary OpenClaw is effective for rapidly discovering the conversation, planning, tool, and specialist behavior the product needs. Production requires explicit state, typed tools, repeatable tests, and a human stop. Forge encodes the workflow as application code: class WorkflowState(StrEnum): RECEIVE_REQUIREMENT = "RECEIVE_REQUIREMENT" VALIDATE_SCOPE = "VALIDATE_SCOPE" INSPECT_REPOSITORY = "INSPECT_REPOSITORY" PROPOSE_PLAN = "PROPOSE_PLAN" APPLY_MINIMAL_PATCH = "APPLY_MINIMAL_PATCH" RUN_TARGETED_TESTS = "RUN_TARGETED_TESTS" SUMMARIZE_EVIDENCE = "SUMMARIZE_EVIDENCE" STOP_FOR_HUMAN_REVIEW = "STOP_FOR_HUMAN_REVIEW" There is deliberately no MERGE or DEPLOY state. In the final code/08 path, the kars MAF Python adapter pins the MAF client to the local router before MAF is imported: from kars_runtime_maf_python import bootstrap bootstrap() from agent_framework import Agent, tool from agent_framework.openai import OpenAIChatClient @tool(approval_mode="never_require") def inspect_release_contract(request_id: str, issue_id: str, revision: str) -> str: # Validate the pinned issue and revision, then return bounded evidence. ... maf_client = OpenAIChatClient(model=MODEL) maf_client.function_invocation_configuration["max_iterations"] = 3 maf_client.function_invocation_configuration["max_function_calls"] = 1 builder = Agent( client=maf_client, name="FabrikamReleaseBuilder", tools=[inspect_release_contract], default_options={"store": False}, ) The resulting path is: OpenClaw Intake -> MAF Agent -> inspect_release_contract @tool -> kars MAF Python adapter -> localhost Router -> GitHub Copilot or the selected provider MAF provides the agent, tool, session, middleware, and workflow programming model. kars provides the identity, network, budget, tool, audit, and lifecycle boundary outside the framework. This is one of kars' most important long-term advantages: the application framework can change without rebuilding the complete security shell. 9. Operational security: do not collapse every event into “the network was blocked” The updated code/06 requires incident response to distinguish: HTTPS; DNS; metadata service; local daemon; operator exec; break-glass access. Every denied channel is recorded in a hash-linked audit chain. Break-glass records without an incident ID are also rejected.\ Controller Conditions + Router Denials + Policy Budgets + Admission Decisions + Workload Recovery = One Investigable Sequence Audit integrity is not audit durability The router can verify the current audit hash chain. The tutorial also validates that pod replacement may reset an in-memory audit log, after which a new valid chain begins. Production deployments must continuously export audit data to an independently controlled durable backend. Otherwise, they can prove that the current chain is intact but cannot prove that pre-incident history survived pod loss. 10. AKS and multi-agent release: separation of authority, not two prompts chatting The AKS promotion separates Forge into Builder and Reviewer roles: Control Builder Reviewer Tokens per request 2048 512 Daily tokens 8192 2048 Tool authority Read, search, patch, test, diff Read diff/evidence and submit a decision Approval mode Never Always Trust threshold 700 800 The Builder may create a patch but cannot approve release. The Reviewer may approve a digest-pinned Builder artifact but cannot change source or approve an artifact created by the Reviewer. The final handoff carries patch, test-evidence, and artifact-manifest digests and still stops for human PR approval. The code/08 final sandbox-escape gate explicitly rejects: self-modified authority; symlink escape; host trust handoff; DNS egress; Builder self-approval; Reviewer source modification; untrusted or expired peer drafts. The release standard therefore becomes: Correct behavior must pass, and credential, network, tool, approval, evidence, and containment boundaries must remain intact. Passing tests alone is not release evidence. 11. What kars does not promise kars strengthens the sandbox, but it does not solve every risk automatically: It does not prove that a generated patch is correct. It does not make untrusted code safe to merge. It cannot protect a credential mistakenly mounted into the agent. Local Docker mode does not become a production boundary. It does not replace tenant RBAC, quotas, image policy, signing, supply-chain controls, or durable audit export. It cannot compensate for a policy that deliberately enables arbitrary shell and unrestricted egress. Confidential isolation does not replace least privilege, tool policy, egress policy, and code review. The sandbox bounds authority and blast radius. Tests, evaluation, independent review, and release policy still determine whether a change is acceptable. 12. An enterprise adoption path with measurable exits Phase 1: define the business and threat contract Specify inputs, outputs, allowed actions, forbidden actions, data boundaries, and the human approval point. Exit: product, platform, and security can all explain the agent's maximum authority. Phase 2: validate one OpenClaw vertical slice Use narrow business MCP tools instead of a general shell, and include hostile repository content. Exit: the normal task succeeds while self-configuration, path/symlink, trust-handoff, and egress tests fail. Phase 3: encode the sandbox as a Kubernetes contract Validate UID separation, root filesystem, capabilities, volumes, service-account tokens, NetworkPolicy, egress guard, and exec admission. Exit: the five boundaries are supported by runtime evidence, not only YAML review. Phase 4: add tool, model, and cost governance Apply McpServer, ToolPolicy, and InferencePolicy. Test unknown tools, dangerous arguments, and token overflow. Exit: violations are denied on the live request path. Phase 5: migrate into explicit MAF code Encode workflow state, typed tools, loop limits, evidence, failure paths, and the human stop. Exit: the MAF runtime preserves the external boundary already proven around the OpenClaw prototype. Phase 6: promote to AKS through GitOps Pin source revision, image digest, and loaded policy digest. Separate Builder and Reviewer authority. Prepare the kill switch, rollback, and durable audit export. Exit: one allowed workflow succeeds, multiple escape and authority-violation scenarios are denied, and all results have correlated evidence. Conclusion: kars does not make the model smarter; it makes agent authority explainable Enterprises will ultimately ask: What can the agent access? Where are the provider credentials? Who defines and changes the tool authority? Can prompt injection move data through DNS, metadata, a daemon, or HTTPS? How many tokens and repair iterations may one task consume? Who may patch, approve, merge, or deploy? Does evidence survive pod loss? If OpenClaw is replaced by MAF, does the security model remain intact? The ByteCraft AI story does not argue for one universal agent framework. It argues for a stable Agent Infrastructure layer: Use OpenClaw to discover valuable behavior quickly, use Microsoft Agent Framework to encode that behavior as explicit and testable application code, and use kars to remove credentials, networking, tools, budgets, sandboxing, audit, and lifecycle authority from the agent application itself. An agent becomes an enterprise workload when it has an independent identity boundary, a budget, a constrained tool surface, controlled egress, exportable evidence, and operational suspension and rollback—not merely when it runs inside a container. References Let's Learn Microsoft kars Microsoft kars