Microsoft's Real Plan: Build Its Own Frontier Models and Stop Renting Intelligence
Microsoft used Build to reveal seven in-house models and an OpenClaw-powered Windows agent called Scout, because its stated goal is true self-sufficiency in AI. Here is what that race toward owning your own intelligence means for a normal business.

Microsoft just told the whole market its real plan, and it has nothing to do with a shinier chatbot. At its Build event the company shipped seven of its own in-house AI models plus an OpenClaw-powered Windows agent called Scout, and its stated reason was blunt: it wants true self-sufficiency in AI instead of renting intelligence from partners like OpenAI. I am Madhuranjan Kumar, and I read that headline less as gossip about big labs and more as a set of instructions for the rest of us. The same logic that pushes Microsoft to own its models is the logic that lets a small business stop overpaying for AI it could run itself. This is a playbook for doing exactly that, one step at a time, using the news as a map.
Sort every AI task into routine and hard
Before you touch a single model, spend an hour listing every place AI already touches your work, then split that list into two piles. The first pile is routine, high-volume, low-stakes work: rewriting supplier copy into your brand voice, drafting answers to common questions, summarizing notes, tagging photos, cleaning up spreadsheets. The second pile is hard, occasional, high-stakes work: analyzing a quarter of sales data to decide what to promote, building a small internal tool, working through a pricing decision that actually moves money.
This sort is the whole game, and it is why Microsoft's move matters to you. The company is not replacing OpenAI for everything. It built a broad lineup so it can route each job to the cheapest model that clears the bar. Its new flagship thinking model is the reasoning workhorse, which Microsoft says is preferred over Sonnet 4.6, though notably it did not compare it against the true state of the art like Opus 4.8. That gap is the tell: even a trillion-dollar company keeps a premium model on hand for the hardest thinking and uses its own cheaper models for everything else. You should copy that discipline exactly. Once your two piles are written down, the rest of this playbook is just deciding which model runs which pile.

Move your highest-volume routine work to a small model
Take the single most repetitive item in your routine pile and move it to a small local or open model this month. Do not try to move everything at once. One task, run to completion, teaches you more than a grand plan.
The tools to do this arrived the same week Microsoft made its announcement, which is not a coincidence. At Computex, NVIDIA showed the RTX Spark, a combined GPU and CPU with up to 128GB of unified compute that can run large, capable models directly on your own machine. On the open side, NVIDIA's Nemotron 3 Ultra ships with 550 billion open weights, Google's Gemma is small enough to run on a laptop, and MiniMax M3 claims to beat GPT 5.5 on SWE-bench Pro with a one million token context. The argument for running these locally is practical, not ideological. Prompts never leave the device, so customer data stays private. A strong model works with no internet. And the per-token cost drops to zero, because you are paying for hardware you already own instead of metering every word. For routine work that you run hundreds of times a day, that difference compounds fast.

Point the strong media models at the work that actually sells
Not every job belongs on a tiny model. Microsoft's standouts this cycle were in media, and they map neatly onto real business work. MAI Transcribe 1.5 is positioned as the world's most accurate transcription model and roughly five times faster than competitors, MAI Voice 2 generates speech across fifteen languages, and MAI Image 2.5 ranked near the top of text to image and number two in image editing behind only GPT Image 2.
The step here is to identify the media work that touches revenue and hand it to these stronger models on purpose. Turning supplier calls and customer service recordings into searchable text is a transcription job. Re-staging and cleaning product photos at scale is an image-editing job. Producing narration for a how-to video in several languages is a speech job. This is also where your paid channels quietly benefit, because sharper product imagery and cleaner ad copy lower the cost per lead on Facebook and Instagram ad campaigns without you spending an extra rupee on media. The point of this step is discrimination: strong models for the media that sells, tiny models for the text that merely needs to exist.
Reserve frontier models for the rare, high-stakes jobs
Now protect your frontier subscription. The mistake I see most often is a business paying frontier prices for work a small model handles for free, then wondering why the AI bill keeps climbing. Reserve the expensive model for the genuinely hard, occasional jobs from your second pile: analyzing which products to promote next quarter, one-shotting a small internal tool that flags low-stock bestsellers before they sell out, or untangling a decision where several variables interact.
Microsoft is teaching this lesson at its own scale. It is not renting frontier intelligence for routine work, and it kept the option to reach for the best available model when a problem truly demands it. Do the same. Keep one frontier plan, use it deliberately, and let the cheap and local models carry the volume. If you want a second reference point on where these routine-versus-frontier lines fall, the GitHub Copilot app that shipped this cycle lets you choose any model from any provider in one place, so once signups reopen you can compare outputs side by side and confirm your routing choices with your own eyes.
Pilot one always-on agent on a single workflow
The last step is the one most owners rush and regret. Microsoft's Scout is an always-on personal agent in a new category it calls autopilots, and under the hood it is literally OpenClaw's open-source technology wired into Windows, Teams, Outlook, OneDrive, and SharePoint at the operating-system level. That kind of always-on agent is powerful precisely because it can act across your files and messages without being asked twice, which is also exactly why you should not turn it loose on your whole business on day one.
Pilot it on a single recurring workflow. Pick one narrow, forgiving job, drafting reorder reminders, or surfacing the day's most urgent customer issues, and let the agent run only that until you trust it. A managed Windows version lowers the barrier a great deal for anyone nervous about setting up an autonomous agent, but the discipline is yours to keep: prove it on one workflow, watch it for a few weeks, then widen its remit. The leads it surfaces should land in your CRM and website stack where your existing follow-up automation handles the next touches, so the agent adds a lane rather than replacing the plumbing you already trust.
Worked example: an e-commerce store's first month
Here is how the whole playbook lands for a mid-size online store, with illustrative numbers to show the shape.
The store starts by sorting its AI work and finds the routine pile is enormous: roughly 400 product descriptions a month, a steady stream of supplier copy to rewrite in the brand voice, and dozens of repeat customer questions a day. Before the change, all of that ran through a frontier model at a metered cost of about 200 dollars a month in tokens. In week one, the owner moves product descriptions and supplier rewrites to a laptop-sized open model. Those hundreds of jobs now cost nothing per token, and because they run on-device, order details never leave the building. Token spend on routine work falls toward zero.
In week two, the store points a top transcription model at its customer service recordings and a strong image-editing model at its product photography, cleaning up and re-staging 60 listings that had weak images. Sharper images and tighter copy feed straight into the store's paid channels, and the effect shows up first in the cost per lead on Facebook and Instagram ad campaigns, where a better creative can pull the number down noticeably. That same refreshed product copy also feeds SEO and organic search with no extra work, since the descriptions were rewritten once and now serve both jobs.
By week four, the store keeps exactly one frontier subscription, reserved for two things: a monthly analysis of sales data to decide which 20 products to push, and the occasional small internal tool. An always-on agent, piloted only on reorder reminders, now drafts restock alerts for the owner to approve. The store sells exactly as it did before. The difference is that the routine writing and support run cheaply and privately on local models, the frontier spend goes only to work that moves revenue, and the monthly AI bill stops being a mystery. On a chart of tasks run locally versus in the cloud, the store moves from about 5 percent local before, to 30 percent by month three, to more than half by month six, and its costs fall the whole way down.
The mistakes that quietly waste this shift
Because this is a strategy rather than a single product, the ways to get it wrong are strategic too, and they cost real money. The first mistake is treating local and frontier models as an either-or choice. Owners hear that a laptop-sized model runs for free and try to force every task onto it, then get frustrated when it flubs a hard analysis. The whole point of Microsoft's approach is a portfolio, not a swap. You keep the frontier model for the jobs that need it and let the cheap models carry the volume. Purity in either direction is a trap.
The second mistake is skipping the sort step and jumping straight to buying hardware or an agent. An RTX Spark or a slick autopilot is exciting, but if you have not written down which tasks are routine and which are hard, you have no idea what to run on it. The hardware follows the sort, never the other way around. Buy or download nothing until your two piles exist on paper.
The third mistake is ignoring privacy as a feature you are paying for. Running a model locally is not only about saving on tokens. It means customer records, order details, and internal notes never leave the building, which for a lot of businesses is worth more than the cost savings. If you route sensitive work to a cloud model out of habit when a local model would have kept it private, you have given up an advantage for no reason. Treat on-device processing as a deliberate choice for anything you would not want to send to a stranger.
The fourth mistake is measuring nothing. This shift only proves itself if you watch the share of work moving to cheap models and the frontier bill shrinking alongside it. Put a simple number on both, revisit it monthly, and let the trend tell you whether your routing is working. Without a measurement, you are guessing, and guessing is how a good strategy slowly drifts back into overpaying by default.
Where the real leverage shows up
The reason this playbook works is that it copies a strategy a giant company just validated in public. Microsoft did not build seven models because owning intelligence is a vanity project. It did it because renting all of your intelligence, forever, at frontier prices, is a bad trade once you have the volume to justify running your own. You will never build seven models, and you do not need to. You need to run the cheap ones for the routine pile, the strong media models for the work that sells, and one frontier plan for the handful of jobs that genuinely deserve it.
Intelligence is leaving the chat window and moving into tools, devices, and machines you control, and the businesses that sort their work early will spend a fraction of what their competitors spend for the same output. You can absolutely run this sort yourself and move your first routine task to a local model this month. If you would rather have someone map exactly which of your tasks belong on cheap local models versus frontier ones, wire up the media models and the agent, and make it pay off from the first week, that is the kind of build I do for clients, and you can bring me in to handle it.
That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.
Book your call →
