Anthropic Says AI Is Starting to Build Itself. Here Is What the Numbers Actually Mean for Your Business
Anthropic's new paper argues AI is increasingly building AI, with task length doubling every four months and most of its own code now written by Claude. Strip out the fear and the real signal is simple: execution is getting cheap, so your judgment is the asset to protect.

In May 2026, Anthropic published a paper stating that more than 80 percent of the code merged into its codebase was written by Claude, up from low single digits just a year earlier. The same paper documented that Claude moved from completing tasks that took a person about four minutes in March 2024 to handling twelve-hour autonomous tasks by 2026, with some Claude Code sessions running forty consecutive hours. An internal research model achieved a 52-times speedup on optimization benchmarks compared to three times a year earlier. I am Madhuranjan Kumar. When I read that paper, I was not primarily thinking about AI development. I was thinking about a specific e-commerce business whose story I want to tell here, because it shows what those numbers mean when translated from an AI lab's research into a business owner's actual week.
The business sells home goods online. Around seventy SKUs, a small fulfillment team, one marketing person, and an owner who handled everything that did not fit clearly into someone else's job. Revenue was not the problem. Margin was reasonable and growing. The problem was capacity. The owner was spending most of each day on execution tasks: writing or revising product descriptions, building the weekly promotion emails, compiling sales data into a summary for the team, updating FAQ content, checking competitor prices, reviewing ad performance and drafting new creative briefs, and processing returns data into a weekly loss report. These were all tasks with clear inputs and definable outputs. None of them required the owner's specific judgment to do passably well. But all of them required someone's time, and that someone was the owner.
The state of the operation before: an owner running on constant execution
The business had grown to a point where the owner recognized the ceiling but not the cause. Revenue kept climbing but it was costing roughly 22 to 25 hours per week of the owner's time in execution tasks. These were tasks the owner was handling either because no one else had capacity or because the tasks were ambiguous enough that outsourcing them felt risky. Product descriptions needed to match the brand voice. Promotional emails needed to reflect the current inventory state. The competitor price check needed consistent structure to be useful. None of these required the owner, but the owner was doing them anyway because the alternative was inconsistency.
The team was capped at three people and adding a fourth did not make economic sense at the current margin. The owner had looked at hiring a virtual assistant and found the onboarding overhead, coordination time, and quality variation added a different kind of execution load. The business was stuck at a capacity ceiling that was not obviously breakable through conventional hiring. There were only so many hours in a day, and too many of them were going to work that had clear rules and consistent structure but still required someone to sit down and do it. The business was profitable and growing, but the owner was running out of week before running out of work. The owner was spending the majority of their most productive hours on work that a capable but rule-following executor could handle, and there was no obvious way to change that without either raising costs significantly or tolerating lower quality.

Reading the paper and reorganizing around the execution-versus-judgment split
The Anthropic paper arrived during this period. The owner read it carefully and came away with one clear insight: the paper was describing a structural shift in what is scarce. Execution, the paper argued, was becoming cheap and fast to automate. Task length agents can handle is doubling every four months. Engineers are writing eight times more code with AI tools, though the paper's own data noted the useful output ratio is closer to four times, because AI-written code is lower quality per line than human-written code and requires a review layer to maintain quality. Research taste and strategic judgment remained firmly in the human domain.
Most of the owner's 22-to-25-hour execution load was execution in precisely the sense the paper described: rule-based, definable, repeatable work that had a clear structure and could be expressed precisely enough for a capable executor to follow without back-and-forth. The decision the owner made was not to subscribe to a tool or hire someone. It was an organizational decision: reorganize the business so that execution tasks go to agents by default and the owner's time is reserved for judgment work that agents cannot yet do reliably. The paper's framing provided a clear sorting criterion: if a task can be fully described in a written brief with enough specificity that a capable assistant could follow it without clarification, it is an execution task. If a task requires weighing competing considerations against unstated business context, it is a judgment task. The sorting criterion sounds abstract until you apply it to an actual list of tasks. Then it becomes obvious quickly. Writing a product description from a template with defined fields is execution. Deciding whether to stock a new category of product is judgment. Drafting a follow-up email from a sequence template is execution. Deciding what the promotional calendar should look like this quarter is judgment. Compiling the weekly sales summary from spreadsheet data is execution. Interpreting why the conversion rate dropped and what to do about it is judgment. The owner found that sorting 54 tasks into those two columns took about 90 minutes and was the most clarifying afternoon they had spent on the business in years.

Mapping 54 recurring tasks into two columns
The owner spent one afternoon listing every recurring task in the business and sorting each one into two columns: execution and judgment. Fifty-four tasks went on the list. Thirty-seven landed in the execution column. Seventeen landed in the judgment column.
The execution column included: product description drafts for new SKUs, weekly promotional email drafts, competitor price monitoring and weekly summary, returns data compilation, ad performance data compilation, seasonal landing page content updates, FAQ content revisions, standard customer service reply drafts for the twelve most common inquiry types, and supplier follow-up email drafts. These tasks had clear inputs, clear success criteria, and consistent enough structure that a well-briefed executor could produce a usable first draft without understanding the business deeply.
The judgment column included: product sourcing decisions, pricing strategy, major customer disputes, supplier relationship management, ad strategy and budget allocation for Google Ads campaigns, inventory forecasting for the holiday period, and hiring decisions. The split was revealing. The owner was spending the majority of their time on the execution column. The judgment column, which represented the highest-value work only the owner could do, was getting crowded into evenings and early mornings when everything else was quiet. The mapping exercise made clear what the Anthropic paper implied: the execution column should not be owned by the owner.
The first 90 days: agents absorb the execution column
Over the first 90 days, the owner deployed agents on the execution column in order of time savings potential. Product descriptions came first: an agent given a template with the brand voice guidelines, the SKU data fields, and three example descriptions produced first-draft descriptions for new products that needed about 15 minutes of review and editing rather than 45 minutes of writing from scratch. Weekly promotional emails came second: an agent given the current inventory highlights, the discount parameters, and the email structure template produced a draft that needed 20 minutes of review versus 75 minutes of drafting. Competitor price monitoring came third: an agent running on a daily schedule compiled a formatted comparison that landed in the owner's inbox each morning, replacing 30 minutes of manual checking.
The Anthropic paper's nuance about 8-times more code but only 4-times more useful output showed up immediately in practice. The agents produced more volume but required a review layer to maintain quality. Product descriptions sometimes needed two rounds of editing rather than one. Email drafts occasionally missed a promotional nuance that required correction. But the review time was dramatically shorter than the drafting time had been, and the editing was more focused because it started from a structured draft rather than a blank page. Skipping the review step produced errors that reached customers.
The discipline of keeping the review step was the difference between a system that improved over time and one that would have broken trust with the team. The first week of running agents on product descriptions without review produced three descriptions that had a factual error about material specifications. Reviewing and catching those before they went live turned a potential problem into a data point: the agent needed the material specifications spelled out explicitly in the brief, not implied from the category name. Adding that field to the template fixed the error pattern. The review step is where the operational knowledge about how to brief agents well gets built, and that knowledge is what makes the system reliable rather than merely fast.
By the end of 90 days, twelve execution tasks were running on agents. The owner spent about four hours per week on agent output review compared to roughly 22 hours per week on the same tasks before. The net time reclaimed was 18 hours per week.
Six months in: the owner is working on different problems
At six months, the owner's description of their week had changed substantially. The 18 hours per week formerly spent on execution tasks was now going into work that had previously been crowded out entirely. Sourcing three new product categories from suppliers they had identified but never had time to pursue. Building relationships with two new suppliers who offered better margin on existing categories. Running properly structured tests on Facebook and Instagram ad campaigns rather than approving whatever the marketing person put together between other tasks. Planning the holiday inventory position with enough lead time to optimize it rather than reacting to what was available in October.
The business added eight new SKUs from the improved supplier relationships. The holiday inventory position was the strongest the business had run. The ad campaigns receiving the owner's direct attention on strategy were outperforming the prior period by a margin the marketing person attributed entirely to having the owner's judgment in the loop rather than at a distance.
The owner's framing of what changed was simple: the execution tasks did not go away. They still happened every week. But the person doing the initial heavy pass on them was now an agent rather than the owner. The owner's time went to finishing and to judgment, which was exactly the reorganization the Anthropic paper implied was available. The paper described AI building AI. The practical implication for this business was AI doing the execution so the owner could build the business. The Anthropic paper's thesis was not primarily about AI labs. It was a description of a structural shift that applies to every business with a meaningful execution load. The owners who read it that way and act on it will find 18 hours per week. The ones who read it as a technology story and wait for it to become relevant will find those 18 hours going to the same execution tasks they are doing today. The practical implication of the paper is not that every business needs to think about AI research or model development. The Anthropic paper's numbers, 80 percent Claude-written code, 52x research speedups, task length doubling every four months, are not primarily a story about what AI labs are doing. They are measurements of how fast the execution layer of any knowledge work is changing. Reading them as a business owner rather than as a technology watcher means asking a single question: is my business organized to take advantage of the fact that execution is becoming cheap, or am I still treating execution as the thing that deserves most of my time? It is that every business with a significant execution load should ask the same question this e-commerce store asked: which of the tasks I own every week are execution tasks that follow clear rules, and which ones are judgment tasks that only I can do well? The answer to that question, applied honestly to a real list of recurring tasks, is what turns an Anthropic research paper into a change in how the business runs.
That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.
Book your call →
