AI DOERS
Book a Call
← All insightsAI Excellence

Why Anthropic Previewing Claude Mythos Without Releasing It Matters

Anthropic previewed Claude Mythos but routed it through a security partnership instead of releasing it, because the same power that crushes coding benchmarks can also find and exploit vulnerabilities. The real lesson for a business is that the model, not the tooling, is what unlocks the next wave of useful AI.

Why Anthropic Previewing Claude Mythos Without Releasing It Matters
Illustration: AI DOERS Studio

A small law firm in a mid-sized city spent fourteen months and roughly $18,000 on AI tools before any of them produced a result worth keeping. The tools changed. The interfaces changed. The workflows the firm designed around them changed three times. What did not change was the outcome: the first draft was unreliable, the partner spent more time fixing AI output than the tool had saved in generation, and the associates stopped using the tools after the first month because the friction was higher than the benefit.

The firm had three lawyers and served primarily small-business clients on contracts, employment matters, and real estate transactions. Its workflows were well-defined. The documents it produced were largely templated at the structural level, with significant variation in the specifics. On paper, it was exactly the kind of operation that should have benefited immediately from AI assistance. In practice, it was not producing consistent AI results in the first year. This is the story of what changed.

Madhuranjan Kumar reviewed the firm's implementation in the second year, after the approach shifted. The failure in year one was not the tools, the interfaces, or the workflows. It was the model.

The first year of AI tools: what kept failing and why the firm almost gave up

The firm's first AI purchase was a general-purpose writing assistant built for legal professionals. The associate who evaluated it ran a trial on a standard commercial lease addendum. The first-draft quality was reasonable for the clauses the model had clearly seen many times: standard rent payment terms, basic maintenance responsibility language, common commercial use restrictions. For those clauses, the model produced language that required light editing.

The failure appeared the first time the associate asked the tool to handle something specific to the firm's actual client base. The firm worked frequently with clients who operated on property with unusual zoning classifications, where the interaction between commercial and agricultural use created edge cases that standard lease templates did not address. The model produced language for those situations that was internally consistent but legally wrong for the applicable jurisdiction. It sounded authoritative. It was wrong.

The associate caught the error because they knew the law. A less experienced reviewer would not have caught it. The partner's conclusion after the trial was that AI assistance saved time on the portions of documents where a careful reviewer was least necessary and created risk on the portions where a careful reviewer was most necessary. The time savings and the risk did not cancel out. The risk was larger.

The second tool the firm tried was a document review system marketed for due diligence. It scanned contracts and flagged clauses that deviated from standard commercial terms. The flagging was accurate for standard deviations. The problem was that about 30 percent of the firm's clients used custom agreements developed over years of negotiation with their counterparties. Those agreements had intentional deviations from standard terms that were protective, not problematic. The AI flagged the intentional deviations at the same rate as the actual problems. The partner was reviewing flagged clauses for 40 minutes per document to identify the 4 minutes of actual issues. The efficiency calculation was negative.

The third tool was a general large language model accessed through a chat interface. The associate began using it to draft initial responses to client questions. The outputs required substantial editing, not because the language was poor but because the model frequently lost track of the context of the matter partway through a longer response. If the client situation involved more than three or four interacting facts, the model would produce a response that was accurate about some of the facts and had quietly dropped or altered others. The associate had to read the output against the full case file every time to catch the omissions, which consumed the time the drafting had saved.

The firm nearly abandoned AI tools entirely after the first year. The partner's assessment was that the tools were helpful for tasks that did not require reliability and harmful for tasks that did. That left very little in a legal workflow where they could be trusted.

How it works (short)

The context capacity problem: why the same workflow produced different results with a stronger model

The decision to try a different approach came from an observation the associate made when comparing outputs from two different models on the same prompt. The prompt was a detailed memo summarizing a multi-party commercial dispute: five parties, four agreements, three years of prior dealings, and a question about which party's notice obligations had been triggered by a specific event.

The weaker model produced a memo that addressed the question but had dropped one of the five parties entirely from its analysis. The memo was internally consistent without that party. It was also wrong, because the dropped party's obligations under one of the four agreements directly affected the answer to the notice question.

The stronger model, accessed through a different interface, kept all five parties present through the full memo. Its analysis was still incomplete, but its incompleteness was visible: it flagged the question it could not resolve rather than silently dropping the complicating factor.

The difference was context capacity. The weaker model was losing track of the full situation as it moved through a longer document. The stronger model was holding more of the situation in view simultaneously and producing output that reflected that larger view. The task was the same. The prompt was the same. The workflow was the same. The model was different, and the output quality difference was large enough to change whether the result was usable at all.

This observation changed how the firm thought about AI tool selection. The question was no longer which interface was cleanest or which product had the most legal-specific features. The question was which underlying model had sufficient capacity to hold a full legal matter in view without losing pieces of it. That was a capability question about the model, not a feature question about the product.

Tasks a firm can safely hand AI as models improve

Loading the model with real firm context: templates, examples, standard positions

Once the firm switched to a frontier model with higher context capacity, the next change was in how it used the tool. The prior approach had been to describe each task from scratch with each prompt: here is the situation, here is what I need, produce the document. The new approach was to load the model with the firm's own context before giving it any task.

The firm assembled what the partner called a context packet: the firm's standard templates for its ten most common document types, annotated with notes about what each section was intended to accomplish and what variations the firm had used for different client situations. The packet also included several examples of documents the partner had personally drafted over the past three years that the partner considered representative of the firm's standard. And it included a summary of the firm's standard positions on the ten issues that came up most frequently in negotiation.

When the associate ran a drafting task with this context packet loaded, the output quality changed substantially. The model was no longer generalizing from training data about what legal documents look like. It was producing drafts that matched the firm's specific approach. The clause language matched the firm's style. The structure matched the firm's templates. The positions taken on common issues matched the positions the firm actually took.

The time the partner spent editing first drafts dropped from an average of 45 minutes per document to approximately 18 minutes per document for standard matters. For complex matters with unusual client situations, the editing time dropped less sharply, from around 90 minutes to about 55 minutes. The complex matters still required substantial partner review. But the base quality was high enough that the partner was refining a usable draft rather than rebuilding an inadequate one.

Six months in: what the firm handed AI and what it kept for itself

After six months of operation with the new approach, the firm developed a practical division of labor. The division was not based on document type or client type. It was based on a distinction the partner described as "settled versus unsettled."

Settled work was any task where the firm had a clear template, standard language, and prior examples that accurately represented the required output. Initial drafts of lease agreements for standard commercial clients. Demand letters for payment in clear-cut cases. Employment offer letters within the firm's standard terms. Engagement letters following the firm's template. For settled work, the AI with the context packet loaded produced a first draft that required editing, not rebuilding. The partner's review time stayed under 20 minutes per document.

Unsettled work was any task where the firm did not have a clear prior example, where the client's situation involved unusual facts, or where the applicable law had changed recently enough that the training data could not be trusted. The firm kept all unsettled work in the hands of the lawyers from the first word. AI did not touch the initial draft for unsettled matters. A human wrote from the context packet, using the templates as structure but departing from them wherever the client's situation required it.

The associates' usage increased substantially compared to the first year. The key change was that the division of labor was clear. The associates knew which tasks they could use the tool for with confidence and which tasks required them to draft independently. In the first year, the ambiguity about when to trust the output had caused consistent friction. In the second year, the settled-versus-unsettled framework removed that ambiguity.

The firm also stopped using AI for document review and concentrated its AI use entirely on drafting. The pattern of false positives in document review had not improved with the stronger model. The associates' time reviewing AI-flagged clauses was still disproportionate to the actual issues found. The firm returned that function to manual review by the responsible associate, which was slower but more reliable.

The number that changed the conversation about AI billing at the firm

Fourteen months into the new approach, the partner ran a calculation that changed how the firm thought about its AI investment. The number was not a cost comparison. It was a capacity number.

Before AI assistance, the firm could handle approximately 85 active matters simultaneously across its three lawyers. That limit was driven by the time required for drafting: the settled work that consumed consistent hours each week without requiring the lawyer's full legal judgment. Standard leases, standard employment matters, standard demand letters. The drafting was not intellectually demanding, but it consumed time that could otherwise have gone to higher-complexity work or client development.

With AI handling the settled drafting, the firm's active matter capacity increased to approximately 110 matters without adding staff. The additional 25 matters per quarter represented roughly $185,000 in additional annual billings at the firm's average matter value. The AI tool costs, including the frontier model subscription, were approximately $4,200 per year. The cost-to-benefit ratio for capacity expansion was not close.

The partner shared this number with Madhuranjan Kumar during a review of the second year. The framing was direct: the question the partner had been asking in year one was whether AI could improve the quality of individual documents. That question had not produced a useful answer because the quality improvement was incremental and hard to measure. The question the partner had started asking in year two was whether AI could increase how many matters the firm could handle with the same three people. That question had a clear numerical answer, and the answer was yes by approximately 29 percent.

The number also changed how the firm described AI to other small firms at professional association meetings. The pitch had shifted from "AI improves your drafts" to "AI expands your capacity without headcount." The second framing was the one that generated serious follow-up questions from other firm partners, because capacity constraints at small firms are universal and headcount decisions are expensive. A technology that addresses a universal capacity problem at a known cost is a different kind of business case than a technology that might improve quality in ways that are difficult to quantify.

The firm's plan for the next year was to expand the context packet to cover an additional fifteen document types, pushing more of the firm's work into the settled category. The unsettled work would stay human-authored. The context packet would grow. The capacity ceiling would rise.

Do it with an expert
You can build this yourself, or have it set up right the first time.

That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.

Book your call →
Madhuranjan Kumar

Madhuranjan Kumar

Founder, AI DOERS · Performance Marketing

Madhuranjan Kumar brings 20 years of performance-marketing experience and has managed over $200 million in Facebook ad spend for brands across the United States and beyond. His expertise spans the full modern marketing stack: Meta, Google Ads, TikTok, email automation, CRM, and the websites that hold it together. At AI DOERS he turns that track record into lead-generation systems for businesses across every industry.

← Back to all insights
Why Anthropic Previewing Claude Mythos Without Releasing It Matters | AI Doers