AI DOERS
Book a Call
← All insightsAI Excellence

Kimi K2 Thinking: A Free Open Model That Runs Hundreds of Steps On Its Own

A new open-weights AI matches top closed models on hard tasks and can chain hundreds of tool calls without a human. Here is what that means and how I would put it to work in a service business.

Kimi K2 Thinking: A Free Open Model That Runs Hundreds of Steps On Its Own
Illustration: AI DOERS Studio

A model anyone can download for free just matched some of the best closed AI systems on the hardest public tests, and it can run two to three hundred tool calls in a row without a human touching it. That is the news, and it is a bigger deal than the usual model-of-the-week churn. Kimi K2 Thinking, an open, open-weights model from a Chinese lab, is not a slightly cheaper alternative to the frontier. On several brutal benchmarks it is standing on the frontier. I am Madhuranjan Kumar, and I want to unpack what actually shipped, why it changes the math for ordinary businesses, and the concrete move I would make with it.

What shipped: open weights that reason for hundreds of steps

The headline capability is stamina. Kimi K2 Thinking is a reasoning model, which means it thinks step by step and uses tools as part of that thinking rather than as a bolt-on afterthought. Give it a goal and it does not just answer. It plans, runs a web search, reads what it finds, reasons about it, runs another search informed by what it just learned, and keeps going. Crucially, it can chain two hundred to three hundred tool calls without a person stepping in to nudge it along.

That long-horizon endurance is the line between a chatbot that answers a question and an agent that finishes a job. Most models lose the thread after a handful of steps. A model that can stay coherent across hundreds of steps can work through a genuinely messy, multi-part task from start to finish. On a hard web-research benchmark it scored well above the human baseline, precisely because that benchmark rewards the search, read, reason, search-again loop it was built for.

How it works (short)

The open-versus-closed gap is closing faster than the story admits

For most of the last few years, the strongest models sat behind a handful of United States labs, and the assumption baked into every business plan was that frontier capability meant paying one of them. Kimi K2 Thinking undercuts that assumption. It is fully open and open-weights, yet it beats leading closed models on some of the toughest public evaluations, including one of the field's hardest exams.

This echoes an earlier moment when another open model shocked the industry, and the pattern is now clear rather than a fluke. Capable AI is no longer something only a few companies own. When the gap between free-and-open and expensive-and-closed shrinks to almost nothing on the hardest tasks, the pricing power of the closed labs erodes, and the range of what a normal business can afford to attempt widens. That is the real story underneath the benchmark numbers.

Hours to prepare a client report (illustrative)

Cheaper to train means cheaper to use for everyone

There is a cost angle that deserves its own spotlight. The base model reportedly trained for only a few million dollars, a fraction of what frontier training used to cost. That matters downstream, because cheaper to train tends to mean cheaper to run and cheaper to access. Every time the cost of a capable model collapses, the floor drops for the businesses that were priced out of the previous generation.

The demos back up that the capability is real, not a benchmark artifact. From a single prompt the model built a working document editor with real formatting and save, an interactive simulation with adjustable controls, and clean data visualizations. In one of the more striking runs it took a research request, pulled public datasets on its own, computed results, ranked the findings, and produced an interactive report complete with a map and charts after just one round of human feedback. That is real analytical work, not a mockup.

Open weights quietly solve the privacy problem

Here is the part of the news that most coverage skips, and it is the part I care about most for real businesses. Because the weights are open, a company can run a model like this on its own infrastructure. Sensitive data never has to leave its control.

For a lot of businesses, that single fact is the difference between AI being usable and being off-limits. Regulated and privacy-sensitive operations often cannot send client data to an outside service, and that rule has kept them on the sidelines while everyone else experimented. An open, frontier-grade model that runs privately removes the objection. You get deep multi-step reasoning and you keep the data in house. That combination is rare, and until recently it did not exist at this level of capability.

Why tools inside the reasoning changes the ceiling

It is worth slowing down on one technical point, because it explains why this feels different from the assistants people are used to. In most earlier systems, a model would produce an answer and then, separately, some outer software would call a tool and hand the result back. The reasoning and the tool use were two different layers stitched together, and the seams showed. The model could not really decide mid-thought that it needed to check one more source and then adjust its plan based on what it found.

Kimi K2 Thinking folds the tools into the thinking itself. A search is not an interruption of the reasoning, it is a move within it, the way a skilled researcher pauses to look something up and then keeps building the argument. That integration is what makes the long chains coherent rather than chaotic. When a model can search, read, reason, and search again all inside one continuous train of thought for hundreds of steps, it can pursue a question the way a person does, following leads, discarding dead ends, and refining as it learns. The result is not just more steps, it is steps that build on each other, which is exactly why the model scored above the human baseline on hard web research. The ceiling on what an agent can accomplish in one run rises sharply when the tools stop being a bolt-on and become part of how it reasons.

The competitive pressure this puts on pricing

Step back from the single model and look at what its arrival does to the market, because that is where the durable effect lives. For years the pricing of capable AI was set by a small group of labs with little to undercut them. A free, open, frontier-grade model changes the negotiation for everyone, even businesses that never touch it. When a genuinely competitive option exists that costs a few million dollars to train and can be run privately, the closed labs lose some of their power to charge whatever they like for the top tier.

That pressure flows downhill to ordinary buyers. It shows up as falling prices on the paid services, as more generous free tiers, and as a wider set of viable options for a task that used to have one expensive answer. The practical lesson for an owner is not to chase every new release, but to recognize that the cost of the AI powering your operation is on a downward trend that these open models accelerate. Budgeting as though today's prices are permanent would be a mistake, and so would locking into a single vendor as if no alternative existed.

Who this changes things for

Any business that does research-heavy or document-heavy work should be paying attention. The strongest fits are the ones where a person spends hours gathering scattered information, cross-checking it, and turning it into a clear summary or recommendation. That describes law firms, accounting practices, real estate research teams, consultants, and insurance agencies. The agentic browsing strength lands squarely on that work, because so much of it is finding the right facts across many sources and reasoning over them.

The private-deployment angle adds a second group: any business that could use frontier AI but is blocked by data rules. Between the two, a large slice of the professional services economy just got access to something it either could not afford or could not legally use a year ago. The concrete move is not to rip out your current tools tomorrow. It is to identify the one research-heavy task where hours disappear and test whether an agent that can chain hundreds of steps can do the first draft.

A worked example: the insurance agency back office

Picture an independent insurance agency where the owner and two agents lose their best hours to reading rather than talking to clients. This is the illustrative shape of the payoff.

The first target is quote preparation and policy comparison. When a client asks which plan fits them, the agent reads the client's stated needs, pulls the relevant policy details, compares coverage and exclusions across carriers, and drafts a plain-language summary that a human reviews before it goes out. Work that used to eat an afternoon of reading fine print becomes a draft you check and refine. Because the model can chain hundreds of steps, it can work through a stack of dense policy documents without constant babysitting.

The second target is claims intake and case research. The agent reads a new claim, gathers the supporting documents, flags missing information, and assembles a tidy file with a summary of what is present and what is still needed. The third is renewals, where each month it reviews the book, flags policies coming due, notes coverage gaps, and drafts personalized outreach for a human to approve.

Say preparing a full client report took four hours before. With this in place it might drop to two hours by week four and one hour by week twelve as the instructions get tuned to the agency's carriers and process. Those numbers are illustrative, but the mechanism is not: the model removes the reading and assembling, and the agent keeps the judgment on risk and relationships. And because it is open-weights, the agency can run it privately so client health and financial details never leave its own systems, which in a regulated business is not a nice-to-have but a requirement.

The move to make now, and where to be careful

Start with one research-heavy task and prove it before scaling anything. For an agency the cleanest first target is policy comparison or claims intake, because the value is obvious and easy to measure. Run it on real cases for a week, keep a human reviewing every single output, and only expand once you trust the quality. If data privacy is a live concern, look seriously at running the open weights on your own infrastructure rather than a public endpoint, since that is one of the concrete advantages this release hands you.

The skill that makes any of this work is writing the instructions so the agent behaves the way your business actually does, and setting clear checkpoints where a person reviews before anything reaches a client. That setup work, plus the judgment about which tasks are safe to automate, is what separates a useful system from a risky one. It is tempting to point a powerful new model at everything at once, but the agencies that get real value do the opposite. They pick one task where a mistake is cheap and easy to catch, they keep a human in the loop on every output for the first few weeks, and they expand only after the quality has earned their trust. The long-horizon stamina that makes this model impressive is also what makes an unreviewed mistake compound, so the review step is not bureaucracy, it is the thing that lets you safely hand the model longer and longer jobs over time. The freed-up hours have a way of compounding across the whole operation, because the same clean client summaries feed your CRM and website stack and the sharper positioning that comes out of better research strengthens both SEO and organic search and the messaging behind your Facebook and Instagram ad campaigns.

You can take the do-it-yourself path above and start proving value this week. If you would rather have the whole thing built, tuned to your carriers and your process, and handed over working with the review checkpoints already in place, that is the kind of setup I do for clients, and you can bring in someone who has already wired this up many times.

Do it with an expert
You can build this yourself, or have it set up right the first time.

That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.

Book your call →
Madhuranjan Kumar

Madhuranjan Kumar

Founder, AI DOERS · Performance Marketing

Madhuranjan Kumar brings 20 years of performance-marketing experience and has managed over $200 million in Facebook ad spend for brands across the United States and beyond. His expertise spans the full modern marketing stack: Meta, Google Ads, TikTok, email automation, CRM, and the websites that hold it together. At AI DOERS he turns that track record into lead-generation systems for businesses across every industry.

← Back to all insights
Kimi K2 Thinking: A Free Open Model That Runs Hundreds of Steps On Its Own | AI Doers