AI DOERS
Book a Call
← All insightsFuture of Marketing

AI Just Showed Its Work on a Hard Problem, and Why That Matters for Your Business

A new AI model solved an unsolved math problem in minutes and, crucially, produced a formula humans can check. The grounded lesson for any owner: you can now trust AI on hard problems, as long as you make it show its work.

AI Just Showed Its Work on a Hard Problem, and Why That Matters for Your Business
Illustration: AI DOERS Studio

The difference between a result you can use and a result you have to trust

I am Madhuranjan Kumar, and I want to begin with a specific distinction that this article is built around, because it changes how you should think about AI outputs on every task that carries real consequences. There are two kinds of AI result. The first is a result you have to trust: the model returns an answer, gives no indication of how it arrived there, and you either accept it or reject it based on gut feeling about whether it sounds right. The second is a result you can use: the model returns an answer along with the reasoning that produced it, and you can verify the reasoning before you act on the answer.

The recent news about AI solving a genuinely unsolved mathematics problem is interesting for the headline reason: an AI found a new result that human researchers had not reached. But the reason it matters for any business owner is a different one. The AI did not just announce the answer. It produced an explicit formula. Researchers could open the formula, examine it, check whether each step was valid, and confirm the result independently. They did not need to trust the AI. They could verify it.

That is the distinction I keep returning to in client work. An answer you can verify is an asset. An answer you have to trust is a liability that has not yet come due. For a quick question that does not matter much, the difference is trivial. For a pricing decision, a scheduling problem, a cost estimate, or a job plan that involves money and commitments, the difference between a verifiable answer and a trusted-on-faith answer is the difference between confidence and hope.

How it works (short)

What "showing work" means when the AI is not a student

The phrase "showing work" comes from mathematics education, where it means writing out the intermediate steps so a teacher can see not just what you got but how you got there. The purpose for a student is to prove that the student understood the process rather than guessing the answer. The purpose for a business owner asking an AI to show its work is completely different.

When you ask an AI to show every assumption and every step in a quantitative problem, you are not testing whether the AI understands the process. You are giving yourself a diagnostic tool. The ability to read the reasoning is what lets you catch the specific step where an assumption was wrong for your situation before you act on the result. An assumption the model made that is generic but correct for most businesses may be wrong for yours. The model used a standard labor hour figure for a type of job but your team consistently runs longer on that job type. The model used current material prices from training data that are three months old. The model applied a waste factor appropriate for flat roofs to a calculation for a complex pitched roof.

None of those errors are visible in the final number. They are only visible in the assumptions section of the reasoning. If the model provides that section and you read it, you catch the error before it becomes a bad quote. If the model does not provide that section, or if you do not read it because the final number looks reasonable, the error propagates into a real commitment.

This is the practical mechanism behind what researchers observed when an AI solved the mathematics problem: the model did not just find a sharp answer at the true boundary of the problem rather than a conservative estimate. It wrote out how it got there. That combination of precision and transparency is what makes an AI result something you can actually build decisions on rather than something you test and hope about.

Pricing scenarios tested per hour

Why the sharp answer matters more than the cautious one

There is a specific pattern in how human experts handle hard quantitative problems under uncertainty: they draw a cautious boundary rather than a precise one. A structural engineer uses a safety factor that is higher than the strict physical calculation requires. A financial estimator adds a contingency buffer. A job cost estimator pads the figure to protect margin. These conventions exist because uncertainty is real and the cost of being wrong is high.

The AI that solved the mathematics problem did something different. Where the human researchers had placed a cautious boundary with a safety margin, the model identified the exact true boundary and drew the line there. It did not need the buffer because it had actually worked through the logic to the precise answer rather than approaching it cautiously.

For a business owner, this property of AI reasoning, the ability to reach for the precise answer rather than the conservative approximation, changes the economics of estimation. When you estimate with a buffer, the buffer protects margin but also costs you competitive bids. If your cost estimate is ten percent higher than the true cost to protect yourself from error, you will lose competitive bids that you could have won with a correct estimate. If your schedule is padded for contingency, you quote longer timelines than competitors who quote more precisely.

An AI that can reason precisely through a cost or schedule problem and show the reasoning so you can verify it gives you the option of submitting sharper estimates where the margin is genuine rather than padding-inflated. You still apply appropriate judgment about contingency for factors the model cannot know. But the baseline calculation can be more precise, which means the contingency you add is real contingency for real uncertainty rather than defensive rounding that makes you less competitive.

The roofing estimator case is the clearest illustration. A roofer who estimates square footage, pitch factor, waste factor, and material cost by eyeballing the job and applying rules of thumb that have been refined over years is producing an estimate that is approximately right on average. Over a hundred jobs, the errors average out. On any specific job, the error could go either way, and on a large job with expensive materials, a few percentage points of error in the wrong direction is a meaningful dollar amount.

An AI that works through those same variables, shows the calculation for each, and produces a verifiable estimate is not replacing the roofer's judgment. It is giving the roofer's judgment a more precise and checkable input. The roofer reads the waste factor the model used, confirms or corrects it based on the specific roof geometry, and gets a sharper starting number for the estimate. The final number includes the roofer's experience with that specific property and those specific materials. The AI contributes precision and transparency. The roofer contributes contextual judgment. The combination is stronger than either alone.

The feedback loop that turns one correct estimate into a sharper system

The most underrated part of the glass-box approach is what happens after the job is done. The model produced an estimate, you verified the reasoning, corrected one assumption, and used the resulting number. The job is now complete. The actual costs are known: real labor hours, real materials used, real waste generated, real time spent on each phase of the work.

If you compare those actuals against the model's estimate at the line level, not just the total, you will find where the model's assumptions were right and where they were wrong. The labor hour estimate for the flat section was accurate. The waste factor for the complex valley section was lower than actual because the model used a generic pitched-roof factor rather than one specific to valleys. That specific finding is the data point that makes the next estimate sharper. You update the assumption, tell the model to use the corrected factor for valley sections going forward, and the next estimate is more accurate without requiring you to supervise every calculation.

Over six months of closing that feedback loop after every job, the AI estimating system learns the specific characteristics of your operation: your team's actual pace on different job types, your supplier's actual prices for the materials you buy most frequently, the waste factors that reflect your crew's specific cutting approach. The system does not learn this automatically. You teach it through the feedback loop. But because the reasoning is explicit, the teaching is targeted rather than broad. You are not re-training a black box. You are correcting a specific assumption in a known place in the reasoning chain.

A roofing company that runs this loop consistently for one full year has a cost estimating system calibrated to its own operation rather than to industry averages. The estimates are more accurate, the margins are more predictable, and the competitive bids are more confident because they are built on data from the company's own jobs rather than generic rules of thumb that apply to the average company in the sector.

Glass box reasoning applied to a pricing decision with real numbers

Let me work through a concrete example from a service business. A carpet cleaning company wants to price a job: a three-bedroom, two-bathroom house with approximately 1,400 square feet of carpet, some high-traffic areas with visible staining. The variables are cleaning time, chemical usage, equipment wear, overhead allocation, and the margin target.

Without a glass-box approach, the owner applies a rule of thumb: multiply the square footage by a per-square-foot rate and adjust for visible condition. The owner has used this rule for years and it works well enough on average. But on a specific job with significant staining, the rule does not capture the extra chemical and extra passes required, which means the quoted price does not fully recover the actual cost. The owner quotes the job, wins it, and finds the margin is lower than the invoice suggested. Over a month, the pattern is invisible because it is buried in aggregate revenue.

With a glass-box approach, the owner feeds the model the job details and asks it to work through the cost structure explicitly, showing each assumption and each calculation. The model returns: base cleaning time estimate for 1,400 square feet at your average productivity rate, plus additional time allowance for the high-traffic areas based on the severity of staining as described, chemical cost calculated from unit price and usage rate per square foot plus stain-specific treatment cost, equipment wear allocated per job-hour at your standard rate, overhead allocated at your monthly fixed cost divided by job count. Total cost: a specific figure. Recommended price at your target margin: a specific number.

The owner reads the reasoning and finds one assumption to correct: the model used a standard stain-treatment chemical usage rate, but the owner knows from experience that this property description implies a rate twenty percent higher than standard. The model recalculates with the corrected assumption and returns a revised price. The owner quotes that price, wins the job, and compares the actual chemical usage afterward. It came in within five percent of the corrected estimate. The feedback loop has confirmed that the corrected assumption is accurate for that staining severity, and the owner records it for use in future quotes with similar descriptions.

The financial difference on a single job is small. The difference across a year of quotes where the staining assumption is correctly applied rather than defaulting to the generic rate is the difference between discovering at the end of the year that a category of jobs was consistently underpriced and never knowing that the problem existed. Glass-box reasoning does not prevent every estimation error. It makes errors visible, correctable, and trackable, which is the property that turns an AI tool from a convenience into a business system.

Do it with an expert
You can build this yourself, or have it set up right the first time.

That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.

Book your call →
Madhuranjan Kumar

Madhuranjan Kumar

Founder, AI DOERS · Performance Marketing

Madhuranjan Kumar brings 20 years of performance-marketing experience and has managed over $200 million in Facebook ad spend for brands across the United States and beyond. His expertise spans the full modern marketing stack: Meta, Google Ads, TikTok, email automation, CRM, and the websites that hold it together. At AI DOERS he turns that track record into lead-generation systems for businesses across every industry.

← Back to all insights
AI Just Showed Its Work on a Hard Problem, and Why That Matters for Your Business | AI Doers