Why 95% of Custom AI Projects Fail and 75% of Companies Still Win
MIT and Wharton are not contradicting each other, they are measuring two different things. Once you see what each number actually counts, you can tell which AI work is a quick win and which is a hard, expensive build.

MIT says 95 percent of custom AI projects fail. Wharton says 75 percent of companies see positive ROI from AI. The popular response is to treat this as a contradiction and ask which study is right. That response draws the wrong lesson from both numbers, and businesses that act on the contradiction framing typically make an expensive decision in the wrong direction.
I am Madhuranjan Kumar, and the position I want to stake here is direct: MIT and Wharton are both right. They are not measuring the same thing, and the gap between their numbers is not noise but a map. That map tells you exactly which type of AI investment pays back fast, which type is genuinely hard to execute, and which mistake most businesses make when they read one of these headlines and try to act on it. The confusion is not the studies' fault. It is a reading problem, and once you resolve it, the right move becomes obvious.
MIT and Wharton are not disagreeing with each other
The reason these two studies appear to contradict each other is that neither one leads with a clear definition of what it measured. MIT's 95 percent failure rate comes from studying custom AI development: the kind where a company commissions a system built around its own specific processes and data. Wharton's 75 percent positive ROI figure comes from studying general AI adoption, a category broad enough to include an employee saving thirty minutes per day using an off-the-shelf AI assistant to draft meeting summaries. Those are different experiments, and you cannot stack their results against each other without first noticing that they are measuring different populations of activity.
MIT was measuring transformation efforts: companies that tried to change how they fundamentally operate using bespoke AI built for their particular situation. Wharton was measuring adoption: companies where people started using AI tools that already existed, were already proven, and required no engineering work to deploy. The populations overlap barely at all. A company using ChatGPT to help the marketing team draft copy is in the Wharton study. A company that hired a development team to build a custom AI system for their inventory workflow is in the MIT study. Those are two completely different experiments with two completely different risk profiles.
The correct reading of the two numbers together is not that AI is a coin flip. It is that AI has a tiered structure. At the base is the generic tool layer, which is low risk and pays back quickly because the tools are already built, already tested, and already have established use patterns. At the top is custom AI development, which is genuinely high risk, slow to pay back, and requires serious technical investment and organizational discipline to succeed. These tiers are not interchangeable, and most of the confusion in the public conversation about AI ROI comes from treating them as if they were the same investment decision.

Custom AI fails for the same reason custom software has always failed
Software projects fail at high rates in general, a fact that was well documented long before AI existed. Studies of large software development projects over the past four decades consistently show failure rates above 50 percent when failure is defined as significantly exceeding budget, missing the deadline, or not delivering the intended functionality. AI projects fail at an even higher rate than conventional software projects, and the reasons are the same plus a few additional ones specific to how AI systems behave in production.
MIT identified three failure patterns that account for the majority of its 95 percent figure. The first is rigidity: systems built to handle one specific configuration of a problem that fail when the real world presents a variation they were not designed for. A human worker adapts to variation because they understand the goal and can reason about edge cases. A rigid AI system does not, because it was built to optimize a specific metric rather than achieve a goal flexibly. The second is human pushback: staff who are afraid of the change, skeptical of the system's reliability, or simply unwilling to alter routines that work for them. The third is skill gaps: development teams who understand machine learning but do not deeply understand the business process they are trying to automate, producing systems that are technically functional but operationally wrong.
These are not new problems. They are the same problems that have plagued ERP implementations, CRM rollouts, and custom software projects for decades. The pattern is: executive sponsor commissions a custom system, development team builds something that works in ideal conditions, the organization is asked to change its habits to accommodate the system, habits do not change, the system sits underused, and the project is judged a failure. AI adds a layer to this pattern because AI systems are less predictable than conventional software and require more active maintenance to stay calibrated as conditions change over time.
The reason external partners succeed at roughly twice the rate of internal builds is not that they have better engineers. It is that they have built these systems before, they have seen where human pushback appears, they know where the edge cases live, and they have a method for the optimization phase that internal teams working from first principles typically do not. Experience with the failure modes is what the external premium is actually buying, and it is worth paying for on projects that genuinely require a custom build.

The generic tool layer is not a consolation prize for businesses that cannot afford the real thing
The way generic AI tools are sometimes discussed implies that they are the accessible version for businesses that cannot justify a real AI investment, and that the serious returns only come from custom builds. That framing is wrong in a specific and important way, and it leads businesses to skip the step that actually produces consistent, measurable results in favor of a step that requires far more to go right.
Roughly 82 percent of employees in the Wharton study use tools like ChatGPT or Copilot weekly. They use them for data analysis, for summarizing documents, for editing drafts, for generating first versions of content, and for answering questions that previously required either a meeting or an extended search. The aggregate lift from those uses is real and measurable. It does not always appear as a named line item in a financial report, but it shows up in hours recovered, in quality improvements, and in the speed at which work moves from draft to finished.
For a team running Meta ads campaigns, the generic tool layer means drafts of ad copy, audience research, and creative briefs can be produced in minutes rather than hours. For a team managing Google Ads accounts, it means budget analysis and keyword research have a faster, better-informed starting point. For a business doing SEO and organic search work, it means content briefs, competitor analyses, and first-draft articles are produced at a pace that was not achievable manually. None of these uses require a custom build. They all require a team with the habit and the prompts to use the tools consistently, which is a habit problem rather than a technology problem.
The businesses that dismiss the generic tool layer as too basic to bother with are also typically the ones whose teams spend the most time on tasks that AI handles well at near-zero cost. The gap between those two behaviors is purely one of adoption. Solving the adoption problem is what the generic tool layer is for. It builds the intuition, the daily workflow, and the understanding of AI's actual behavior that makes every downstream investment in custom work land better. Skipping the generic layer to go straight to custom is not a shortcut to bigger returns. It is a way to miss the foundation that makes custom returns achievable in the first place.
Most companies are attempting the hard version before they have earned it
The sequence that most companies follow when they take AI seriously for the first time is to identify the highest-value automation opportunity, scope a custom project to address it, and spend the budget on a build. That sequence skips the generic tool adoption phase that would have told them whether their teams are capable of trusting an AI system's output, whether the process they want to automate is actually consistent enough to be automated, and whether the problem they think is highest-value is actually the one where AI performs best under realistic conditions.
Small and mid-sized companies hit positive ROI from generic AI tools in roughly 90 days. Large enterprises with more stakeholders and slower change management take roughly 9 months to see the same result. Both timelines are for the base layer of adoption. When a company skips that layer and goes straight to a custom build, the team is using an AI system without having developed the daily habits that make AI systems trustworthy tools rather than black boxes. That lack of familiarity produces exactly the human pushback that MIT identified as a primary failure mode, because people resist systems they do not understand and have no personal history with.
The companies that succeed at custom AI builds are almost never starting from zero with AI. They are the ones whose teams have been using off-the-shelf tools for months, who understand from daily experience what AI is reliable for and what requires human review, and who can tell the difference between a system behaving as intended and a system behaving oddly. That background is what allows them to evaluate a custom system's output intelligently during the testing and optimization phases, where the real calibration happens and where most projects either survive or quietly fail.
The implication for any business considering a custom AI investment is direct: before scoping the custom build, run a 90-day serious adoption sprint on generic tools. Document where the time savings appear, where the tools struggle, and where the team's habits change most readily. That 90-day sprint is not a delay. It is the due diligence that tells you whether the custom project you are considering is solving the right problem, whether your team will adopt the output, and whether the budget for optimization is realistic given what you have learned about how AI actually behaves in your specific workflow.
The optimization phase is where every custom project lives or dies
If there is a single reason MIT's 95 percent failure rate is as high as it is, and it is not inflexible systems or human pushback alone even though those are real, it is that most custom AI projects are scoped and priced without an honest accounting of what the optimization phase costs. The development phase is the visible part: the weeks of engineering work that produce a system that functions in a controlled test. The optimization phase is the part that is routinely underestimated, underfunded, or omitted from the project scope entirely.
The development phase of most custom AI projects takes 4 to 6 weeks. That is the phase that generates the deliverable, the thing the stakeholder was expecting when they signed the contract. The optimization phase takes 4 to 8 additional weeks at minimum, requires consistent feedback from the team actually using the system day-to-day, and involves changes to the model's configuration, its inputs, and often the business process itself. That is the phase where the actual value of the build is produced, because it is where the system is tuned from something that works in ideal conditions to something that works in the conditions the business actually operates under.
Vendors who quote custom AI projects without including an optimization phase are either unaware of how these projects actually go, or they are pricing low to win the engagement and expecting to sell the optimization as a separate project once the client is already committed. Either outcome produces a client who pays for a development phase, receives something that does not quite work in their real environment, and concludes that AI did not deliver on its promise. That client appears in the MIT data, and they are right to be skeptical of AI custom builds, because the project they bought was never scoped to succeed.
The fix is to structure custom AI projects with a clear setup phase, a clear delivery, and a defined period of active management and optimization afterward. A client who understands that 4 to 8 weeks of post-launch tuning is part of the product they are buying will have realistic expectations for how the build goes. A client who was told the project ends at delivery will not. The businesses that get real value from custom AI are almost always the ones whose partners were honest about the second phase from the first conversation, and priced for it rather than hoping the client would not notice it was missing.
That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.
Book your call →
