Article Image

Why 95% of Enterprise AI Projects Fail

Samuel Wyndham · 24th August 2026

The figure doing the rounds is that 95% of enterprise AI projects fail. It comes from MIT Project NANDA's The GenAI Divide: State of AI in Business 2025, and the finding underneath it is more specific than the headline: across some $30–40 billion of enterprise spend, 95% of generative AI pilots produced no measurable P&L impact. Only 5% were extracting real value.

Not that they broke. Not that the technology did not work. That nothing showed up in the numbers.

I think there are three reasons, and none of them are technical.

One: AI removes the thinking and hands back the checking

This is the big one and it is almost never scoped for.

The pitch is that AI does the work. What actually happens is that AI does the thinking part of the work and then quietly transfers a new obligation to you — checking whether it was right.

That checking is not free. On any task where being wrong matters, someone has to read the output with enough attention to catch a confident mistake, which is a harder kind of reading than writing it yourself would have been.

Run the honest sum on a lot of enterprise deployments and the promised time saving comes out around 15% of what was projected. The rest went into verification. Nobody put verification in the business case, so the business case was wrong before anyone typed a prompt.

You can get that number back up. But you get it back by designing the checking step deliberately — deciding who checks, how, against what, and where an agent can do the first pass — not by pretending it does not exist.

Two: the self-improvement story is not true yet

There is a persistent claim that these systems get better on their own once they are running. Point it at your process, let it learn, walk away.

I have not seen it. What I have seen is a manual checking loop that never ends, run by someone whose job it quietly became.

Systems do improve, but the improvement comes from a human noticing a failure pattern and changing the instructions, the process document, or the data the thing is working from. That is a real job with a real cost, and if nobody owns it, the deployment does not get better — it just accumulates unexamined errors until someone loses confidence in it and stops using it. Which is what most "failed" AI projects actually look like from the inside. Not a crash. A quiet abandonment.

Three: it was oversold, so it could only disappoint

The promise was unlimited productivity. Anything less than unlimited then reads as failure, even when the delivered result is genuinely good.

I have watched a deployment that saved real money get written up as a disappointment, purely because the original pitch had been unbounded. If you promise the moon and deliver a very good bridge, you have still failed against the brief you set.

The vendors did this. Plenty of consultants did this. It is the single easiest of the three to fix, because it costs nothing but honesty at the start.

What careful adoption actually looks like

The projects I have seen work all share a shape, and it is unglamorous. They are scoped as a sentence:

Reduce time on X to get Y, without Z.

  • X is one named process. Not "our operations." A specific, repeatable thing that happens weekly, that someone can describe start to finish.
  • Y is the outcome you are buying with the time — capacity for a person to do something better, a faster turnaround the customer actually notices, a queue that stops backing up.
  • Z is what you are not willing to trade. Accuracy below a threshold. Client data leaving the tenancy. A human disappearing from a decision that needs one.

Every part of that sentence does work. X keeps the scope small enough to finish. Y stops you celebrating a time saving that went nowhere. Z is the one people leave off, and it is the one that keeps you out of trouble.

Here is the scale that succeeds, honestly stated: on a recent piece of work for a large enterprise, the result was roughly half a full-time employee's worth of time returned. Not a department. Half a person.

That is not a disappointing number. On an enterprise cost base it is substantial, it recurs every year, and it was delivered on a scope small enough that we could prove it. Ten of those is five people. Nobody gets ten of those by starting with a project that tries to be all ten at once.

The real lesson

AI succeeds on limited, specific problems, in businesses honest enough to measure the result against a promise that was reasonable in the first place.

The 95% is not a story about the technology being bad. It is a story about scope being infinite, checking being invisible, and expectations being set by people who were not going to be there at the end.

If you would like a scope written as an actual sentence — with the X, the Y and the Z in it — come and have a chat with us. Or get this kind of thinking weekly.

Want to see how you can upgrade your IT and be even more productive?