Operational notes Method

95% of AI Experiments Never Reach Production: How to Land in the 5%

3 min read

A technician testing a model aircraft in a wind tunnel
Testing in real-world conditions: the difference between an experiment and a demonstration.

In the summer of 2025, an MIT study (the NANDA project) made headlines across the world’s press with a brutal figure: roughly 95% of corporate generative AI experiments produce no measurable return and never move beyond the demonstration stage. The exact number is open to debate — the scope of the survey, what counts as a “return” — but anyone working in the field knows the order of magnitude is right. The useful question is not whether the statistic is precise: it is why this happens, and what the 5% do differently.

The real causes (not the ones you hear about)

It isn’t the model. Models are the most mature part of the whole chain. The recurring causes lie elsewhere:

  1. The use case is vague. “Let’s put AI on customer service” is not a use case: it’s a wish. A use case is: this type of request, these data to answer with, this metric before and after. Vague experiments don’t even fail — you simply can’t tell whether they worked, which is worse.
  2. The data isn’t ready, and nobody budgeted for that. The model has to answer from company data, but that data sits in seven systems that don’t talk to each other, with mismatched records. The demo works because it uses ten hand-picked documents; production doesn’t, because it has to use everything else. The hard part — connecting the data into an operational model — gets discovered halfway through the project, once the budget and the enthusiasm have run out.
  3. No process around it. The experiment produces answers: and then what? Who receives them, who approves them, what changes in the workflow? An AI system with no process to use it is a demo in perpetuity. Adoption is organisational work, not technical — and it almost always has zero budget.
  4. No metrics. If the cost of the process wasn’t measured before, nobody can prove the saving after. The project dies at the first budget review, not because it failed but because it can’t prove it worked.

What the 5% do

The MIT study also contains the less-quoted, plain-spoken part: the initiatives that work are mostly narrow in scope, integrated into existing processes and often built with specialised external partners, rather than improvised in-house on generic tools. In practical terms:

  • One process, measurable, that hurts. Not “transform the company”: remove the 12 hours a week the technical office spends searching for documents. Win that one, then expand.
  • Data first, model second. Week one is spent on the systems and the records, not on the prompt. If the data you need isn’t reachable, better to find out on day three than in month four.
  • Into production early, at small scale. Twenty real users in three weeks beat a perfect demo in six months: only real use reveals where the system gets it wrong. This is the principle behind our operational trial: one real use case, on your data, in production — and then you grow.
  • The business case written beforehand. Hours saved, errors avoided, planned downtime: figures agreed with whoever signs off the budget, measured the same way before and after.

The right question to ask whoever is pitching you AI

Not “which model do you use?” but: “What’s the first process, with which metric, and in how many weeks will we see it working on our data?” Whoever has a precise answer to this question is working to put you in the 5%. Whoever answers with a slide is working for the statistic.

Do you have an experiment stuck in neutral, or do you want to avoid adding one more to the list? Half an hour to take stock.

Sources