AI 19 August 2026

Why most AI projects in small companies never reach production

The pilot works. The demo gets applause. Then it quietly dies in month three. The failure is almost never technical, and it is predictable enough to design around.

Hugo Cardellach 5 min read

The pattern is boringly consistent

A director sees a demo. Someone builds a prototype in a fortnight. It works on the ten examples they tested, everyone is impressed, then nothing changes in the business.

Six months later the tool has three users, all of them the person who built it. Nobody kills the project officially. It just stops being mentioned.

We have walked into this room many times. The diagnosis is rarely the model.

Four reasons, in order of how often we see them

1. It was never attached to a number

A pilot that saves undefined time cannot be defended in a budget meeting. If nobody agreed the baseline before it started, nobody can prove it worked afterwards.

Baseline first, always. How long does this take today, how often does it happen, how often is it wrong? Twenty minutes of measurement beats a quarter of argument.

2. It lived outside the tools people use

The prototype sits on its own web page. Your team lives in a CRM, an inbox and a chat window. Asking them to open a fourth thing costs more attention than the tool gives back.

This is the single most common reason a good pilot dies. Adoption is a distance problem before it is a quality problem.

3. It was 90 percent right, with no plan for the other 10

Software that is correct 90 percent of the time is excellent. That same accuracy with no exception queue is a trap, because a human has to check all of it to find the tenth case.

Design the failure path on day one. What does the machine do when it is unsure, who sees that, how fast?

4. Nobody owned it after the demo

The consultant left. An internal champion changed roles. Then a model provider shipped an update, the output drifted, and somebody quietly went back to the old spreadsheet.

A pilot proves the technology can work. Production proves your company can hold it. Those are different projects, and only one of them pays.

A worked example of the cost

A 35 person recruitment firm spent four months building an AI screener for inbound CVs. Budget was $28,000 in external work plus roughly 120 hours of internal time.

$28kExternal spend
120 hInternal hours
11People who tried it
2Still using it at month four

The screener was accurate. Its problem was that recruiters had to upload a CV to a separate app, wait, then copy a summary back into the applicant tracking system.

Added steps per candidate: three. Time saved per candidate: four minutes. The tool asked two of those minutes back in copy and paste, so the real gain was thin enough to ignore under pressure.

We rebuilt it in nine days inside the tracking system itself. Screening now happens on arrival, the summary appears in the candidate record, and nobody uploads anything. Usage went from two recruiters to nine because the work moved to zero extra clicks.

The lesson in one line

The rebuild cost a fraction of the original and changed nothing about the model. It changed where the model lived.

What a project that survives looks like

The shape is different from the start. Not more careful, just aimed at production rather than at approval.

  • One process, one measurable number, one named owner inside the company.
  • Built where the work already happens, even when that platform is duller to build in.
  • A sandbox with real historical data before anything touches a live client.
  • An exception queue from day one, with a person who reviews it daily.
  • Something running in the business inside five weeks, not a roadmap for month six.

That last point is the one owners push back on, usually because a longer plan feels more serious. Six months of building without production contact is not thoroughness. It is a bet placed on assumptions nobody has tested.

Where AI genuinely earns its place

We use models where rules cannot reach: reading unstructured documents, classifying free text, drafting in a trained voice, summarising a call into a decision. Everything else is better as plain automation, which is cheaper to run and far easier to debug.

A system that only uses AI for the language part tends to survive, because the boring 80 percent keeps working when a model has a bad day.

Before you approve the next one

Ask four questions in the meeting. They take a minute each and they are uncomfortable on purpose.

  1. What number moves, and what is it today?
  2. Which tool does the user already have open when this happens?
  3. What happens when the output is wrong, and who sees it?
  4. Who owns this in month six, by name?

A project that cannot answer all four is not ready. That is a cheap thing to discover now.

The month three test

Put a date in the calendar three months after go live, before anybody builds anything. On that day you check two numbers.

How many people used it last week, and whether the baseline moved. Usage below half your intended users means the tool sits in the wrong place, whatever the accuracy report says.

A baseline that has not moved means the process was never the constraint. Either answer is cheap to fix in month three and painful to discover in month nine.

We book that meeting at kickoff and hold it even when the project is obviously working. The habit is what makes the bad news arrive on time.

What to do with this

Five things you can do tomorrow.

  1. Write the baseline number for your current pilot on one line. If you cannot, stop and measure it this week.
  2. Open the tool your team uses most. Ask whether the pilot could live inside it instead.
  3. Define the exception path: what happens when the AI is unsure, and who reviews it.
  4. Put one name against ownership in month six, and tell that person.
  5. Set a date five weeks out for something in real use. Cut scope until that date is credible.
AI & Operations Audit

Want this done on your company, with your numbers?

The audit is $150, and it comes off your first service in full. You get a written plan in two working days: the processes that cost you, ranked by return, with a KPI on each one.