AVODA Group

Why Your AI Pilot Died in Week Six

Why Your AI Pilot Died in Week Six

It started well. There was a demo, a champion, some genuine excitement, and a budget line that went through without much argument. By week three people were still using it. By week six they were not, and nobody said so out loud. Eight months later the project is neither alive nor dead, the licence renewal is coming up, and the honest answer to “did it work” is that nobody can say, because nobody measured anything before it started.

This is the most common shape of AI failure in East African organisations, and almost none of it is caused by the technology.

Key Takeaways

  • Pilots fail on process, not on capability. The five failure modes below account for most of what we see, and four are decided before the tool is chosen.
  • The single largest cause is verification time. A pilot that halves drafting and adds twenty minutes of checking has not halved anything, and this is almost never measured.
  • The second largest is that only the champion could make it work, which means it dies with their annual leave.
  • Kill criteria agreed after the fact are never triggered. Agreed on day zero and signed by the sponsor, they end projects on schedule and cheaply.
  • Stopping a pilot on evidence is a successful outcome, not a failed one. Organisations that cannot stop things pay for that inability for years.

Failure one: nobody measured the task before touching it

Ask a team what the task took before the pilot and you get an estimate. Ask how the estimate was produced and you get a shrug. Without a baseline there is no way to demonstrate improvement, which means the project’s fate is decided by whoever is most enthusiastic in the room, which means it survives when it should stop and stops when it should survive.

The fix costs a week. Time the task as it is actually done, not as the procedure says. Three real runs, timed, by the people who do it. Collect the last ten real outputs, good and bad, as a test set. Write down what good looks like, specifically enough that two people would agree. If week one of a pilot ends without a number, the pilot has not started.

The organisations that skip this always give the same reason: it feels like delay. It is not delay. It is the only part of the pilot that produces a defensible answer.

Failure two: verification time was never counted

This is the quiet one, and in our experience it kills more pilots than any other single factor.

A report that took four hours to draft now takes ninety minutes. Everyone is delighted. What nobody logs is that a senior person now spends forty minutes checking figures that used to be correct because the person who compiled them had produced that report for three years. The task went from four hours of one person to two hours and ten minutes across two people, one of whom is more expensive. On the timesheet that is a saving. In the organisation it is close to a wash, and the senior person’s forty minutes came out of something else.

Measure the total, not the drafting. Minutes on the task, plus minutes checking, equals the number that matters. A pilot that reports a saving without a verification column is reporting half a result.

Failure three: only the champion could make it work

Every AI pilot has a person who wanted it. They are usually good at it, often unusually so, and they carry the project further than its design deserves. Then they take two weeks of leave and the whole thing stops, and everyone learns something they could have learned in week two for free.

A pilot run only by its champion tells you about the champion. Put the tool in the hands of two or three ordinary users doing real work, and pay attention to what they find confusing rather than explaining it away. If the answer is that they need the champion sitting next to them, that is the finding, and it is a finding about the design rather than about the people.

Failure four: the scope was three things instead of one

The proposal names a department, or a function, or a strategy. Customer service. Reporting. Operations. Within a fortnight the pilot is doing three things badly instead of one thing well, and there is no single measure that can move.

Pick one task with a name a colleague would recognise. Not “admin”. Compiling the monthly donor report. Answering order status questions on WhatsApp. First-pass data entry from delivery notes. And pick the boring one, not the exciting one. Repetitive, high-volume, low-judgement tasks have clean measurements and forgiving failure modes. The exciting task usually involves judgement, which is exactly where these tools are least reliable and where a wrong answer costs most.

Failure five: nobody wrote down what would count as failure

This is the one that turns a dead pilot into a budget line. Without agreed kill criteria, a pilot that did not work becomes a pilot that “needs more time”. More time becomes a phase two. Phase two becomes a line in next year’s budget, and by then nobody remembers that the original measure never moved.

Write the criteria on day zero, get the sponsor to sign them, and test them out loud in the final week with the sponsor in the room. A workable standard set:

  • The measure did not move at all.
  • The output still needs a full human rewrite every time, which means a step was added rather than removed.
  • Only the pilot owner can make it work.
  • Checking costs more time than the task saved.
  • It produced a confident wrong answer that reached a customer or a donor, and not as a one-off.
  • The staff in the pilot stopped using it in week three without being told to. Behaviour is more honest than a feedback form.
  • Doing it properly would cost more than the annual value of the time saved.

The regional failure modes nobody puts in the proposal

Two constraints show up in East African deployments that proposals written elsewhere ignore entirely.

Connectivity. These tools are close to useless on a link that drops mid-session. A pilot that runs beautifully in a Kampala head office can fail completely at a district branch, and the failure will be reported as “the system does not work” rather than “the connection is unstable”. Test where the work actually happens.

Power. An outage during a working session costs the session, and in some organisations that is several sessions a week. Where power is unreliable, backup for the machines doing this work is part of the AI budget whether or not anyone calls it that.

Neither of these is a reason not to proceed. Both are reasons to choose a pilot task that tolerates interruption, and to notice when a technology failure is actually an infrastructure failure.

What a pilot that survives looks like

WeekWhat happensWhat it produces
1Time the task three times as actually done. Collect ten real outputs. Set up business-tier accounts. Brief the users for thirty minutes, including what they may not enterA baseline number and a test set
2Real staff do real work with the tool. Expect it to be slower. Log everything that breaks, and log every fixA list of what broke, which is the specification for the real build
3Apply the fixes. Time it again, same method. Have someone outside the pilot judge ten outputs blind. Record checking time separatelyA second number, and an honest verification figure
4Compare, test the kill criteria out loud with the sponsor, cost the real version with what you now know, write one pageA decision with a date on it

Four weeks, one task, one part-time internal owner, and a handful of paid accounts. For most organisations of twenty to sixty people that is a cost somewhere between the price of a laptop and the price of a junior hire. It is the cheapest way to find out that exists.

Budget for two, and expect one to fail

An organisation that plans a single pilot and calls it the AI strategy has bet the whole programme on choosing correctly the first time, with no information. Two smaller pilots cost less than one large one and produce a real answer.

And when one of them fails, write two paragraphs and file them where the next person will find them: what you tried, and what you would do differently. An organisation that runs four honest thirty-day pilots and stops three of them is in a far stronger position than one that ran a single eighteen-month project nobody could end. The three that stopped cost a month each. The one that ran cost a year, a budget line, and the organisation’s appetite for trying again.

That last cost is the one nobody prices, and it is the largest. The real damage of a pilot that drifted is not the money. It is that the next good idea now has to fight the memory of this one.

Leave a Comment

Your email address will not be published. Required fields are marked *