AVODA Group

Does AI Pay Rent? The SME Evidence Review

Yes, AI can pay for itself in a small business — but the payment is conditional, not automatic. The three strongest studies we have show that AI lifts productivity by roughly 15% when it is scoped, grounded and reviewed; that it actively harms weaker operators who follow its advice uncritically; and that 95% of corporate AI pilots produce no measurable profit-and-loss impact at all. The difference between the firms that profit and the firms that pay tuition is not the model they choose. It is the management discipline they bring to it.

Key Takeaways

  • The best peer-reviewed evidence — a study of 5,172 customer-support agents published in the Quarterly Journal of Economics — found AI assistance raised productivity by about 15% on average, with the largest gains (over 30%) going to the least experienced workers (1).
  • A randomized controlled trial with 640 Kenyan small-business owners found no average effect from a GPT-4 business mentor on WhatsApp: high performers gained just over 15%, while low performers did roughly 8% worse — because they implemented generic advice without judgment (2)(3).
  • MIT’s “GenAI Divide” report found 95% of enterprise generative-AI pilots delivered no measurable P&L impact; the 5% that succeeded bought scoped tools, empowered line managers, and integrated AI into one workflow at a time (4).
  • Purchased, narrowly scoped AI solutions succeeded roughly twice as often as ambitious internal builds in the MIT data — a finding that favors small firms, which cannot afford to build anyway (4).
  • Regulators are now prosecuting AI hype: the US FTC’s 2025 actions against vendors who exaggerated AI business claims confirm that the market will cheerfully sell you into the failing 95% (6).
  • The practical answer for an East African SME: treat every AI deployment like a tenant — give it one room, a written lease, a monthly receipt, and an eviction date if the rent stops arriving.

Why does this question matter more in Kampala than in California?

A venture-funded startup in San Francisco can afford a failed AI experiment. A five-person firm in Kampala, Kisumu or Mwanza cannot. When the OECD surveyed SME adoption across twelve countries, it found that while 61% of small firms had touched at least one AI tool, 76% of those users were “AI novices” using basic tools in isolated corners of the business — activity without transformation (5). Meanwhile the cost of getting it wrong is borne in the hardest currency there is for an African founder: cash flow, customer trust, and the founder’s own scarce attention.

So the question “does AI pay rent?” is not rhetorical. It is the most important due-diligence question in East African business this decade, and for the first time we have evidence good enough to answer it. Three bodies of research — one gold-standard workplace study, one randomized trial run with Kenyan entrepreneurs, and one unsparing audit of corporate pilots — triangulate to a single conclusion. Let us take them in order, because each one corrects a different illusion.

What does the strongest evidence say AI actually does to productivity?

Start with the most rigorous workplace study yet conducted. Erik Brynjolfsson (Stanford), Danielle Li and Lindsey Raymond (MIT) tracked 5,172 customer-support agents at a Fortune 500 software firm as a generative-AI assistant was rolled out in stages — a natural experiment published in the Quarterly Journal of Economics in 2025 (1). The headline: agents with AI access resolved about 15% more issues per hour on average.

But the average hides the finding that should make every East African founder sit up. The gains were wildly uneven — in the opposite direction from what most people assume. The least experienced, lowest-skilled agents improved both speed and quality dramatically; the most experienced agents gained little and in some cases saw slight quality declines. AI compressed the learning curve: an agent with two months of experience plus AI performed like an agent with six months of experience without it (1). The assistant was, in effect, distributing the embedded know-how of the firm’s best performers to its newest ones.

This is the strongest empirical case for what I call the brilliant-intern model of AI: a tireless, fast, occasionally overconfident junior colleague whose value depends entirely on the quality of the instructions and the institutional knowledge you hand it. The Fortune 500 firm in the study did not give its agents a blank chatbot. It gave them an assistant trained on the company’s own successful conversations — grounded in what the firm already knew to be true. That detail is not incidental. It is the whole game, and it is why AI that misquotes its sources is a business problem before it is a philosophical one.

Why did AI make some Kenyan businesses worse off?

Now to the study closest to home — and the most important AI study almost no East African founder has read.

Between 2023 and 2024, researchers Nicholas Otis, Rowan Clarke, Solène Delecourt, David Holtz and Rembrand Koning ran a randomized controlled trial with about 640 small-business owners in Kenya — fast-food joints, poultry farms, salons, cybershops, retail stalls. The treatment group received a GPT-4-powered business mentor delivered through the channel Kenyan business actually runs on: WhatsApp (2)(3).

The result, on average, was nothing. The researchers could not reject the hypothesis that AI access had zero effect on business performance. But the average concealed a sharp divergence. Entrepreneurs who were already high performers at baseline saw profits and revenues improve by just over 15%. Low performers did about 8% worse than their control-group peers (3).

Why? Not because the AI gave different advice to different people. The researchers found the advice was broadly similar. The divergence came from selection and implementation — high performers treated the AI as a thinking partner, filtered its suggestions against their knowledge of their own customers, adapted what fit and discarded what did not. Low performers followed generic advice literally: spending on advertising they could not afford, expanding inventory ahead of demand, applying textbook solutions to non-textbook problems (2)(3). MIT Sloan Management Review summarized the pattern in a title that deserves to be painted on the wall of every accelerator on the continent: “How AI Helps the Best and Hurts the Rest” (7).

Read carefully, the Kenya RCT is not a verdict against AI for African SMEs. It is a verdict against unsupervised AI. The same tool, in the same economy, on the same channel, made disciplined operators measurably richer and undisciplined ones measurably poorer. AI is not a strategy. It is an amplifier of the judgment already present in the firm — which is why the scarcest input in East Africa’s AI economy is managerial judgment, not code.

Why do 95% of AI pilots fail to move the P&L?

The third leg of the evidence comes from MIT’s Project NANDA, whose 2025 report “The GenAI Divide: State of AI in Business” examined 300 public AI deployments, interviewed 150 leaders and surveyed 350 employees. Its finding ricocheted through boardrooms: 95% of enterprise generative-AI pilots delivered no measurable P&L impact. Only about 5% produced rapid, attributable value (4).

Note what the report did not say. It did not say AI fails to work — over 80% of the organizations studied had piloted tools, and individual employees widely reported personal productivity gains. The failure was a failure of translation: activity that never became accountable business results. The report’s authors called the gap between high adoption and low transformation “the GenAI Divide” (4).

The autopsy of the failing 95% reads like a checklist of skipped disciplines:

  1. No baseline. Firms could not say what the process cost before AI, so they could not say what AI saved.
  2. No scope. Pilots aimed at “transformation” rather than one measurable workflow.
  3. No grounding. Generic tools answered from model memory rather than the company’s own verified knowledge, so output required so much correction that the savings evaporated.
  4. No owner. Central AI committees ran pilots; the line managers who owned the actual workflow were bystanders.

And the success factors of the 5% invert the list precisely: purchased, scoped tools succeeded roughly twice as often as internal builds; line managers, not labs, drove adoption; tools were embedded into a single workflow and measured against it (4). A note of method-honesty the hype cycle skipped: the NANDA report was a preliminary study, criticized for its loose definition of “failure,” and even skeptical reviewers concluded its core finding was directionally right — pilots without measurement disciplines do not produce measurable returns, almost by definition (8). That tautology is the lesson.

There is one more actor in this story: the vendor. In 2025 the US Federal Trade Commission brought actions against companies — including business-opportunity sellers Ascend Ecom and Air AI — for deceptive claims about AI-powered business returns (6). The market will sell you into the failing 95% with a straight face and a testimonial video. The evidence base is your only defense.

What separates the firms where AI pays rent?

Put the three studies side by side and a single pattern emerges with unusual clarity for social science:

StudySettingResultThe variable that decided it
Brynjolfsson, Li & Raymond (1)5,172 support agents, Fortune 500+15% avg., +30%+ for novicesAssistant grounded in the firm’s own best conversations
Otis et al. Kenya RCT (2)(3)~640 Kenyan SME owners on WhatsAppHigh performers +15%, low performers −8%Judgment in selecting and implementing advice
MIT GenAI Divide (4)300 enterprise deployments95% no P&L impactScope, ownership and measurement

Grounding. Judgment. Measurement. Three studies, three continents, one conclusion: AI pays rent where a human manager makes it pay, and squats where nobody collects.

This should be the most encouraging research finding an East African founder reads all year. The binding constraint is not capital — the marginal cost of capable AI is now within reach of a realistic Kampala SME budget. It is not infrastructure, which improves every quarter. It is a set of management behaviors that cost nothing but discipline — and discipline is the one input a resource-constrained founder can supply in unlimited quantity.

The Rent Ledger: a four-line test every AI deployment must pass

Here is the framework I teach founders, distilled from the evidence above. Before any AI tool earns a permanent place in your business, open a Rent Ledger for it — four lines, one page, reviewed monthly.

Line 1 — The Baseline. Write down, in numbers, what the target workflow costs you today before AI touches it. Hours per week answering customer questions. Days from inquiry to quote. Percentage of follow-ups that never happen. The MIT data shows this is the line 95% of pilots skip (4) — and without it, every later claim of value is testimony, not evidence.

Line 2 — The Lease. Define the single room the AI occupies: one workflow, one scope, one written brief stating what it may answer, from which approved knowledge, and what it must escalate to a human. The Brynjolfsson study’s grounded assistant and the MIT report’s scoped purchased tools both succeeded on exactly these terms (1)(4). An AI with the run of the whole house pays rent on none of it.

Line 3 — The Receipt. Every month, collect the rent in the same units as the baseline: hours recovered, response time cut, conversion lifted, errors caught. If you cannot express the benefit as a number a skeptical spouse would accept, you have a hobby, not a tenant.

Line 4 — The Eviction Date. Set it in advance — I recommend 90 days. If by then the receipts do not cover the full cost (subscription, data bundles, and the founder-hours spent reviewing output), the tenant leaves. No sentimentality. The Kenya RCT’s low performers lost money precisely because nothing in their process forced a reckoning between advice received and results obtained (2).

The Rent Ledger sounds almost insultingly simple. That is the point. The evidence says the failing 95% are not failing on model selection, prompt sophistication or GPU access. They are failing on lines one, three and four — the parts that require no technology whatsoever.

What should an East African founder do with this evidence this quarter?

First, refuse the two fashionable errors. Error one is AI maximalism — adopting tools because competitors are, vendors call, or the feed says you are behind. The Kenya RCT shows that posture has a measurable cost: −8% (3). Error two is AI dismissal — reading the MIT 95% stat as permission to wait. The same report shows the 5% achieving rapid, compounding gains (4), and the QJE study shows the biggest winners are the least experienced operators (1) — which describes most of the East African SME economy. Waiting forfeits exactly the advantage the evidence says is yours.

Second, start where the learning is leaking, not where the demo is shiny. The highest-yield first deployments in the data are unglamorous: grounded customer-response assistants, follow-up that stops dying in the WhatsApp scroll, records that turn mobile-money flows into readable books. East Africa’s commercial life already runs through chat and mobile money — which is why the most rent-paying deployments on this continent will be WhatsApp-first, grounded and human-escalated, not enterprise dashboards.

Third, run the readiness diagnostic before you spend a shilling. A small firm does not need an enterprise maturity matrix; it needs three answerable questions — where learning leaks, what must always be true, and who reviews before send. If you can answer those, the evidence in this article says you are statistically positioned among the winners before you have chosen a single tool.

The age of AI faith is over; the age of AI evidence has begun, and the evidence is on the side of the disciplined small firm. AI’s returns in this decade will not go to the businesses with the biggest budgets or the earliest adoption dates. They will go — they are already going — to the operators who treat AI the way a good landlord treats a tenant: welcomed warmly, briefed clearly, and asked, every single month, for the rent.

Frequently Asked Questions

Does AI actually increase productivity in small businesses?
The evidence says yes, conditionally. A Quarterly Journal of Economics study of 5,172 support agents found a 15% average productivity gain, with novices gaining most. But a Kenyan RCT found low-performing business owners did 8% worse with AI — gains depend on grounding, judgment and review, not on access alone.

What was the Kenyan AI small-business study?
Researchers from Berkeley and Harvard gave roughly 640 Kenyan SME owners a GPT-4 business mentor on WhatsApp. On average it changed nothing; high performers gained just over 15% while low performers lost about 8%, because weaker operators implemented generic advice without filtering it against local reality.

Why do most AI pilots fail to show ROI?
MIT’s GenAI Divide report found 95% of enterprise pilots produced no measurable P&L impact — not because models failed, but because firms skipped baselines, scoped nothing, grounded tools in no verified knowledge, and assigned no owner. Scoped, purchased tools with line-manager ownership succeeded about twice as often.

How long should an SME give an AI tool to prove itself?
Ninety days is a defensible standard. Record a numeric baseline first, then measure monthly in the same units — hours saved, response time, conversion. If receipts do not cover the full cost (subscription, data, and review time) by day 90, retire the tool and redeploy the attention.

Is it safer for a small firm to wait until AI matures?
The data argues the opposite. The largest documented gains went to the least experienced operators, and scoped tools are already cheap. Waiting forfeits the one advantage small firms hold — speed of disciplined adoption — while competitors bank compounding 15% improvements in customer-facing workflows.

Related Reading

Sources and Evidence

  1. Brynjolfsson, E., Li, D., & Raymond, L., “Generative AI at Work,” Quarterly Journal of Economics 140(2), 2025 (NBER Working Paper 31161). https://www.nber.org/papers/w31161 — Peer-reviewed in a top-five economics journal; 5,172-agent field study; the gold standard on AI workplace productivity.
  2. Otis, N., Clarke, R., Delecourt, S., Holtz, D., & Koning, R., “The Uneven Impact of Generative AI on Entrepreneurial Performance,” SSRN Working Paper 4671369 / Harvard Business School. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4671369 — Pre-registered randomized controlled trial with ~640 Kenyan SME owners; the most directly relevant causal evidence for African small business.
  3. Berkeley Haas Newsroom, “Gen AI field experiment shows mixed results in helping small businesses grow.” https://newsroom.haas.berkeley.edu/research/gen-ai-experiment-shows-mixed-results-in-helping-small-businesses-grow/ — Institutional summary by the authors’ university; source of the +15%/−8% headline figures.
  4. Fortune, “MIT report: 95% of generative AI pilots at companies are failing,” August 18, 2025, reporting MIT Project NANDA’s “The GenAI Divide: State of AI in Business 2025.” https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/ — Major business outlet reporting an MIT research initiative; methodology: 150 interviews, 350 surveys, 300 deployment analyses.
  5. OECD, “AI adoption by small and medium-sized enterprises,” OECD Publishing, December 2025. https://www.oecd.org/en/publications/ai-adoption-by-small-and-medium-sized-enterprises_426399c1-en.html — Intergovernmental statistical authority; 12-country D4SME survey data on SME adoption depth and skills barriers.
  6. Bovis, Friedman & Verner LLP, “AI Washing: How Exaggerating Your Artificial Intelligence Can Get You (and Your Business) in Trouble,” analyzing 2025 FTC enforcement actions including Ascend Ecom and Air AI. https://www.bfvlaw.com/ai-washing-how-exaggerating-your-artificial-intelligence-can-get-you-and-your-business-in-trouble/ — Law-firm analysis of public FTC enforcement records.
  7. MIT Sloan Management Review, “How AI Helps the Best and Hurts the Rest.” https://sloanreview.mit.edu/article/how-ai-helps-the-best-and-hurts-the-rest/ — Editorially reviewed management journal; authors’ own discussion of the Kenya RCT mechanism.
  8. Marketing AI Institute, “That Viral MIT Study Claiming 95% of AI Pilots Fail? Don’t Believe the Hype.” https://www.marketingaiinstitute.com/blog/mit-study-ai-pilots — Industry critique included for methodological balance on the NANDA report’s limitations.

Leave a Comment

Your email address will not be published. Required fields are marked *