AVODA Group

Accelerator Outcomes Measurement: Survival and Revenue

An accelerator that cannot report its graduates’ two-year survival rate, revenue growth, jobs created, and follow-on capital raised is not measuring acceleration — it is measuring its own busyness. The honest scoreboard for any entrepreneur support program is four numbers tracked over 24 months, plus a comparison group that answers the only question that matters: what would these ventures have done without you? Building that scoreboard costs far less than most directors believe, and the programs that build it first will own the next decade of funding, because every serious buyer of acceleration — government, donor, or investor — is learning to demand it.

Key Takeaways

  • The sector’s measurement reckoning is explicit: practitioner analysis now distinguishes outputs (sessions held, mentors recruited, applications received) from outcomes (revenue, survival, follow-on capital), and ecosystem analysts argue programs that cannot report the latter are “managing an event calendar” (1, 2).
  • The templates exist and are free: GALI’s longitudinal method — application-time baselines, annual follow-ups, rejected applicants as a comparison group — has tracked 23,000+ ventures, and ANDE’s impact-measurement hub publishes the instruments (3, 4).
  • Revenue growth, employment, and outside investment are the three most consistently measured outcomes across the GALI dataset — the de facto standard — with survival as the foundation beneath all three (3, 5).
  • The counterfactual is the difference between evidence and theater: 2025 NBER research finds most accelerators add negative value against a no-accelerator benchmark, a result invisible to any program that only tracks its own alumni (6).
  • Instrumentation is cheap now: lean-data phone surveys pioneered by Acumen’s 60 Decibels reach founders in days at a fraction of traditional evaluation cost, and baseline data costs nothing if collected at application (7).
  • Context matters in East Africa: average Ugandan business survival runs roughly 4.85 years, with about 30% of SMEs closing within three years — so a program reporting 80% two-year survival against that base rate is making a measurable, fundable claim (8).

Why Do Programs Count Heads Instead of Businesses?

Because heads are what they were paid to count. Two decades of donor funding trained the entrepreneur-support sector — nowhere more thoroughly than in East Africa — to report beneficiaries trained, workshops delivered, gender ratios achieved, and grants disbursed, since those were the deliverables in the contract. The program optimized what the funder inspected. AfriLabs’ surveys of the continent’s 500+ innovation hubs document a sector fluent in reach numbers — lives touched, founders trained — and largely silent on venture performance (9). This is not dishonesty; it is obedience. The reporting culture did exactly what it was designed to do, and what it was designed to do was the wrong thing.

The result is the event-calendar trap. A program that measures activity will, with complete sincerity, maximize activity: more workshops, more mentors onboarded, more demo days, more memoranda signed. Every one of these is an input masquerading as an achievement. The calendar fills; the question of whether any business is better off remains not just unanswered but unasked. As one ecosystem analyst puts it, a program that cannot report survival, revenue growth, follow-on capital, and founder retention is “managing an event calendar,” not running an accelerator (2). The trap has a cruel second jaw: when the funding environment turned hostile — the story I documented in the ESO funding crisis — programs with ten years of attendance data and zero outcome data discovered they had nothing to argue with. The funders could not tell the good programs from the busy ones, so the cuts could not distinguish them either.

And beneath the incentives sits a quieter failure: most programs never instrumented because they assumed measurement was an evaluation expense — something donors commission consultants to do after the fact — rather than an operating system. That assumption is the thing this article exists to break.

What Belongs on the Honest Scoreboard?

Four outcomes, one timestamp, one comparison. The outcomes are not exotic; they are the three most consistently measured variables across the GALI dataset — revenue, employment, outside investment — plus the foundation beneath them, survival (3, 5).

Survival at 24 months. Is the business still operating two years after the program? This is the floor metric: nothing else means anything for a dead company. It is also the cheapest to verify — a phone call, a mobile-money transaction record, a market visit. In Uganda, where average business survival is roughly 4.85 years and about 30% of SMEs close within three (8), survival against the base rate is a real, defensible claim of value.

Revenue trajectory. Not a snapshot — a slope. Revenue at application (the baseline), at graduation, at 12 and 24 months. Self-reported is acceptable if collected consistently; banded figures (e.g., “UGX 50–100M”) raise response rates and are good enough for program-level inference. The point is not audit-grade precision; it is honest direction and magnitude.

Jobs. Full-time equivalents, including the founder, at the same four timestamps. Jobs are the outcome governments actually buy — the unit of account in nearly every public tender — and the easiest number for a funder’s board to understand.

Follow-on capital — counted, not worshipped. Equity, debt, grants, and supplier credit raised after the program, by instrument. In capital-scarce markets this is a noisy signal of quality — plenty of excellent East African SMEs will rightly never raise institutional money — which is why it is one number among four, not the headline. A program whose only outcome metric is “capital raised by alumni” has merely upgraded its vanity.

The timestamp discipline matters more than any single metric: application, graduation, 12 months, 24 months. The application form is the most undervalued measurement instrument in the industry — it captures the baseline for free, from every applicant, including the ones you reject. Which brings us to the comparison.

The comparison group. “Our alumni grew 40%” is unfalsifiable theater until you can say what comparable non-alumni did. GALI’s pragmatic answer is the rejected-applicant comparison: ventures that applied, were screened by the same process, but did not enter (3). Track a sample of your near-miss rejects with the same 12- and 24-month surveys. It is imperfect — selection effects survive — but it converts your claims from anecdote to evidence, and it is the design insight behind the 2025 NBER finding that most programs add negative value relative to a no-accelerator benchmark (6): a result no program tracking only its own alumni could ever detect about itself. The right tail of programs, as I argued in the evidence on whether acceleration works, is defined by exactly this willingness to be measured against the counterfactual.

How Do You Instrument Outcomes on a Small Budget?

This is where most directors overestimate the cost by an order of magnitude. The full stack, for a program of 20–40 ventures a year, is a spreadsheet, a phone, and discipline.

Baseline at application — cost: zero. Add ten questions to the application form: revenue (banded), employees, capital raised to date, registration status, founder contacts (two numbers, one next-of-kin — contact decay is the great enemy of follow-up). Every applicant fills it; your comparison group is enrolling itself.

Follow-up by phone — cost: cents per venture. The lean-data revolution settled this. 60 Decibels, spun out of Acumen to industrialize its Lean Data method, demonstrated that short, structured phone surveys reach low-income customers and founders in days, at a fraction of traditional evaluation cost, with response quality that satisfies institutional investors (7). A program officer with a call script and an afternoon a month can run 12- and 24-month follow-ups for two cohorts plus a rejected-applicant sample. The fuller case for this method — and what the impact-investing world has learned from it — is the subject of impact measurement the 60 Decibels way.

Verification triangulation — cost: minimal. Self-report drifts optimistic. Spot-check 20% of responses against something harder: mobile-money statements, URA/KRA registration status, a supplier reference, a site visit bundled into mentor travel. You are not auditing; you are calibrating the drift.

One register, one owner. A single venture registry — every applicant, every status, every timestamp — owned by one named staff member. ANDE’s impact-measurement hub publishes instruments and templates for precisely this (4). The institutional failure mode is not bad tools; it is orphaned spreadsheets.

The entire apparatus costs perhaps 2–3% of a typical program budget. Set against what it buys — the ability to answer every funder, every government tender, and every board meeting with evidence — it is the highest-return line item a director controls. The Kauffman Foundation’s measurement brief made the same point for US programs years ago: the barrier is not method or money; it is that nobody required it (5).

The Five-Rung Evidence Ladder: How Honest Is Your Program?

Audits need a scale. I use a five-rung ladder — the Evidence Ladder — to grade any program’s measurement maturity, including the awkward rung most of the industry stands on.

Rung 0 — Activity counts. Workshops held, founders trained, mentors recruited. This is the event calendar. It proves budget was spent.

Rung 1 — A living venture registry. Every applicant and alumnus, with baseline data and current status, in one maintained system. Most programs claiming to be data-driven fail here — they cannot produce a clean list of who they have served and whether those ventures still exist.

Rung 2 — Outcomes at 24 months. The four scoreboard numbers, collected at the four timestamps, for every graduating cohort. This is where “we think we help” becomes “here is what happened.”

Rung 3 — A comparison group. Rejected-applicant tracking or a matched external sample. This is where “here is what happened” becomes “here is what we caused” — the rung the NBER work proves most programs would not enjoy climbing (6).

Rung 4 — Publication. Outcomes and methods published annually, before anyone demands them. This is where measurement stops being reporting and becomes strategy: a published scoreboard is a moat, because competitors without numbers cannot answer it.

Two uses for the ladder. First, internally: a director who knows the program stands at Rung 0 can reach Rung 2 within one cohort cycle — the application-form baseline starts working the day it is added. Second, externally: funders should fund by rung. A program at Rung 3 with mediocre results is a better steward of the next grant than a program at Rung 0 with a beautiful brochure, because the first one can learn and prove it.

The deeper payoff is managerial, not reputational. Measurement is the steering wheel, not the scorecard: a director who watches two-year survival and revenue per graduate designs a different program — longer horizons, fewer stage-mismatched workshops, more working-capital instruments — than one who watches attendance. The metric chooses the program. Choose the metric on purpose. (It will also, in time, choose the calendar: outcome data is the strongest argument for the milestone-based designs I make the case for in the critique of the three-month batch.)

What Should Funders and Governments Demand?

The buyer’s side of this market is where the leverage lives, because programs measure what payers inspect. Four demands, all reasonable, all cheap to comply with for any program actually creating value:

Demand the scoreboard, timestamped. Survival, revenue trajectory, jobs, and follow-on capital at 24 months, for the last two cohorts, with collection methodology stated. Not projections. Not testimonials. A program that has never collected this should be funded — once — explicitly to build Rungs 1 and 2, with renewal conditional on producing them.

Pay partly on outcomes. Tie a meaningful tranche — 20–30% — to verified venture outcomes rather than activities delivered. The design details, and the international evidence behind outcome-based contracting of operators, belong to the procurement argument I make in how governments should buy acceleration; the measurement point is simpler: outcome payment makes the honest scoreboard a revenue requirement, which is the only force stronger than reporting habit.

Fund the comparison group. A few thousand dollars a year of survey cost converts an entire portfolio’s reporting from theater to evidence. No single line in a development budget buys more truth per dollar.

Refuse beneficiary inflation. Reject “lives reached.” Reject cumulative double-counting. Insist that the unit of account is the venture, identified, dated, and re-contactable. The sector’s credibility deficit is, at root, a unit-of-account problem.

And for the directors: do not wait to be asked. Publish first. In a sector where almost nobody can prove anything, the first program in a market to publish two-year survival and revenue against a base rate acquires an authority no marketing spend can buy — with founders, with funders, and with the governments now deciding, post-aid, which operators deserve to survive. The numbers do not need to be perfect. They need to be real, dated, and yours. The event calendar is full everywhere. The scoreboard is still nearly empty, and it is there for the taking.

Frequently Asked Questions

What metrics should an accelerator track?
Four outcomes at four timestamps: venture survival, revenue trajectory, jobs (full-time equivalents), and follow-on capital by instrument — at application, graduation, 12 months, and 24 months. These are the most consistently measured variables in the GALI dataset and the de facto standard for program evaluation (3, 5).

What is the difference between outputs and outcomes?
Outputs are what the program did: workshops held, founders trained, mentors recruited, demo days staged. Outcomes are what changed for the ventures: survival, revenue growth, employment, capital raised. Programs that report only outputs are documenting their own busyness, not their effect on any business (1, 2).

How can a small program afford impact measurement?
By collecting baselines free at application and following up with short structured phone surveys — the lean-data method 60 Decibels industrialized. For 20–40 ventures a year, the full system is a maintained registry plus a few staff-days per quarter: roughly 2–3% of program budget (4, 7).

Why does an accelerator need a comparison group?
Because “our alumni grew 40%” is unfalsifiable without knowing what similar non-alumni did. Tracking rejected applicants — GALI’s pragmatic design — converts claims into evidence. NBER research shows most programs add negative value versus a benchmark, something alumni-only data can never reveal (3, 6).

What should funders require before renewing a program?
The four-number scoreboard for the last two cohorts with stated methodology, a maintained venture registry, and a plan to track a comparison sample. Funders should pay a tranche on verified outcomes and reject beneficiary counts — “lives reached” — as a unit of account (2, 5).

Related Reading

Sources and Evidence

  1. Achinonu, K., 2025. “Measuring What Matters: The Role of Metrics in Designing an Accelerator Programme.” Medium. https://medium.com/@kelechiachinonu/measuring-what-matters-the-role-of-metrics-in-designing-an-accelerator-programme-9ab24418f88f — African practitioner analysis articulating the outputs-versus-outcomes distinction inside program design.
  2. O’Brien, S. “Startup Cities Need to Measure and Expect Outcomes, Not Activity.” https://seobrien.com/startup-ecosystem-metrics — Ecosystem-development analyst; source of the “event calendar” framing and the outcome expectation list (survival, scale, capital, revenue).
  3. Global Accelerator Learning Initiative (GALI). Reports & Publications. https://www.galidata.org/publications/ — The ANDE/Emory University collaboration whose application-baseline, annual-follow-up, rejected-applicant methodology across 23,000+ ventures is the field’s measurement template.
  4. ANDE Impact Measurement Hub. Accelerators & Incubators. https://imm.andeglobal.org/category/accelerators-incubators/ — Sector body’s open library of instruments and templates for ESO outcome measurement.
  5. Kauffman Foundation, 2020. “Measuring Accelerator Performance: Potential Metrics and Considerations” (Issue Brief). https://www.kauffman.org/wp-content/uploads/2021/07/Kauffman_Issue-Brief_Measuring-Accelerator-Performance_2020.pdf — Foundation research establishing standard metric sets and the case that measurement barriers are institutional, not methodological.
  6. National Bureau of Economic Research, 2025. “Beyond Demo Day” (Working Paper 35063). https://www.nber.org/papers/w35063 — Academic working paper finding negative median accelerator value-added against a no-accelerator counterfactual; the strongest argument for comparison-group measurement.
  7. Acumen, 2019. “Acumen Launches 60 Decibels to Make Lean Data an Impact Measurement Standard for Impact Investing.” https://acumen.org/news/acumen-launches-60-decibels-to-make-lean-data-an-impact-measurement-standard-for-impact-investing/ — Primary record of the lean-data method: rapid, low-cost phone surveys as institutional-grade impact measurement.
  8. Academic Journals (African Journal of Business Management), 2023. “Survival of Uganda’s small and medium businesses in a Cox model.” https://academicjournals.org/journal/AJBM/article-full-text-pdf/797CACF70843 — Peer-reviewed survival analysis; source for the ~4.85-year average survival and ~30% three-year closure rates used as base rates.
  9. AfriLabs, 2024. 2024 Impact Report. https://www.afrilabs.com/afrilabs-releases-2024-impact-report-over-500-innovation-hubs-280000-lives-reached-and-a-1-trillion-vision-for-africas-digital-future/ — The continent’s largest hub network; illustrative of the sector’s reach-based (rather than venture-outcome-based) reporting culture.
  10. ScienceDirect (Journal of Business Venturing Insights), 2024. “No substitute for strong institutions: Impact of accelerators on new venture performance.” https://www.sciencedirect.com/science/article/abs/pii/S235267342400043X — Peer-reviewed evidence that accelerator impact varies with institutional context, underscoring why local base rates matter in evaluation.

Leave a Comment

Your email address will not be published. Required fields are marked *