IE Insights / Series C · The Evidence / C·03
    Entrepreneurship Zone🇾🇪

    A Lottery in Sana'a

    In January 2014, two rooms in Yemen decided which firms received a grant by public draw. It produced better evidence about enterprise support than most programs generate in a decade.

    Derived from a randomized controlled trial, 820 applicant firms · 1 figure · 1 table · Yemen, 2014 to 2015 · 12 min read
    Executive Summary

    Almost every enterprise support program reports its results as a count of firms served. Almost none can say what would have happened to those firms without it. That gap is usually explained as a practical impossibility in difficult environments. Yemen ran the experiment anyway, in the year before its civil war, and the design that made it possible is simpler than the excuse for avoiding it.

    1. Oversubscription is a design choice, not luck. The application was made deliberately easy. 820 firms applied for 200 grants, slightly more than four times the funding available.
    2. The draw was held in public. Randomization events in Sana'a and Aden, open to attendees including applicants and media. Transparent allocation and a valid control group are the same act.
    3. The measured effect was large. Firms receiving grants were 37.1 percentage points more likely to undertake an innovative activity, took 1.8 more such activities than the control group, and were 48 percentage points more likely to report sales growth.
    4. The limits are stated, not buried. Civil conflict drove 45 percent attrition, canceled the second round, and made long-run effects unknowable. The authors say so plainly, which is what makes the rest credible.

    The hardest question in enterprise support is not whether the firms you funded did well. It is whether they would have done well anyway. Every program has beneficiaries who grew, and every program can produce them on request. What almost none can produce is a comparable group of firms who wanted the support, qualified for it, and did not receive it, because that group is usually never constructed. Yemen constructed one, in Sana'a and Aden in January 2014, by holding a public lottery.

    01The Problem With Oversubscription

    There is a standing argument that rigorous evaluation of business support is impractical: firms are few, applications are thin, and there is no honest way to construct a control group without denying help to someone who needs it. The first half of that argument has real history behind it. Earlier matching grant programs repeatedly failed to generate enough applications to support experimental analysis, and the small numbers received doomed plans built on oversubscription designs.

    Yemen inverted the problem by treating take-up as something to engineer rather than something to observe. The program was deliberately designed to make applying easy, on the reasoning that a difficult application process filters for administrative capacity rather than for business potential. It worked. In total 820 applications arrived for 200 available grants. Nineteen were rejected, either because the firm was not eligible or because it would not agree to cover its share of the cost. What remained was an eligible pool slightly more than four times the size of the funding.

    That ratio is the whole trick. Once qualified demand exceeds supply by a wide margin, someone has to be turned away regardless of the selection method. The only question is whether the rejected firms are chosen in a way that makes them a useful comparison group. Rank them by committee and they are systematically different from the winners, which is precisely what makes the results uninterpretable. Draw them at random and the two groups are alike in expectation, and every subsequent difference between them is attributable to the grant.

    820 → 200Eligible applications received against grants available. Oversubscription of slightly more than four to one made random allocation both necessary and analytically valuable.

    02A Draw Held in Public

    Selection took place at public randomization events, in Sana'a on 9 January 2014 and in Aden on 12 January, with the draw stratified by city. The events were open, and attendees included applicants themselves, members of the project advisory committee, and media including television. One hundred firms were drawn in each city, along with a reserve list.

    It is worth pausing on why this matters beyond methodology. A public lottery answers a political problem and a statistical problem with a single act. The political problem is that any discretionary allocation of public benefit invites the suspicion, often justified, that connections decided the outcome. A televised draw is the most legible possible answer to that suspicion. The statistical problem is the absence of a counterfactual. The same draw solves it. An agency that runs one gets fairness it can demonstrate and evidence it can defend, and it gets both for the cost of hiring a room.

    The randomization did what randomization is supposed to do. A joint test of orthogonality across baseline characteristics could not reject balance between the treatment and control samples, at a p-value of 0.318. The two groups of firms, before any grant was paid, were statistically indistinguishable.

    A public draw answers the political question and the statistical question with one act. Fairness you can demonstrate, and evidence you can defend, for the price of a room.

    03What the Grants Did

    The follow-up survey was conducted in March 2015, as civil conflict was breaking out. Against the control group, firms that received the matching grant showed changes across the range of activity the program was intended to provoke.

    Outcome measured against control groupEffect
    More likely to have undertaken an innovative activity+37.1 pts
    Additional innovative activities undertaken+1.8 activities
    More likely to report sales grew over the past year+48 pts
    Balance test at baseline (joint orthogonality)p = 0.318
    Take-up among firms selected57.8%
    Attrition in follow-up survey45%
    <b>Measured effects and their limits, side by side.</b> Effects are differences against the randomly selected control group. The balance test confirms the two groups were statistically indistinguishable before treatment. Take-up and attrition are reported because they bound what the effects can be taken to mean. Source: McKenzie, Assaf and Cusolito (2016).

    Grant recipients also introduced more new products, carried out more marketing, were more likely to introduce a new accounting system, trained more workers and made more capital investments. The authors summarize it as additionality: the program produced activity that would not otherwise have occurred. That is a stronger claim than any beneficiary count can support, and it is available only because somebody built the comparison group first.

    Four-step diagram of the Yemen evaluation design: easy application produces 820 applicants for 200 grants, public draw in Sana'a and Aden allocates them, unselected eligible firms become the control group, and the difference between groups is the measured effect
    The design, in four steps. Make the application easy so qualified demand exceeds supply. Draw winners in public. The eligible firms not drawn become the control group at no additional cost. Measure the difference. Each step solves a problem the next one depends on, and none of them requires a research budget larger than the program itself. Source: design as described in McKenzie, Assaf and Cusolito (2016).

    04The Caveats Are the Credential

    What distinguishes this study from most program reporting is not the size of the effects. It is the candor about what the effects cannot tell you, and the list is not short. Conflict broke out as the follow-up survey was in the field, producing 45 percent attrition, far above what a study of this kind would ordinarily tolerate. The second year of the program was canceled outright, halving the intended sample. Long-run impact is therefore unknowable. The reduced sample leaves no statistical power to examine which kinds of firm benefited most. And the analysis measures additionality for the firms that received grants without measuring whether the gains came partly at the expense of competitors who did not.

    Take-up is reported with the same directness. Of the firms selected, 57.8 percent actually used the grant, with drop-outs citing the security situation and a fuel crisis that made them unwilling to risk paying their half of the cost. Comparing those who took the grant with those who did not shows the two groups statistically similar on most characteristics, so take-up was not simply the larger firms self-selecting in.

    Every one of those limitations weakens the headline. Reporting them is what makes the headline worth anything. A result presented without its bounds is a claim; a result presented with them is a finding, and the difference is visible to anyone deciding whether to spend public money on the strength of it.

    05The Record Since

    Yemen's experiment is best understood against what came before it. The reason this design was notable is that earlier attempts at the same thing had failed for a mundane reason: matching grant programs across several countries could not attract enough applications to support experimental analysis, and the small numbers received doomed plans built on oversubscription. The methodological problem was never the statistics. It was that too few firms applied.

    Matching grants themselves are not a fading instrument. They appear across the institutional record from 2000 through 2026, in agriculture, in innovation, in enterprise support. What has not appeared, in our holdings, is a second public-lottery evaluation of this kind at comparable rigor. A decade after Sana'a and Aden, the design that solved the problem remains a demonstration rather than a standard. That is the gap this piece is really about: not whether the method works, which it did, but why a method that works and costs almost nothing has not been adopted more widely.

    06The Builder's Reading

    The transferable lesson is that evaluation is a design decision made at the start of a program, not an exercise bolted on at the end. By the time a program is running and allocating on merit, the opportunity to know whether it worked has already been spent.

    1. Engineer oversubscription. Make the application genuinely easy. Historic evaluation failures in this instrument were failures of take-up, not of statistics.
    2. Set eligibility, then randomize among the eligible. The draw does not lower the bar. It decides only among firms that have already cleared it.
    3. Hold the draw in public. The transparency is not a nicety. It is the answer to the accusation that will otherwise follow the program for its whole life.
    4. Track the firms you turned down. They are the control group and they cost nothing to create. Losing contact with them is the single most common way this evidence is thrown away.
    5. Publish the attrition, the take-up and the power limits. A finding with stated bounds outranks a headline without them, and every reader who matters knows it.

    The wider point is about what counts as a hard environment. Yemen was the poorest country in the Arab world, with a per-capita GDP of 1,473 US dollars in 2013 and a private sector in which an estimated 88 percent of firms employed fewer than five workers, and it was months from war. If a credible randomized evaluation of enterprise support could be run there, the claim that it is impractical elsewhere is doing other work. Most programs do not lack the conditions to find out whether they work. They lack the decision to.

    Figures are from the randomized evaluation of Yemen's business development services matching grants, covering the first and only completed round. Effects are intention-to-treat comparisons against the randomly assigned control group. The 45 percent attrition rate is unusually high and is a direct consequence of the outbreak of civil conflict during the follow-up survey; the authors report that the interviewed sample remained balanced between treatment and control on observed characteristics. Long-run effects, heterogeneity across firm types, and effects on non-participating competitors are outside what this study can establish. No IEPA engine outputs are used in this piece.