IE Insights / Series C · The Evidence / C·09
    Entrepreneurship Zone🇾🇪

    Ask, or Measure

    Every claim in this series rests on knowing what would have happened otherwise. There are only two ways to find out. One study did both, on the same firms, and the answers were not the same.

    Traced across 1985 to 2016 · the extreme test against a randomized trial · World Bank investment climate and enterprise evidence · 11 min read
    Executive Summary

    Behind every finding in this series sits the same question: would this have happened anyway? It is the question that separates an instrument that works from an instrument that pays for something already in motion, and there are exactly two ways to answer it. You can ask the people involved, or you can construct a comparison and measure. Both are used. They are not equivalent, and the difference is not academic.

    1. Asking is the older method and it carries most of the evidence. The extreme test, devised in 1985, produces the redundancy ratios of 70 to 85 percent that this series has relied on. It is cheap, it scales, and it is the only method available retrospectively.
    2. Measuring means building a comparison group before the program starts. That is a design decision taken at the outset or not at all, which is why so few programs can do it.
    3. One study did both on the same firms. Asked whether they would have acted without the grant, 43 percent of Yemeni recipients said no, 41 percent said they would have acted at a smaller scale, and 16 percent said they would have done the same thing anyway.
    4. That middle category is the finding. Four firms in ten sat between yes and no. A method that forces a binary answer cannot see them, and most of what we know about these instruments comes from methods that force a binary answer.

    A program reports that it supported four hundred firms and that those firms grew. Both statements can be entirely true and tell you nothing about whether the program worked, because the firms that apply to a program are not a random sample of firms. They are the ones with the ambition to apply, the capacity to complete the paperwork, and often the prospects that made them attractive in the first place. Selecting well and treating well produce the same annual report. Distinguishing between them requires knowing what those four hundred firms would have done without you, and that is a fact about a world that did not happen.

    01Two Ways to Know

    The first way is to ask. Devised in 1985 and applied through the surveys covered earlier in this series, the extreme test puts the counterfactual directly to the participant: would you have made this investment if everything else were identical and the incentive had simply not been offered? Firms answering no are the ones the instrument reached. Firms answering yes received a payment that changed nothing.

    Its strengths are real. It can be run after the fact, which means it works on programs that were designed without any thought of evaluation, and that is most of them. It scales across countries cheaply. Nearly everything this series has been able to say about incentive redundancy exists because somebody asked.

    The second way is to measure. Construct a group of firms who wanted the support, qualified for it, and did not receive it, then compare. That group has to be built before the program allocates anything, which makes it a decision taken at the design stage or never. The framing in the Yemen study is exact: randomized controlled trials can provide the counterfactual needed to answer the additionality question, but efforts to experiment with matching grant programs have often failed.

    Asking can be done afterwards, which is why most evidence uses it. Measuring has to be decided before anything is allocated, which is why most programs cannot.

    02Both, on the Same Firms

    The Yemen matching grant experiment is unusual because it ran both methods on the same population. Firms that received grants were asked directly, in the follow-up survey, whether they would have undertaken the activities without the grant. Alongside that, the randomized design produced what the authors describe as a more rigorous measurement of additionality.

    Asked directly: would you have done this without the grant?Share of recipients
    No, would not have done it43%
    Would have done it, but at a smaller scale41%
    Would have done the same activity anyway16%
    <b>Self-reported additionality among firms that used the grant.</b> The middle row is the one a binary counterfactual question cannot represent: these firms were not caused to act, and were not unaffected either. Source: McKenzie, Assaf and Cusolito (2016).

    Read the three rows as a distribution rather than a verdict. Only 16 percent describe themselves as pure deadweight. Only 43 percent describe the grant as decisive. The largest single group, 41 percent, sits between: they would have acted, but smaller. The grant changed the size of the thing, not whether it happened.

    41%Firms that would have acted anyway, but at a smaller scale. Neither caused nor unaffected, and invisible to a yes-or-no question.

    03What the Middle Row Does to the Rest

    This is where the argument turns back on the series itself. The redundancy ratios in the earlier pieces, the 70 to 85 percent, come from asking investors a question with two answers. Yemen shows that when firms are given a third option, four in ten take it.

    That does not invalidate the redundancy findings, and it is worth being careful about why. A firm that would have invested at a smaller scale is genuinely different from a firm the incentive brought in, and counting it as redundant is defensible: the investment was coming. But it is also different from pure deadweight, because something did change. A binary instrument has to assign that firm to one side or the other, and whichever side it picks, the number it reports is a simplification of a gradient.

    The direction of the simplification is what matters for policy. If a survey codes the partial cases as redundant, redundancy is overstated and the instrument looks worse than it is. If it codes them as additional, redundancy is understated and the instrument looks better. Neither error is obviously more likely, which means the honest reading of any single redundancy ratio includes a band of uncertainty that the headline figure does not display.

    04Why Measurement Is Rare

    Given that measurement is stronger, the question is why so little of it exists. The Yemen study is direct about the practical obstacle, and it is not statistical sophistication. Earlier attempts at experimental evaluation of matching grants failed because too few firms applied, and the small numbers received doomed plans built on oversubscription designs. The binding constraint was demand for the program, not method.

    There is a second obstacle visible elsewhere in this series. Ex-post evaluations identifying the impact of an intervention can only take place if the proper collection systems were established at the outset and maintained throughout implementation. A program that did not plan to be evaluated has, by the time anyone asks, destroyed its own ability to answer. Not through negligence, simply by never recording the thing that would have made the comparison possible.

    The Yemen authors also record the smaller constraints honestly, and they are the ordinary kind. Logistical and time pressures limited how many questions the follow-up survey could ask. A direct question about sales drew 51 percent item non-response among firms that answered the survey at all. The heterogeneity of the applicant pool left insufficient power to examine which kinds of firm benefited most. These are not exotic problems. They are what fieldwork looks like, and any program planning to measure should expect them.

    05The Builder's Reading

    The practical question is not which method is better in the abstract. It is which one your situation permits, and what you are entitled to claim afterwards.

    1. If the program is already running, ask. It is the only method available retrospectively and it produces real information. Do not treat the impossibility of an experiment as a reason to report nothing.
    2. Offer three answers, not two. Would not have done it, would have done it smaller, would have done it anyway. Yemen's middle category held 41 percent of firms, and a binary question would have forced every one of them into a misleading bucket.
    3. If the program is being designed, decide now. Oversubscription plus random allocation among the eligible is the whole mechanism, and it has to be chosen before anything is awarded.
    4. Record the losers. The applicants you turn down are the comparison group and they cost nothing to create. Losing contact with them is the most common way this evidence is discarded.
    5. State which method produced your number. A self-reported additionality figure and a measured one are different claims. Presenting them interchangeably is how a soft finding acquires a hard reputation.

    The larger point is about intellectual posture. This series has spent eight articles pressing agencies on whether they know what their instruments accomplished. The same standard has to apply to the evidence used to press them, and applied honestly it says this: most of what we know about these instruments comes from asking, asking is the weaker of the two available methods, and the one program that did both found a large middle group that asking alone would have mislabeled. The findings hold. The confidence attached to them should be calibrated to the method that produced them, and usually is not.

    The 43, 41 and 16 percent figures are self-reported responses from firms that used the matching grant in Yemen, collected in a follow-up survey conducted in March 2015. They are distinct from the experimental estimates reported elsewhere in this series, which compare selected firms against the randomly assigned control group. The extreme test is the method introduced by Guisinger and Associates in 1985 and applied in the investor surveys covered in the first article of this series. No IEPA engine outputs are used in this piece.