Every claim in this series rests on knowing what would have happened otherwise. There are only two ways to find out. One study did both, on the same firms, and the answers were not the same.
Behind every finding in this series sits the same question: would this have happened anyway? It is the question that separates an instrument that works from an instrument that pays for something already in motion, and there are exactly two ways to answer it. You can ask the people involved, or you can construct a comparison and measure. Both are used. They are not equivalent, and the difference is not academic.
A program reports that it supported four hundred firms and that those firms grew. Both statements can be entirely true and tell you nothing about whether the program worked, because the firms that apply to a program are not a random sample of firms. They are the ones with the ambition to apply, the capacity to complete the paperwork, and often the prospects that made them attractive in the first place. Selecting well and treating well produce the same annual report. Distinguishing between them requires knowing what those four hundred firms would have done without you, and that is a fact about a world that did not happen.
The first way is to ask. Devised in 1985 and applied through the surveys covered earlier in this series, the extreme test puts the counterfactual directly to the participant: would you have made this investment if everything else were identical and the incentive had simply not been offered? Firms answering no are the ones the instrument reached. Firms answering yes received a payment that changed nothing.
Its strengths are real. It can be run after the fact, which means it works on programs that were designed without any thought of evaluation, and that is most of them. It scales across countries cheaply. Nearly everything this series has been able to say about incentive redundancy exists because somebody asked.
The second way is to measure. Construct a group of firms who wanted the support, qualified for it, and did not receive it, then compare. That group has to be built before the program allocates anything, which makes it a decision taken at the design stage or never. The framing in the Yemen study is exact: randomized controlled trials can provide the counterfactual needed to answer the additionality question, but efforts to experiment with matching grant programs have often failed.
Asking can be done afterwards, which is why most evidence uses it. Measuring has to be decided before anything is allocated, which is why most programs cannot.
The Yemen matching grant experiment is unusual because it ran both methods on the same population. Firms that received grants were asked directly, in the follow-up survey, whether they would have undertaken the activities without the grant. Alongside that, the randomized design produced what the authors describe as a more rigorous measurement of additionality.
| Asked directly: would you have done this without the grant? | Share of recipients |
|---|---|
| No, would not have done it | 43% |
| Would have done it, but at a smaller scale | 41% |
| Would have done the same activity anyway | 16% |
Read the three rows as a distribution rather than a verdict. Only 16 percent describe themselves as pure deadweight. Only 43 percent describe the grant as decisive. The largest single group, 41 percent, sits between: they would have acted, but smaller. The grant changed the size of the thing, not whether it happened.
This is where the argument turns back on the series itself. The redundancy ratios in the earlier pieces, the 70 to 85 percent, come from asking investors a question with two answers. Yemen shows that when firms are given a third option, four in ten take it.
That does not invalidate the redundancy findings, and it is worth being careful about why. A firm that would have invested at a smaller scale is genuinely different from a firm the incentive brought in, and counting it as redundant is defensible: the investment was coming. But it is also different from pure deadweight, because something did change. A binary instrument has to assign that firm to one side or the other, and whichever side it picks, the number it reports is a simplification of a gradient.
The direction of the simplification is what matters for policy. If a survey codes the partial cases as redundant, redundancy is overstated and the instrument looks worse than it is. If it codes them as additional, redundancy is understated and the instrument looks better. Neither error is obviously more likely, which means the honest reading of any single redundancy ratio includes a band of uncertainty that the headline figure does not display.
Given that measurement is stronger, the question is why so little of it exists. The Yemen study is direct about the practical obstacle, and it is not statistical sophistication. Earlier attempts at experimental evaluation of matching grants failed because too few firms applied, and the small numbers received doomed plans built on oversubscription designs. The binding constraint was demand for the program, not method.
There is a second obstacle visible elsewhere in this series. Ex-post evaluations identifying the impact of an intervention can only take place if the proper collection systems were established at the outset and maintained throughout implementation. A program that did not plan to be evaluated has, by the time anyone asks, destroyed its own ability to answer. Not through negligence, simply by never recording the thing that would have made the comparison possible.
The Yemen authors also record the smaller constraints honestly, and they are the ordinary kind. Logistical and time pressures limited how many questions the follow-up survey could ask. A direct question about sales drew 51 percent item non-response among firms that answered the survey at all. The heterogeneity of the applicant pool left insufficient power to examine which kinds of firm benefited most. These are not exotic problems. They are what fieldwork looks like, and any program planning to measure should expect them.
The practical question is not which method is better in the abstract. It is which one your situation permits, and what you are entitled to claim afterwards.
The larger point is about intellectual posture. This series has spent eight articles pressing agencies on whether they know what their instruments accomplished. The same standard has to apply to the evidence used to press them, and applied honestly it says this: most of what we know about these instruments comes from asking, asking is the weaker of the two available methods, and the one program that did both found a large middle group that asking alone would have mislabeled. The findings hold. The confidence attached to them should be calibrated to the method that produced them, and usually is not.
The 43, 41 and 16 percent figures are self-reported responses from firms that used the matching grant in Yemen, collected in a follow-up survey conducted in March 2015. They are distinct from the experimental estimates reported elsewhere in this series, which compare selected firms against the randomly assigned control group. The extreme test is the method introduced by Guisinger and Associates in 1985 and applied in the investor surveys covered in the first article of this series. No IEPA engine outputs are used in this piece.