IE Insights / Series C · The Evidence / C·06
    Alignment Zone🇦🇲 🇮🇳 🇬🇪 +11

    Evidence of Results

    Fourteen innovation agencies were studied by the institution that helps fund them. Under a heading marked Impact Evaluations, most of what appears is a count of beneficiaries.

    Derived from 14 innovation agency case studies · 1 figure · World Bank comparative agency research, 2019 · 11 min read
    Executive Summary

    An innovation agency is the institutional centre of most national innovation strategies: the body that runs the grants, the vouchers, the matching schemes and the incubation programs. In 2019 fourteen of them across developing and emerging economies were examined together. The study is candid about what it found when it asked how the agencies knew their programs worked.

    1. The finding is stated in the source's own words. Only few of the cases studied in the report had undergone rigorous impact evaluations of their programs.
    2. The results sections count instead. Under headings marked Impact Evaluations and Evidence of Results, agencies report beneficiaries supported, companies funded and start-ups assisted. Those are outputs, not effects.
    3. One attribution claim shows the problem exactly. One agency records a contribution of 5 percent of its country's annual economic output. No counterfactual accompanies it.
    4. Robust monitoring and evaluation is named as a prerequisite. It appears as one of seven building blocks the study identifies for agency success, alongside diagnostic-based interventions and capable staff.

    There is a particular kind of document that is more useful than it intends to be. A comparative study of fourteen innovation agencies, published to help governments build better ones, sets out each agency's structure, mandate and results. Read for its intended purpose it is a competent survey of institutional design. Read for what its results sections actually contain, it is the clearest available account of how little the field knows about whether any of this works.

    01Fourteen Agencies

    The study covers agencies across a wide span of economies: the Enterprise Incubator Foundation in Armenia, BIRAC in India, GITA in Georgia, HAMAG-BICRO in Croatia, ICTA in Sri Lanka, the Innovation Fund in Serbia, iNNpulsa in Colombia, Kafalat in Lebanon, MTDC in Malaysia, NCBR in Poland, RDB in Rwanda, SPRING in Singapore, TIA in South Africa and TTGV in Turkey. Each case draws on interviews with agency staff and reviews of program materials including annual reports, budget briefs and impact evaluations, with cross-case analysis to find common patterns.

    That is a serious research design, and it is worth saying that the resulting cases are genuinely informative about how these institutions are built. The observation that follows is not a criticism of the study. It is a criticism of the field the study documents, and the study is the source of it.

    02What the Results Sections Contain

    Each case carries a section headed Impact Evaluations and Evidence of Results. The heading sets an expectation about what will follow. What follows, in case after case, is a tally.

    AgencyReported under 'Impact Evaluations / Evidence of Results'
    BIRAC, IndiaApproximately 830 beneficiaries in total; 384 SMEs supported
    Enterprise Incubator Foundation, ArmeniaOver 600 companies supported; contributed 5% of annual economic output
    GITA, Georgia175 start-ups supported; physical infrastructure investments
    ICTA, Sri LankaGovernment network established; over 50 e-services implemented
    HAMAG-BICRO, Croatia2011 evaluation of RAZUM: research capacity improved, positive spillovers
    <b>Outputs reported under an impact heading.</b> A representative selection of what appears in the results sections of the fourteen cases. Counts of beneficiaries, companies and projects predominate. Croatia's 2011 evaluation of the RAZUM program is among the few instances where an actual evaluation is reported. Source: Kapil and Aridi (2019).

    None of these numbers is false, and none is unimportant to the agencies that produced them. They are simply answers to a different question. Eight hundred and thirty beneficiaries tells you the scale of an operation. It does not tell you what happened to those beneficiaries that would not have happened otherwise, which is the question the heading promises to answer and the only one that justifies the budget.

    A count of beneficiaries measures the size of a program. It is silent on the one thing a program is funded to change.

    03The Attribution Problem, In One Line

    One entry deserves separate attention because it shows the failure mode at full extension. Armenia's Enterprise Incubator Foundation is recorded as having supported over 600 companies representing an annual turnover of 500 million US dollars, and as having contributed 5 percent of Armenia's economic output, alongside benefits to more than 15,000 employees, 60 universities and 9,000 students.

    Consider what would be required to establish that 5 percent figure. You would need to know what those 600 companies would have produced without the foundation's involvement, which means knowing which of them would have existed, at what scale, on what timeline. No such counterfactual is offered, and constructing one after the fact is close to impossible. The number is not so much wrong as unfalsifiable: there is no observation that could contradict it, which means it carries no information about performance even though it reads as the strongest possible claim in the document.

    The pattern is familiar from elsewhere in this series. An incentive regime counts investment that would have arrived anyway. A guarantee scheme counts loans that would have been made anyway. An innovation agency counts firms that would have grown anyway. In each case the instrument takes credit for the whole of an outcome it contributed some unmeasured fraction of, and in each case the alternative was available and not taken.

    04The Study Says So Itself

    What makes this document unusually valuable is that it does not leave the reader to infer the conclusion. Discussing evaluation, it states that while only few of the cases studied in this report have undergone rigorous impact evaluations of their programs, their importance cannot be emphasized enough, and recommends that innovation agencies develop their own internal monitoring and evaluation capabilities.

    It also makes the sequencing explicit, which is the practical part. Ex-post evaluations identifying the impact of an intervention can only take place if the proper collection systems were established at the outset and maintained throughout implementation. Evaluation is not something that can be commissioned once a program has run. If the data was not designed in at the start, the question is permanently unanswerable, which is why so many programs of long standing cannot say what they achieved.

    Seven building blocks identified for innovation agency success: clear adaptable mission, capable staff, effective governance, diagnostic-based interventions, robust monitoring and evaluation, sustainable funding, strategic partnerships
    The seven prerequisites the study identifies. Drawn from cross-case analysis of the fourteen agencies. Two are worth noting against the rest of this series: diagnostic-based interventions, which is binding-constraint reasoning stated as institutional practice, and robust monitoring and evaluation, which the same document reports most of the studied agencies had not achieved. Source: Kapil and Aridi (2019).

    The seven building blocks the study derives are a clear but adaptable mission, capable staff, effective governance and management structures, diagnostic-based interventions, robust monitoring and evaluation, sustainable funding, and strategic partnerships and networks. Two of those are load-bearing for the argument of this series. Diagnostic-based intervention is binding-constraint reasoning restated as an institutional requirement. And robust monitoring and evaluation appears on a list of prerequisites for success in the same document that reports most of the agencies studied did not have it.

    05The Record Since

    This study is from 2019, and it remains the most comprehensive comparative examination of innovation agencies in our holdings. That is worth stating plainly, because the finding it reports is precisely that these institutions are not rigorously evaluated, and the study documenting that has not itself been followed by a systematic re-examination.

    Innovation agencies have not become less common in the years since. They appear across country analysis through 2026, and regional reviews of productive development policy continue to treat them as a standard instrument. What the record does not show, on the evidence available to us, is a subsequent effort to establish whether the fourteen agencies studied, or the many not studied, produced effects that would not otherwise have occurred. The recommendation the study made, that agencies build internal monitoring and evaluation capability so that ex-post evaluation becomes possible, was made because that capability was largely absent. Whether it has been built is not something the record currently answers.

    06The Builder's Reading

    1. Design the data collection before the program launches. Ex-post evaluation is only possible where collection systems existed at the outset. After launch, the window has closed.
    2. Ban unfalsifiable attribution. A share-of-GDP contribution claim with no counterfactual should not survive internal review, however favourable it sounds.
    3. Separate the output count from the impact claim. Report beneficiaries as reach and label them as such. The confusion is caused by putting them under an impact heading.
    4. Build internal evaluation capability. The study's own recommendation, and it implies a standing function rather than a periodic consultancy.
    5. Treat diagnostic-based intervention as the design test. Which constraint is this program addressing, and what evidence says it binds? An agency that cannot answer is choosing instruments by availability.

    The uncomfortable reading is what this implies about the wider advisory literature. Recommendations about innovation agencies circulate widely, are cited in national strategies, and shape institutions that spend public money for decades. This study suggests much of that guidance rests on operational description and reported outputs rather than measured effects, because measured effects were mostly never produced. That is not an argument against innovation agencies. It is an argument for knowing which parts of the received wisdom have been tested and which have only been repeated.

    All agency figures, the evaluation finding and the seven building blocks are as reported in the 2019 comparative study of fourteen innovation agencies. Reported outputs are reproduced as recorded and are not independently verified here. The characterization of the 5 percent economic output claim as unfalsifiable refers to the absence of a stated counterfactual in the source, not to any assessment of the agency's actual contribution. No IEPA engine outputs are used in this piece.