Are We Guessing Well or Knowing Well? Evidence Gaps in R&D Institutional Design
Decisions made based on instinct, analogy, and a strong theory of change have a way of becoming retroactively reframed as evidence-based once they work out. That’s part of why “well-reasoned” and “evidence-based” are often used interchangeably in policy conversations – and why it can be hard to tell, looking back, which parts of our policies, programs, and institutions were actually designed on rigorous grounds and which just got lucky. The distinction becomes more important than ever in moments of disruption; design features that were never truly grounded in evidence can become precedent, and underexamined ones get a pass because they haven’t failed in a visible way.
This isn’t an abstract worry. In early July, FAS convened a group of experts for a design exercise, imagining the federal R&D institutions they would build if they could start from scratch. The exercise was designed to help answer the questions: Are the institutions we have today fit for purpose? If not, do we know enough to redesign them well? We found that the answer to both of these questions is no – not because good ideas are scarce, but because the evidence base underneath institutional design isn’t quite as mature as most confident policy proposals assume. The design of federal R&D institutions as they exist today is part inherited guesswork and part accumulated “stop energy” , layered so many times over that it’s impossible to remember what we’ve rigorously tested (if anything). As we look toward the future, we have a chance to design better – here’s how.
Where are the gaps in our current evidence base?
We’ve made strides in understanding what our R&D institutions are producing: patents, citations, journal articles, and more. But hitting every metric can happen by chance, and doesn’t guarantee impact.
Our existing evidence base has a few structural problems:
Bias in the metrics themselves. The limited set of metrics that the current literature on the ROI of R&D relies upon show the survivors, not the failures. We know a lot about the DARPA programs that produced GPS and the internet; we know much less about the DARPA bets that didn’t pan out; the evidence is even thinner on other agencies that tried a similar model. These metrics are also limited in what they can tell you about impact – some research might not lead to a high-impact journal article, but might help a specific community solve a big problem or provide training for early-career researchers, and that might be just as valuable.
We’ve written much more about the types of questions the research community should invest in answering, and the outcomes they should seek to measure, to have the type of robust understanding of the impact of R&D spending decision-makers need today. Without a clearer picture of the impact of our current institutions’ approaches – i.e., an understanding of what works and under what conditions – we run the risk of replicating approaches that won’t drive impact.
Findings may not generalize. Differences between design models clearly matter — the portfolio-driven approach of ARPA models produces different outputs than investigator-initiated funding’s broader model. And different agencies that have implemented ARPAs have employed different approaches to their talent model, decision-making, and timelines for funding. But knowing that the models differ isn’t the same as knowing, rigorously, which conditions produce breakthroughs versus reliable, incremental progress. A mechanism that works in one scientific domain, funding environment, or political moment doesn’t automatically transfer to another.
A limited comparative infrastructure. Institutional design knowledge is strikingly under-documented, with few good manuals and little (but not zero!) cross-case comparison. Much of the literature on key ingredients in institutional design, like hiring authorities, describes the available options or relies on anecdotal evidence rather than rigorous evaluation of the impact of various approaches. Toolkits and templates have started to emerge, such as The Institutional Architecture Lab’s guides for new and renewing institutions and the Overedge Catalog of innovative research organizations. But the conversation at the workshop showed us that these resources are still more limited than is needed, and the resources that do exist aren’t yet relied upon by those who are imagining the future of R&D institutions. If the opportunity to design a new R&D institution emerged today, it would be designed largely from analogy and instinct, not from an accumulated, testable body of knowledge.
The potential implication of these evidence gaps becomes more consequential when you look at how institutions actually make decisions day to day. For example: in one of the workshop’s exercises, participants found that funding instruments usually get locked in by habit, not fit – program managers default to whatever mechanism they’ve used before, and there’s rarely real capacity to reassess mid-course. The gap isn’t just a missing tool. It’s a missing muscle for choosing, and re-choosing, deliberately. That’s just one example, but it’s exactly the kind of decision that’s usually defended after the fact as evidence-based when it works out, when it was really just default behavior dressed up in retrospective justification.
What would actually closing the gap require?
The answer isn’t to slow everything down until perfect evidence arrives — that evidence may never fully arrive.
But there’s a few concrete moves that can be made starting today:
- Document design logic as it happens, not after the fact. The same legibility gap that leaves institutional design under-documented also means hard-won lessons from any given redesign effort risk being lost before the next one starts. Capturing the reasoning behind decisions – not just the decisions themselves – should be its own deliverable. New technological advances make it easier than ever before to continuously synthesize data and evidence as it emerges, which could give institutional founders a greater understanding of what design decisions have been made in the past, why those decisions were made, and the impact they led to. That data and evidence should be a key input into institutional design conversations, such as agency reauthorization.
- Build the comparative infrastructure that future institutional founders will need. Simulation, cross-case comparison, and AI-assisted tools could let designers stress-test structures proposed designs before committing to them, rather than relying on analogy to a handful of famous examples. For example, the Institute for Progress’s Atlas of Innovation is a new tool that could offer needed insights on how different funding instruments work in practice, or new approaches like digital twin technology could help simulate the impact of different design decisions.
- Design for evaluability and course-correction from day one. Rather than assuming the founding design is correct and static, new institutions have a rare opportunity to build in real feedback loops and the flexibility to course-correct, treating the initial design as a hypothesis rather than a final answer. In the R&D context, proposed metascience units are a promising way to tackle this – but if, and only if, mechanisms are constructed to ensure that the evidence they generate has a meaningful impact on decision making.
Temper confidence, don’t abandon action
None of this is an argument for paralysis; Americans deserve institutions that produce high-quality R&D. It is an argument for being honest about the difference between “well-reasoned” and “rigorously evidenced”.
Some tensions in institutional design – autonomy versus accountability, curiosity versus utility, efficiency versus equity – aren’t problems to solve, just permanent features to navigate. The evidence gap is like that too. It’s not fully going away before the next round of decisions gets made. The best available response isn’t to wait for certainty that may never come, but to be honest about how much of the confidence behind any given proposal is actually earned – and to build institutions flexible enough to find out they were wrong.
Every new institution inherits something – existing incentives, existing staff, decades of prior investment that shaped what’s possible.
The digital government field has an opportunity to build a more responsive and resilient government by pushing into new frontiers, with new tools, approaches, and even organizations that don’t exist yet. This is the time for radical experimentation, delivery, and exploration.
This is a tremendous opportunity to redefine what people expect from government, and in doing so, inspire cities across the country to raise their own ambitions. We are excited to see this initiative lead the way and look forward to cheering your success.
Let’s see what rules we can rewrite and beliefs we can reset: a few digital service sacred cows are long overdue to be put out to pasture.