What Survives the Trip From Blue-Sky to the Real World
No institution gets built on a clean sheet. Every new institution inherits something – existing incentives, existing staff, decades of prior investment that shaped what’s possible. A new Advanced Research Projects Agency (ARPA)-style agency has a foundation partly because of the Defense Advanced Research Projects Agency (DARPA) model that came before it and the foundational research funded by more traditional R&D institutions that makes its agenda possible. A coordinating body proposed today would inherit the turf, budgets, and statutory authority of whatever agencies already touch that problem. A darker version of this premise is that while market incentives may clear out failed models in the private sector, in government very little forces that sort of reckoning, so institutions built for a 1960s problem are still standing when problems of the 2020s show up, wearing their old habits and assumptions.
Federal R&D faces a particularly stark version of this challenge. The basic architecture of how the federal government funds and fosters science, the peer review system, the relationship between agencies and universities or labs – all of those have roots in Vannevar Bush’s post-war vision. This (at the time) radical and successful system was not built for some of today’s disruptions, like AI acceleration of research pipelines, the changing role of the private sector, vastly different legislative-executive roles, and an increasingly tangled federal coordination architecture. In the 80 years since, we’ve layered, worked around, bloated, and streamlined this inheritance, but have not really made an effort to start from scratch.
That’s what makes blue-sky thinking exercises valuable. Strip away the inheritance of the past for a day and people can make unencumbered choices on decisionmaking, feedback loops, talent models, innovation, coordination, resource access and allocation, and more. Most institutions never get a chance to make these choices on purpose. Left alone, a bureaucracy will choose stability over exploring renewal almost every time, whether or not the status quo is working, because nothing inside the system forces the other choice.
FAS ran such an exercise in July, pulling together founders, government-capacity practitioners, academics, and metascience researchers in groups that each imagined a new R&D institution from scratch around a specific real-world problem. They made concrete decisions on governance, instruments, talent, and learning rather than describing aspirational values, pulling together institutions focused on topics from housing affordability or a federal grand-challenges portfolio.
The exercise itself was not the hard part (or at least, not the hardest part). Give a group of smart, experienced people a blank sheet and a real problem, and they’ll design something better than what exists today. The real design challenge starts the moment you try to put ideas back into a world full of prior claims, getting any of the great ideas generated into a bill, an agency reorganization, or a budget request. From this exercise, we think there may be strong candidates worth breaking the hold of prior models: seven groups worked through seven unrelated problems with no way to compare notes, and a few of the same decisions still showed up across the designs. As we consider future R&D institutional design opportunities (from NSF reauthorization to relevant planning for future presidential administrations), these are worth taking on the trip from blue-sky to reality.
What we learned from the designs
We asked groups to surface what’s broken today and the critical ingredients in future R&D institutions: instruments, credibility systems, public engagement, talent models, founding decisions. They took those diagnoses, which highlighted challenges in agility, balancing tradeoffs, superficial public engagement, lax accountability, and used a basic rubric to build something new and aspirational. Here are the themes that showed up again and again across groups tackling different challenges.
- R&D coordination problems don’t get solved by convening power alone – they get solved by convening power PLUS authority.
This was the most consistent thread across institutional designs, building on a diagnosis that critiqued poor coordination among agencies with overlapping science missions and emphasis on designing institutions around mechanism rather than the mission or market failure they are meant to solve. One group built their entire proposal around granting real coordinating authority to hold other agencies accountable; their design was explicitly based on the premise that no single existing agency owns “grand challenge” R&D problems, and the agencies that touch them are too siloed to solve this on their own. Another group reached the same conclusion from a different angle, proposing to centralize innovation services for federal R&D that are currently fragmented across delegation and coordination, backed by continual dedicated funding and maximum hiring authority. A third group wanted to build a “new credibility/adaptation layer” and described its function as closing feedback loops between the agencies that fund research and the agencies that would act on its findings.
What each of those groups said, with a slightly different twist, is that when tackling cross-agency alignment and coordination, the missing ingredient wasn’t a body that could bring people to the table, but one that could make a decision stick once it got there. That’s a meaningfully different task than “more coordination”, which is often invoked for purposes of transparency or as advisory conversations to tee up recommendations. Groups wanted to imagine what it looked like when there is real authority to force choices, hold other agencies accountable for outcomes, or centralize functions that are currently scattered across delegation and hand-offs – the kind of authority that can survive a disagreement.
Notably, none of these models turned to existing center-of-government bodies (or other intra or interagency models) to take this role. Andrew Greenway highlights this feature in his work on creative destruction, recognizing that task-force style bodies in the UK are most successful, deliberately built outside the host agency’s normal chain so they couldn’t be smothered by its process, then given real authority to bind decisions rather than just recommend them. Writing with Jen Pahlka, Greenway also notes that before the Consumer Financial Protection Bureau (CFPB) existed, seven separate agencies nominally held pieces of consumer protection, and none had it as a primary mission; consolidating that fragmented authority into one body with a clear mandate let it return more than $21 billion to consumers.
One useful exception: sometimes “coordination” can mean linking incentives, not designating governance through a single authority over everyone else. One group split the R&D pipeline into three separate entities – discovery, application, and commercialization – linked by shared governance and co-ownership rather than a hierarchy. In their design, profit generated from downstream success flows back to whichever entity originated the science, so nobody has a reason to let a good result die at the handoff between stages. This offered a different answer to the same problem the authority-based designs solved: instead of one body with the power to make a decision stick, this design pointed all of the entities’ interests in the same direction so that hierarchy wasn’t the only thing holding it together.
Takeaway for the trip: Goodwill and convenings are ingredients but insufficient for effective coordination of federal R&D power; what experts reached for was authority to set direction and make it stick, or incentives to make cooperation the default. Hiya Jain, Stuart Buck, and Aishwarya Khandija make an intersecting point in their new essay on taxonomies of R&D organizations: two organizations can be identical on funding, timelines, and who they’re serving, and still produce completely different results depending on nothing but how decisions get made.
- Instruments should be matched to where a problem sits in the R&D pipeline, not a single default tool.
Today, R&D funding instruments are often set by habit and comfort, rather than fit – a key diagnostic of this exercise – with little inside-government capacity to stipulate what works or to propose shifts in approach. These break down in predictable ways when science moves faster than the plan anticipated, like cooperative agreements that don’t fit commercialization or milestone-based contracts that break down under R&D uncertainty, driving constant contract modifications to keep up. To address these limits, one design treated the choice of funding mechanism as something that should shift as a project matures — grants for early discovery, more flexible investment vehicles once there’s something to build on, a different structure again once the goal is commercialization or deployment — rather than picking one instrument and forcing every stage of a problem through it. Their approach overcomes the square-peg-round-hole instrument challenge by letting their institution ask “which tool fits this specific point in the process,” with the handoffs between stages treated as seriously as the stages themselves.
Most of the other designs didn’t formally stage instruments this way, but still resisted restricting themselves to a single funding tool. Their logic was that the tool should follow the problem rather than the other way around – for example, building a portfolio that explicitly seeks to make investments on different time horizons or that uses different funding instruments depending on the purpose and end-user of the research.
Takeaway for the trip: an instrument chosen out of habit isn’t neutral, it’s a bet that this problem looks like the last one. The Institute for Progress’s Atlas of Innovation is a great start to this challenge, routing funders to one of thirteen mechanisms by answering three questions: how well-defined the problem is, what a solution would look like, and which team is positioned to build it?
- Public input shouldn’t just be a feedback mechanism; it should be a design input, too.
The diagnosis of current public input mechanisms into R&D was a vociferous failing grade; from notice and comment, to community advisory bodies, to efforts at transparency – most were seen as well-intentioned box checking rather than efforts for legitimate engagement. The workshop repeatedly came back to the theme that improving legitimacy and effectiveness in the eyes of the public is vital – and currently broken. Effective engagement lives upstream of today’s formal input mechanisms, both as a means to effectively inform public science and to increase the legitimacy of public investments. If current mechanisms were once well-intentioned, they’ve suffered the fate of many other institutional innovations, calcifying into a procedural or compliance check rather than a value.
As a consequence, the more interesting designs didn’t propose better PR or storytelling that explains funding decisions to the public after they were made. Instead, they asked how to let public and community input shape which problems agencies should tackle and how research got framed in the first place. That’s meaningfully different from most of our current approaches that embed and elevate the same voices at a point too late for useful input. For example, one group explicitly said the biggest departure in their design from what R&D institutions have today is “giving community feedback real influence over problem choice and research design”; another proposed a review process where individual project choices rely on both local/community impact and scientific expertise.
Takeaway for the trip: public input isn’t neutral: without deliberate design, the loudest or best-resourced voices, such as business interests, might crowd out everyone else, so the mechanism itself has to be built to ensure it gathers input from a representative sample.
- Talent is the design variable that does the most work – and the challenge is far more sophisticated than hiring fast.
“Nobody is in charge of innovation and anyone who claims to be probably isn’t.”This sharp diagnostic comment from one participant highlighted the root of the government’s R&D talent problem. Innovation comes from motivated, curious people incentivized to learn in a system that roots for them, not from an organizational chart, and federal agencies have struggled to find or support that kind of person.
The recurring throughline in designs wasn’t simply “hire good people;” it was a specific kind of comfort with ceding control and creating systems that let them thrive: leaders willing to hire people who know more than they do and are empowered to act decisively and creatively. Several groups proposed a model where program managers were empowered to move across agency lines (so-called “boundary spanners”) rather than staying in a lane; others proposed hiring staff whose job is explicitly to translate between worlds (science and market, or science and community) rather than sit inside just one of them. Pahlka and Greenway make a related point about boundary-spanners: the strongest founding teams pair insiders who chafe against institutional norms with outsiders drawn by the mission, since too many career officials will stifle disruption and too many people who don’t understand how bureaucracy works can be just as fatal.
Hiring the right people (and the appropriate mechanism to do so) isn’t the only key concern. This sort of talent is both rare and challenging to entice into overly bureaucratic roles, necessitating not just rethinking of role definition but growth, cross-training, and incentives. Continuous learning for program managers and other key staff to ensure they don’t become entrenched in the approaches that existed when they arrived is just as important. Superstar talent will struggle to stay engaged on a junior varsity field – DARPA’s reverence for program managers only works because the agency can act fast around them, not wasting a talent recruit on a six month hiring process or multi-month decisions processes.
Ultimately, this means caring about talent in ways government has struggled with, particularly at R&D institutions: investing in people willing to take risks, in the people who build the infrastructure that lets them take those risks for public impact, and in the recruitment infrastructure that finds both.
Takeaway for the trip: government keeps trying to solve challenges related to innovation talent as a hiring problem when it’s actually three problems stacked on top of each other, whether the right person exists, whether the system is built for them to succeed, and whether anyone was looking for them in the first place. Fixing only the first one is why so many great hires become average or former.
- Political durability should be scoped and designed for up front.
Publicly funded science gets held to short-term political scrutiny that other government money escapes. Political cycles move faster than science can absorb, which means priorities swing before results are even in, and the accountability hits hardest on exactly the kind of research that takes the longest to pay off.
R&D institutions get a similar problem: new institutions with new leaders or missions get more carve-outs, room to maneuver, and the halo effect from overseers; older institutions are stuck with residual procedural barriers and longstanding scrutiny.
Rather than treating the risk of political turnover, earmarks, or administration changes as an implementation problem to worry about after the institution existed, several groups treated it as a first-order design constraint. That shows up as a deliberate push for insulation and independent authority in some designs, and as open, unresolved questions about whether existing structures (the National Science Board, existing agency directorates) are the right vehicle for that durability in others. For example, one group proposed a coordinating body where governance would be led by career staff from various R&D agencies rather than political appointees.
Takeaway for the trip: Political change and risk isn’t a threat to manage, it’s a design spec; realistic considerations around insulation, governance, and succession need answers before the institution launches.
- The most common “biggest departure” from what we have today wasn’t a new tool, it was shifting power.
Ask each group what its design changed most from how things work today, and the answers split by which direction power moved. Three groups pushed power down, toward the communities and end-users closest to a problem: real influence over problem choice, devolved leadership, feedback loops built into the structure instead of added on. Two more pushed it up, into a coordinating authority that doesn’t currently exist and would need real teeth over agencies that already do.
The split tracks back to two different diagnoses from the day before. The groups pushing power down were answering the public-engagement critique directly: existing mechanisms, comment periods, advisory boards, oversight, all operate downstream of the decisions that actually matter. The groups pushing power up were answering the coordination critique: agencies that touch the same science don’t coordinate well, and nobody currently has the authority to make a cross-agency call stick.
It would be easy to treat these patterns as interesting but premature – worth revisiting once the right big legislative vehicle comes along. Founding moments are the point of maximum leverage in an institution’s life: decisions about mission, succession, culture, and evaluation made at the outset shape an institution for decades, yet rarely get made consciously. But waiting for a hypothetical clean-slate moment misses the actual openings that exist right now, in whatever form they take – and by the time a bigger opportunity arrives, the important decisions made at a smaller one will already be locked in, for better or worse.
“Right now” doesn’t have to mean a brand new institution. Pahlka and Greenway lay out a real range of options between doing nothing and starting from scratch. Universal Credit’s turnaround happened entirely inside the UK’s existing welfare department, with new leadership and a new mandate. Colorado’s digital service took a middle path, planting a new team inside a broken agency until the new way of working displaced the old one outright. New institutions like NASA or DARPA are the right call when no existing agency could plausibly take on the job, but they’re one option on a spectrum, not the only lever available to whoever’s holding a smaller opening right now.
Takeaway for the trip: before you design anything, consider which failure modes scare you; and match your ambition to the transformational opening you actually have.
Some tensions aren’t problems to solve. They’re just what you navigate.
Four tensions sit underneath almost any institutional design, and none of them get solved, only managed. Curiosity versus utility: research pursued because it’s interesting, versus research aimed at an immediate need. Political accountability versus scientific autonomy: the demand for relevance against the independence that lets evidence lead somewhere inconvenient. Efficiency versus equity: doing the most with limited money against reaching whoever needs the work most, not just whoever’s easiest to serve. And an organization’s stated mission versus what actually motivates the people inside it, which drift apart more often than anyone wants to admit.
The toughest version of all four shows up at the budget level: full-cost accountability versus flexibility. You can have rigorous, fully-costed oversight, or you can have room to move fast and change course; wanting both in full is how institutions end up with neither.
Groups refined these tradeoffs to understand what will always be a challenge; they’re not a checklist to clear before a design counts as finished. The question isn’t how to solve the tension but which side you’re choosing on purpose, and whether the people accountable for the institution know that’s the choice being made.
Where do we go from here?
The type of overhaul that results in a brand-new science institution may not be on the table anytime soon, but plenty of the ideas generated by groups are still ripe for implementation today.
Here’s where we can meaningfully learn from their designs to improve our existing institutions in the near-term:
Increase the capacity of existing coordinating bodies to effect change. The number of groups that proposed new forms of convening and coordination across R&D agencies is a clear diagnosis that our existing bodies, such as the National Science and Technology Council (NSTC), aren’t working. And frankly, it’s not their fault: the Congressional Research Service recommended in 2023 that Congress meaningfully interrogate the resource constraints on the Office of Science and Technology Policy. Ensuring that existing policy coordination committees are appropriately resourced and have the capability to provide actionable recommendations to agencies might solve some of the problems that groups pointed to without requiring a full overhaul.
Experiment with new funding approaches. The recently announced X-Labs program at NSF is an important proof point that agencies can try out new ways to fund research that are explicitly designed to be fit-for-purpose. This is an opportunity to try something different and learn what is working in the process. Other R&D agencies should think just as creatively – for example, the Department of Energy underutilizes its Other Transactions (OT) Authority; an updated approach at DOE could expand its use when the flexibility of OTs is a better match than traditional grantmaking strategies, aligned to workshop groups’ recommendations around instrument mix.
Incorporate more public input. Grantmaking agencies can adopt approaches that allow a broader set of voices to influence grantmaking strategy. This can take a number of forms – from making requests for information about new research programs more accessible to new types of research awards. One compelling strategy to learn from is the Institute of Education Sciences’ recent announcement of changes to how it supports research “Network Hubs”; while these have historically served as ways to connect researchers, they will now be asked to conduct activities that allow them to better understand Americans’ research needs.
Leverage live policy windows. While agencies can do plenty without a legislative overhaul, there are some upcoming opportunities to enact some of the more dramatic changes groups pointed to – in particular, the National Science Foundation’s current authorizing legislation will expire next year. This is a unique opportunity to holistically consider whether NSF’s current design is well-aligned to its three-prong purpose.
As legislators consider priorities for reauthorization, there are a few specific design areas they should explore:
- Recommitment to Catalyzing Foundational Research: Aligning on the purpose for NSF, with a primary focus on building a foundation, but not constraining the agency to grant-only funding mechanisms and preserving the Directorate for Technology, Innovation, and Partnership (TIP)’s focus on accelerating innovation.
- STEM Workforce Mandate: NSF could take a more strategic approach to its role in building the STEM workforce of tomorrow. For example, it could house a workforce intelligence program that helps explore the specific needs of emerging S&T fields. Building upon existing approaches like the Rotator program, NSF might also adopt more use of personnel mechanisms that allow them to bring in specific expertise when it’s needed, including some of the more translational roles that many groups proposed in their designs.
- Agility and Flexibility: This might include devolved authority over grantmaking, beginning or expanding programs that give program officers greater discretion in awarding funding, similar to ARPA models or “fast grants” that accelerate the pace of grantmaking. Additionally, NSF reauthorization could expand Other Transactions authority to extend agency-wide beyond a single directorate (currently, the Technology, Innovation, Partnerships directorate) as a mechanism for enhanced flexibility.
Fix what accountability truly measures. R&D spending gets judged by a stricter and stranger standard than most government money. It faces short-term political scrutiny that other spending largely escapes, on timelines that don’t match how long real research takes to pay off, and is tracked more by geography rather than by outcome or impact, so even when something is working, there’s no clean way to show it. The fix isn’t asking science to move faster, but instead evaluating it on a timeline and with measures that match how the work happens in reality, and tracking it by what it produces and the impact that work can have.
What’s haunting these designs but never gets named directly: peer review. The workshop’s diagnosis flagged that credibility systems are built for accountability and transparency, but poorly equipped to reward agility or value negative results. This resurfaced almost as an aside in the workshop’s conversations, where one group noted flatly that traditional peer review “drives funding decisions toward safe, incremental research” and needs to reorient toward risk management instead. No group built a fix into an actual institution. But look back at what the other recommendations require: instruments that shift with a problem’s stage, talent willing to take real risks, coordinating bodies with authority to back ambitious bets. All of it still runs through a review process built to reward a safest version. Fixing peer review isn’t a fifth item on this list so much as a thing that determines whether any of the others work.
What doesn’t survive the trip?
Not everything blue-sky thinking produces will translate cleanly. The same groups that worried about political resilience were also, implicitly, wrestling with a harder tension: the least-friction option in institutional design tends to produce the worst result, and building real coordinating authority is inherently high-friction. It means taking on existing agencies’ turf, existing appropriators’ preferences, and existing staff’s incentives. A design that looks clean on a whiteboard because it assumes away that friction probably isn’t quite ready for a policy window. It is impossible to resolve all tradeoffs in shaping and incentivizing public science and anyone who tells you otherwise is selling something. Tough choices will always need to be made.
The value of an exercise like this was never the designs themselves (though they are impressively detailed for the hour we gave groups to come up with them). It was building a sharper, shared vocabulary for what well-designed R&D institutions actually require – coordination with real authority, the right talent and leaders, feedback loops built in from the start – and testing that vocabulary against problems specific enough to force real tradeoffs.
The next step is doing the harder work of testing those same patterns against whatever windows are open, one at a time, wherever they show up – a reauthorization law due for renewal, a new agency proposal with room to build these patterns in from day one, or simply an existing agency choosing, when its own authority makes it possible, to pursue new funding strategies and build in real feedback loops without waiting for a bill at all.
At a period where the federal government is undergoing significant changes in how it hires, buys, collects and organizes data, and delivers, deeper exploration of trust in these facets as worthwhile.
What if low trust was not a given? Or, said another way: what if we had the power to improve trust in government – what would that world look like?
A proposal to build trust in the health data ecosystem.
With thoughtful policy action, it is still possible to build systems that are fair, transparent, and accountable, and to earn the public trust that will ultimately determine AI’s future. We hope policymakers are ready to act.