A Focused Research Organization to Measure Complete Neuronal Input-Output Functions
Measuring how neurons integrate their inputs and respond to them is key to understanding the impressive and complex behavior of humans and animals. However, a complete measurement of neuronal Input-Output Functions (IOFs) has not been achieved in any animal. Undertaking the complete measurement of IOFs in the model system C. elegans could refine critical methods and discover principles that will generalize across neuroscience.
Systems neuroscience aims to understand the complex interplay of neurons in the brain, which enables the impressive behaviors of animals. A critical component of this understanding is the Input-Output Function (IOF) of neurons, which characterizes how a neuron integrates its inputs and responds to them. However, despite its importance, a complete measurement of IOFs in any animal has not been achieved, creating a significant blind spot in our understanding of brain function. While parts of IOFs have been measured through various experiments, these efforts only control a narrow subset of the inputs to any given neuron, providing only a small slice of the true IOF. Furthermore, the output of neurons is a complex nonlinear function of their inputs, adding to the complexity of the problem. To truly understand the computation and function of the brain, we need a detailed functioning of the IOFs, which requires controlling all of the inputs and observing the factors that shape IOFs. Pursuing this in a single animal model would uncover new methods, tools, and scientific principles that could catalyze large-scale innovation across neuroscience.
Project Concept
The project aims to unlock a deeper understanding of how neurons in the brain process information through the complete mapping of neuronal IOFs using the model organism C. elegans. This comprehensive mapping is a vital step in understanding how the brain’s neurons receive, integrate, and respond to signals. The project will use a combination of advanced techniques including optogenetics, modern microscopy, and microfluidics to control and observe the nervous system of the worms, and will develop models to predict how neurons respond based on their inputs. The team will also explore how different factors, such as chemicals, drugs, and non-neural cells influence these responses. To ensure maximal benefit to the field and the scientific community, all data and findings will be shared openly and resources will be allocated to promoting collaboration with outside experts on experimental design, technology development, and computational and theoretical analysis.
What is a Focused Research Organization?
Focused Research Organizations (FROs) are time-limited mission-focused research teams organized like a startup to tackle a specific mid-scale science or technology challenge. FRO projects seek to produce transformative new tools, technologies, processes, or datasets that serve as public goods, creating new capabilities for the research community with the goal of accelerating scientific and technological progress more broadly. Crucially, FRO projects are those that often fall between the cracks left by existing research funding sources due to conflicting incentives, processes, mission, or culture. There are likely a large range of project concepts for which agencies could leverage FRO-style entities to achieve their mission and advance scientific progress.
This project is suited for a FRO-style approach because it requires a level of coordinated development and engineering that is too big for a single academic lab, too complex for a loose multi-lab collaboration, and not directly profitable enough for a venture-backed startup or industrial R&D project. The project also aims to create a suite of public goods through its commitment to open science, with plans to share data and code as they are developed. More broadly, the work lends itself well to the development of a new set of tools and methodologies rather than the products or papers incentivized by traditional research models.
How This Project Will Benefit Scientific Progress
This project aims to revolutionize neuroscience by providing the first complete measurement of neuronal Input-Output Functions (IOFs) in the model organism C. elegans. Identifying causal interactions between neurons will provide unprecedented insight into brain function, and the methods developed in the process pave the way for understanding more complex nervous systems.
Key Contacts
Authors
- Konrad Kording, University of Pennsylvania, koerding@gmail.com
Referrers
- Jordan Dworkin, Federation of American Scientists, jdworkin@fas.org
A Focused Research Organization to Develop a Modular and Scalable Platform for Human Molecular Monitoring
Wearable health electronics are now ubiquitous, but continuous molecular monitoring is only widely available for glucose. Decades of research have expanded continuous monitoring to other molecules, but these techniques are restricted to research labs and remain disconnected from daily human use. We propose a platform to translate and distribute these emerging technologies, enabling the mapping of the time-varying human metabolome and the design of closed-loop devices for personalized health.
Humans are the best model organisms for humans, yet we have few tools to study human biochemistry in situ and in real time. The Human Metabolome Database lists ~20,000 detected compounds, of which only ~3,000 have been quantified. Even fewer of these biomolecules have been studied with time resolution in longitudinal human studies.
Cardiovascular, metabolic/endocrine and drug pharmacokinetic phenomena are driven by biomolecules varying on the time scale of seconds to minutes to hours, impacting behavior and well-being on similar time scales. Currently, the only way to measure these molecules is through laboratory testing, which is obtrusive to daily life. Moreover, laboratory testing cannot be conducted at the frequency necessary to capture all of these variations.
In contrast, wearable monitors are user-friendly and enable continuous measurements at the correct time scale. These monitors require specially engineered biosensors, since there are a limited number of naturally-occurring enzymes that generate continuous, time-varying electrical signals like those used by continuous glucose monitors, which are currently the only commercially available device of this kind. However, most labs that pioneer biosensing strategies do not develop human-compatible devices and vice versa, creating a chasm between these two areas of research and development.
Project Concept
This project aims to achieve minimally invasive continuous monitoring of 100+ analytes in the human body and deliver devices to researchers. This project will develop a medical device testbed that can use the myriad biosensing strategies pursued by academic labs to develop devices for human experiments. Our technical approach combines synthetic ion channels and conformational switches coupled with a CMOS array to assess many analytes in parallel. We already have access to customized fabrication techniques, and the cost and barriers in designing and manufacturing proteins and silicon sensors are trending downwards.
The project will progress through four interdependent stages:
- Survey and prioritize metabolic, hormone, and immune targets that provide the greatest explanatory power for well-being. Select biosensor and transduction systems reported in the literature to interface with our platform.
- Develop an integrated circuit functionalized with biosensors for parallel multi-analyte sensing packaged in a wearable form factor. Test devices in humans to validate against conventional blood sample analyses.
- Specify and share validated devices in bulk to catalyze large-scale human field research.
- Curate a time-varying human metabolome.
The project will progress through four interdependent stages.
What is a Focused Research Organization?
Focused Research Organizations (FROs) are time-limited mission-focused research teams organized like a startup to tackle a specific mid-scale science or technology challenge. FRO projects seek to produce transformative new tools, technologies, processes, or datasets that serve as public goods, creating new capabilities for the research community with the goal of accelerating scientific and technological progress more broadly. Crucially, FRO projects are those that often fall between the cracks left by existing research funding sources due to conflicting incentives, processes, mission, or culture. There are likely a large range of project concepts for which agencies could leverage FRO-style entities to achieve their mission and advance scientific progress.
This project suits a FRO-style model because the four research stages require tight feedback loops and standardization and the medical device industry is not incentivized to pursue this kind of research. Private companies typically focus on a handful of molecules most relevant to diabetes care, taking advantage of proven biosensors and predictable insurance reimbursement. A stand-alone, non-profit institute is best suited to standardize and derisk experiments on molecular monitoring to catalyze the formation of a consortium of experimenters. Just as the nonprofit AddGene has standardized and democratized access to genetic material, we seek to develop the analog institution for medical devices.
How This Project Will Benefit Scientific Progress
Human physiology is currently a poorly explored, high-dimensional space, and scientific labs lack the tools to measure human biochemistry over time. The devices developed through this project will first help scientists study specific questions in domains such as disease etiology, human behavior, and drug discovery. Study validity will be enhanced by providing additional molecular time courses which will clarify the relationships between conditions, specific biomarkers, and interventions. Next, these studies will enable the development of a human molecular atlas, similar to the Human Metabolome Database but with time resolution. This database could act as a powerful tool for developing nuanced models of human physiology to uncover previously overlooked phenomena. Finally, similar devices will ultimately become accessible as consumer health products, enabling the next generation of personalized health.
Key Contacts
Authors
- Anand K. Muthusamy (amuthusa@caltech.edu) and
- David Garrett (dgarrett@caltech.edu), Caltech
Referrers
- Alice Wu, Federation of American Scientists, awu@fas.org
Learn more about FROs, and see our full library of FRO project proposals here.
The overarching strategy involves converting biochemical interactions into a time-varying electrical signal and isolating specific interactions into separate channels on a CMOS chip. We envision a universal platform that can accept the myriad of biosensors from labs worldwide. Advances in synthetic biology and protein engineering have unlocked various modes of biosensing. Oxidoreductases, like glucose oxidase used in CGMs, provide a direct conversion of the local ligand concentration to an electrical current. Recent advances in deep learning protein engineering have shown proof of concept for the de novo design of enzymes, but some classes of biomolecules such as nucleic acids and proteins would be out of reach for enzymatic detection. In turn, we can either mimic nature’s ion channels or create our own electronic transduction systems. Aptamers are amenable to high-throughput evolution to bind a variety of targets and provide a hairpin motion that can move redox probes to and away from an electrode surface (cf. example of a cortisol sensor). Recent advances in DNA origami have enabled synthetic, self-assembling nanopores whose current can be modulated by the binding of other biomolecules. Both biological and solid-state nanopores are increasingly used in commercial scientific equipment, providing a tailwind for medical devices.
All of these biosensor and transduction systems can be integrated into field-effect transistors (FETs). Even strategies involving membrane proteins can now be interfaced with transistors thanks to organic electrochemical transistors with supported lipid bilayers harboring ion channels and nanopores. Modern complementary metal oxide semiconductor (CMOS) processes allow us to parallelize many sensors onto a single device at a low per-unit cost. Each sensor consists of a FET sensing unit, where the presence of molecules of interest induces a surface potential in the FET gate. This modulates the current in the FET, which is then measured with read-out circuitry. Each sensor’s current is then digitized and used as an indicator of molecular concentrations. Since relatively few channels are required (~100s) compared to conventional CMOS devices, coarser fabrication processes could be used (e.g. 130 nm) which would reduce cost and time while still allowing for a compact active area (~mm2).
The CMOS device will be packaged in a wearable form factor miniaturized for on-body measurements using a battery and wireless data transmission to a smartphone (e.g. using Bluetooth low-energy). We will draw interstitial fluid via reverse iontophoresis to reach the CMOS chip sitting next to the skin surface. Since different analytes have different rates of variation and concentration ranges, we will model the appropriate temporal resolution and SNR we can achieve. This device form factor and reverse iontophoresis method has been validated for ~mM glucose measurements with ~min resolution. In contrast, hormones may be present at ~6 orders of magnitude lower concentrations but only vary over days-weeks. Using our CMOS configuration, different channels can be read out individually. To efficiently capture these changes, different sampling rates will be used to ensure we capture their peaks and troughs while minimizing energy consumption. For analytes with lower baseline concentrations, techniques like averaging (over multiple samples and CMOS channels) may be required to improve sensitivity. The mean absolute relative difference (MARD) will be computed against blood LC-MS ground truth.
We are in a position where our data science methods outpace quality datasets. For example, geometric deep learning is useful for mapping interactions in a metabolome, and long short-term memory is well-established for analyzing and predicting time series data. These data can become actionable in a closed-loop device where a biosensing event triggers an alert or pharmacological intervention as in an artificial pancreas. Here again, control theory approaches to physiology are sufficiently mature.
A Focused Research Organization to Systematically Study Bacteriophage Genes and their Functions
Systematically sequencing the genome and studying the function of genes from all viruses that infect a set of model bacteria with significant scientific, biotechnological, and human health relevance will enable the development of phage-gene libraries that can in turn enable the faster development of genetic tools for advancing molecular biology.
Problem Statement
Viruses have been evolving host-modifying factors for billions of years. This wealth of naturally engineered proteins holds the key to unlocking the full potential of the cell. Virus-derived genetic tools have driven most key advances in molecular biology, from recombinant DNA to CRISPR genetic engineering. Although most such transformative discoveries have resulted from the study of bacteriophages (viruses of bacteria), phage research has relied primarily on inferential work rather than systematic approaches to discern the functions of phage genes. With serendipity as the primary engine of discovery, experimental approaches have not kept pace with phage genome sequencing over the past decade. Consequently, the vast majority of phage genetic diversity is still entirely unexplored.
Project Concept
We have built a high throughput screening platform to characterize phage genes and completed a pilot of the entire pipeline, from gene selection through functional screening and mechanistic follow-up (manuscript in preparation and available upon request). Our FRO will scale this platform and use it to chemically synthesize and test all non-redundant phage genes from two clinically relevant families of bacteria (Enterobacteria and Mycobacteria) which collectively host ~40% of all isolated phages, allowing us to test a large swathe of phage genetic diversity in a set of model species. The phage-gene library generated from this process will enable us to pursue the following objectives: 1) discover new molecular tools with revolutionary potential (eg. broadly understanding the principles of protein detection in antiviral immunity could yield a generalizable protein-targeting framework without some of the pitfalls of antibodies), 2) develop therapeutic avenues for antimicrobial resistant infections inspired by natural antiviral defense and counter-defense strategies, 3) build an inventory of phage design principles and engineering methods for therapeutic, industrial, and microbiome-directed applications, and 4) gain a complete understanding of interactions between phage and their hosts.
What is a Focused Research Organization?
Focused Research Organizations (FROs) are time-limited mission-focused research teams organized like a startup to tackle a specific mid-scale science or technology challenge. FRO projects seek to produce transformative new tools, technologies, processes, or datasets that serve as public goods, creating new capabilities for the research community with the goal of accelerating scientific and technological progress more broadly. Crucially, FRO projects are those that often fall between the cracks left by existing research funding sources due to conflicting incentives, processes, mission, or culture. There are likely a large range of project concepts for which agencies could leverage FRO-style entities to achieve their mission and advance scientific progress.
This project is suited for a FRO-style approach because to achieve our scientific goals, we will need to scale our platform ~10,000-fold from 104-5 assays in the pilot to ~108-9 assays at the FRO. Massively parallelizing these assays will involve a highly systematic effort with a tightly coordinated and dedicated team, a substantial initial investment in gene-library synthesis and platform engineering, and long publishing timelines, which are qualities unsuitable for traditional grant funding. For these reasons, an FRO is the ideal (and probably the only viable) structure for this project.
How This Project Will Benefit Scientific Progress
Paradigm shifts in biology have often started with the humble bacteriophage. With 108-9 prospects across the oldest and most diverse host-pathogen interface in the biosphere, our FRO presents abundant opportunities for making impactful discoveries, and will pioneer a new field of functional metaviromics. Moreover, the phage-gene libraries we will create are analogous to small-molecule screening libraries, consisting of 104-105 phage-derived natural products that can be used to find potentiators or suppressors of any cellular stressor of interest. We expect these resources to enable discovery far beyond the scope and timeline of our FRO.
Key Contacts
Authors
- Sukrit Silas, University of California San Francisco,
Referrers
- Alice Wu, Federation of American Scientists, awu@fas.org
Learn more about FROs, and see our full library of FRO project proposals here.
A Focused Research Organization to Reduce Antibiotic Resistance In Aquaculture
Research and engineering to reverse antibiotic resistance in aquatic bacteria, through the application of a well-validated CRISPR-based genetic system, can help catalyze safer, more sustainable land-based aquaculture as a nutritious and affordable food source.
The growing human population needs affordable, healthy sources of protein. With overfishing putting severe pressure on global fish stocks, aquafarming presents a potential alternative. The U.S. currently imports about 80% of its seafood, and most imports are produced by foreign aquaculture; expanding domestic aquaculture could help to close the $17 billion seafood trade deficit. But domestic aquafarming poses its own challenges, including the potential for environmental contamination near ocean-based operations. In such scenarios, high concentrations of fish within netted areas lead to bacterial and other waste contamination spreading beyond the arena of fish confinement. The alternative strategy of raising fish in isolated inland enclosures may pose less environmental risk, but also requires maintenance of water quality, frequent water filtration and, often, the use of high antibiotic concentrations mitigate bacterial fish pathogens that thrive in such overcrowded conditions. In practice, aquafarmers often try to reduce the level of antibiotics added to the water in the last few weeks of fish growth to drop their concentrations below mandated health standards for commercial fish, but these efforts are only partly effective and create significant logistical burdens.
Project Concept
We proposed the development of genetic systems to reduce the prevalence of antibiotic resistance in land-based aquafarming enclosures. We will develop harmless strains of environmental bacteria capable of transferring self-copying genetic cassettes to pathogenic bacterial strains of concern in aquaculture. With these strains, we aim to reduce virulence of those bacterial pathogens in high-density fish enclosures and scrub their antibiotic resistance.
The heart of the project is to apply a well-validated self-amplifying genetic system, referred to as Prokaryotic-Active Genetics (Pro-AG), to the task of scrubbing virulence and antibiotic resistance factors from bacterial pathogens in aquaculture facilities. Since publication of the seminal study describing this CRISPR-based system for reversing antibiotic resistance (Valderrama et al., 2019, Nat. Comm. 10, 5726), we have further advanced the Pro-AG platform by combining it with means of spreading between bacteria through horizontal transfer systems such as conjugal transfer elements or bacteriophage. We have also incorporated new genetic features to the Pro-AG toolkit including a system to cleanly and efficiently delete genetic elements such as virulence factors responsible for antibiotic resistance. Building on these core achievements, we will transfer the Pro-AG framework and novel integrated phage-based systems to several bacterial strains of concern to aquaculture with the goal of diminishing their antibiotic resistance (AR) genes and virulence potential.
What is a Focused Research Organization?
Focused Research Organizations (FROs) are time-limited mission-focused research teams organized like a startup to tackle a specific mid-scale science or technology challenge. FRO projects seek to produce transformative new tools, technologies, processes, or datasets that serve as public goods, creating new capabilities for the research community with the goal of accelerating scientific and technological progress more broadly. Crucially, FRO projects are those that often fall between the cracks left by existing research funding sources due to conflicting incentives, processes, mission, or culture. There are likely a large range of project concepts for which agencies could leverage FRO-style entities to achieve their mission and advance scientific progress.
This project is suited for a FRO-style approach for three reasons. First, it would be very difficult to attract VC or industry funding for this effort. The expected timeline is too long for most VCs who want to see a shorter horizon on return for their investments (on the order of 2-3 years). Second, the project has significant technical risk since we do not know how the Pro-AG systems will perform in the context of large enclosures densely packed with fish, which is a daunting environment for any anti-microbial intervention. Third, the scale of just the laboratory component of the project exceeds the level of funding normally available through standard channels of support for academic science, since Pro-AG delivery systems would need to be engineered in parallel for several different species of fish pathogens. This will also require more “applied” work than is typically supported by many academic research programs. For these reasons, the project fits perfectly in the sweet spot for a FRO.
How This Project Will Benefit Scientific Progress
If successful, our systems would greatly reduce the necessary frequency and concentrations of antibiotics to control bacterial fish pathogens. Solving or attenuating this central challenge to land-based aquaculture should help foster safe, sustainable and affordable sources of nutritious, uncontaminated fresh fish and help catalyze a shift away from unsustainable overfishing practices in the open ocean and environmentally hazardous practices in ocean-based aquafarms. This project could also have broader knock-on effects by enabling similar technical advances to reduce antibiotic resistance prevalence in other environmental settings (e.g., livestock, sewage treatment), which are also substantial sources of worldwide antibiotic resistance.
Key Contacts
Author
- Dr. Ethan Bier, UCSD, ebier@ucsd.edu
Referrers
- Alice Wu, Federation of American Scientists, awu@fas.org.
Learn more about FROs, and see our full library of FRO project proposals here.
A Focused Research Organization to Design and Synthesize Spiroligomer Catalysts
This FRO will design and synthesize a library of spiroligomer enzyme-like catalysts to enable the development of new industrial processes for the production of green fuels and chemicals.
Humanity needs catalysts to create fuel, feedstocks to make materials, and fertilizers to grow food. Catalysts allow us to arrange atoms into the molecules we need with extremely high selectivity, cleanliness, and low energy input. Ever since Emil Fischer first conceived of the “lock and key” hypothesis of enzyme function, scientists have dreamed of rationally designing enzyme-like molecules. In 2021 the Nobel prize was awarded to Benjamin List and David MacMillan for developing organocatalysis – organic molecules that demonstrate basic catalytic function in enzymes. However, organocatalysts demonstrate only a small fraction (1/1,000,000,000) of the natural activity of the most capable enzymes because they are too small and do not display the deep, complex pockets of enzyme active sites needed to stabilize the transition states of reactions.
Project Concept
Spiroligomer nanostructures enable the creation of deep, complex, structured pockets that will allow us to design, assemble, and understand much more capable active sites. Using spiroligomer synthesis technology developed and scaled up over the last three years, this FRO will design, synthesize, and screen a library of spiroligomer catalysts. This will require the additional work of X-ray structure determination of active catalysts, measurement of activity using chemical kinetics, and computational modeling of active sites and reaction transition states.
What is a Focused Research Organization?
Focused Research Organizations (FROs) are time-limited mission-focused research teams organized like a startup to tackle a specific mid-scale science or technology challenge. FRO projects seek to produce transformative new tools, technologies, processes, or datasets that serve as public goods, creating new capabilities for the research community with the goal of accelerating scientific and technological progress more broadly. Crucially, FRO projects are those that often fall between the cracks left by existing research funding sources due to conflicting incentives, processes, mission, or culture. There are likely a large range of project concepts for which agencies could leverage FRO-style entities to achieve their mission and advance scientific progress.
This project is suited for a FRO-style approach because…Developing enzyme-like catalysts requires the engineering of highly structured molecules that are at least ten times larger than the kinds of molecules that synthetic chemists commonly create (5,000 Daltons for enzyme-like catalysts rather than 500 Daltons for typical small molecule therapeutics or organocatalysts). This will require a large engineering team with complex automation, instrumentation, and computation capabilities and professional synthetic chemists.
How This Project Will Benefit Scientific Progress
Unlike natural proteins that unfold and lose their activity when removed from their optimal temperature and solvent conditions, the spiroligomer-based catalysts we develop will be much more robust and valuable as industrial catalysis. Using synthetic chemistry, they can be produced at scale and made with much better quality control than natural proteins. It will also be easier to modify and tune their properties to create desired products. These catalysts display a much wider variety of chemically reactive groups than proteins do. This will open up the possibility of entirely new chemical production processes, such as artificial photosynthesis to create clean fuels and the production of other green chemicals and feedstocks.
Key Contacts
Author
- Christian Schafmeister, Temple University, meister@temple.edu
Referrers
- Alice Wu, Federation of American Scientists, awu@fas.org
Learn more about FROs, and see our full library of FRO project proposals here.
A Focused Research Organization to Develop RNA Sequencing Technologies
RNA therapeutics are gaining popularity since they are cheap and easy to make, but sequencing technologies today rely on converting RNA back to cDNA, which collapses information on the more than 150 different chemically modified bases for RNA into just four bases. To address this knowledge gap, this project aims to develop direct sequencing tools specifically for RNA including chemical modifications in order to enable the complete sequencing of RNA bases and improve RNA therapeutics.
RNA encodes regulatory information that directs cellular functions. This dynamic information is written in four canonical bases, each of which can be chemically modified to create more than 150 different bases. Today, sequences of RNA are mostly obtained by converting RNA back to cDNA which is then sequenced; in that process of reverse transcription, all the modifications on RNA are lost. This knowledge gap significantly weakens our understanding and treatment of human diseases.
Project Concept
As a Focused Research Organization, this project will recruit an interdisciplinary team to develop a low-cost tabletop device that can sequence RNA, including chemical modifications, as accurately as we sequence DNA, and make this product available to researchers at large. We propose to 1) improve high-resolution imaging technologies, 2) develop single-cell mass spectrometry for RNA sequencing, and 3) advance nanopore technology to enhance their accuracy and resolution. Ultimately, the imaging technology will be coupled with the mass spectrometry or nanopore technology to extract full-length RNA transcripts from subcellular locations and then sequence them. In parallel, we will develop algorithms to deconvolute data and build data displays to quantitatively reflect the locations of RNA transcripts and their sequences with all modifications.
What is a Focused Research Organization?
Focused Research Organizations (FROs) are time-limited mission-focused research teams organized like a startup to tackle a specific mid-scale science or technology challenge. FRO projects seek to produce transformative new tools, technologies, processes, or datasets that serve as public goods, creating new capabilities for the research community with the goal of accelerating scientific and technological progress more broadly. Crucially, FRO projects are those that often fall between the cracks left by existing research funding sources due to conflicting incentives, processes, mission, or culture. There are likely a large range of project concepts for which agencies could leverage FRO-style entities to achieve their mission and advance scientific progress.
This project is suited for a FRO-style approach because academia is too siloed to support the kind of large-scale, interdisciplinary research project that this would require, and companies that currently provide DNA-based RNA sequencing tools have no motivation to innovate beyond tweaks to the current methods.
How This Project Will Benefit Scientific Progress
The development of direct RNA sequencing technologies with all chemical modifications will enable the execution of the Human RNome Project, which is anticipated to result from the the National Academies of Sciences, Engineering and Medicine Report (NASEM) on RNA Sequencing. A similar NASEM report resulted in the creation of the Human Genome Project. The Human RNome Project aims to identify all the possible chemical modifications of RNA and create a true sequence of RNA, i.e. the RNome. Direct RNA sequencing technologies will also enable scientists to sequence the complete genome of RNA viruses, such as SARS-CoV-2, the hepatitis C virus, and influenza viruses and accelerate the development of new RNA therapeutics for these viruses and other diseases.
Key Contacts
Authors
- Vivian Cheung, University of Michigan, Life Sciences Institute, vgcheung@umich.edu
- George Langford, Syracuse University
Referrers
- Alice Wu, Federation of American Scientists, awu@fas.org
Learn more about FROs, and see our full library of FRO project proposals here.
A Focused Research Organization to Characterize Antibodies Through Open Science
Many antibodies that scientists purchase from commercial manufacturers to conduct their research do not work as advertised, because most have never been validated properly. This project brings together the public and private sectors to conduct independent, third-party testing of commercial antibody manufacturers’ catalogs and publish the results in the public domain, such that no scientist ever uses an ineffective antibody again.
Thousands of scientists use antibodies – each of which targets one of the 20,000 human proteins – to develop fundamental theories of human biology, and to identify targets for new medicines. These antibodies are often purchased from commercial antibody manufacturers, whose combined catalog contains between 3.5 million and 4.8 million products. But for more than 30 years, the scientific community has been aware that many of these antibodies do not work as advertised, meaning that they do not recognize the intended protein target, or recognize the target but also recognize non-specific targets that confound their use. This occurs because many if not most antibodies have never been validated, or have been validated using inferior or outdated scientific methods, and because academics do not have resources or skill sets to test them themselves. When an antibody binds to a non-targeted protein, a researcher may believe that the target protein, perhaps a drug target, is present in a particular cell type or subcellular organelle when in reality it is not. These erroneous results lead to a vast waste of time, resources, and human capital.
Project Concept
The science on the optimal antibody testing methodology is largely settled: using an appropriately selected wild type human cell and a CRISPR knockout version of the same cell as the basis for testing yields the most rigorous and broadly applicable results. However, the cost of testing for an individual target or antibody is often prohibitive for any individual academic lab or company. Our organization, YCharOS (Antibody Characterization through Open Science), couples the settled science with a unique open science business model, in which a consortium of antibody manufacturers provide, in-kind, all their renewable antibodies (i.e. monoclonal or recombinant, which once tested are of value in perpetuity) to any given target to YCharOS for use in direct, head-to-head comparisons. This centralized testing model creates massive economic efficiencies for the sector while also providing immense scientific benefit to the public. Moreover, since all data will be released into the public domain using the principles of open science, the benefits accrue to all. We envision a world where no scientist ever uses an antibody that has not been rigorously tested by an independent third party. We believe that renewable antibodies for all 20,000 human proteins can be knockout validated in many applications for a one-time total budget of approximately $100 million.
What is a Focused Research Organization?
Focused Research Organizations (FROs) are time-limited mission-focused research teams organized like a startup to tackle a specific mid-scale science or technology challenge. FRO projects seek to produce transformative new tools, technologies, processes, or datasets that serve as public goods, creating new capabilities for the research community with the goal of accelerating scientific and technological progress more broadly. Crucially, FRO projects are those that often fall between the cracks left by existing research funding sources due to conflicting incentives, processes, mission, or culture. There are likely a large range of project concepts for which agencies could leverage FRO-style entities to achieve their mission and advance scientific progress.
This project is suited for a FRO-style approach because antibody characterization is a time limited project that, once completed, will identify high-performing antibodies that can be produced and used in perpetuity. Antibody validation itself is unlikely to generate papers, but will create a public good that enables the production of new research results using properly validated antibodies.
How This Project Will Benefit Scientific Progress
Academic and pharmaceutical scientists laboring to advance our understanding and treatment of human disease will be able to save time and money and produce higher quality research using validated antibodies. Monetarily, scientists spend an estimated $1 billion per year on ineffective antibodies that could otherwise be spent on conducting further research. Furthermore, there is a not insignificant volume of faulty research publications that have resulted from scientists unknowingly using ineffective antibodies.
Key Contacts
Authors
- Chetan Raina, CEO of YCharOS, chetan.raina@ycharos.com
- Al Edwards, CEO of the Structural Genomics Consortium, aled.edwards@utoronto.ca
- Peter McPherson, McGill University, peter.mcpherson@mcgill.ca
Referrers
- Alice Wu, Federation of American Scientists, awu@fas.org
Learn more about FROs, and see our full library of FRO project proposals here.
A Focused Research Organization to Quantify Ocean Carbon
The [C]Worthy Project will create first-of-its-kind, open-source software infrastructure for ocean carbon measurement, reporting and verification (MRV) to help drive the nascent marine-based carbon dioxide removal market.
There is scientific consensus that industrial-scale Carbon Dioxide Removal (CDR) will be necessary to meet the Paris Agreement’s goal of keeping the rise in global temperature to within 1.5°C or 2°C. By enhancing the ocean’s natural capacity for carbon sequestration and long-term storage, ocean-based CDR is one of the few strategies with the potential to remove carbon at the necessary scales. While investments are pouring into the design and early-stage deployment of ocean-based CDR technologies, the ocean is a complex ecosystem, a constantly moving fluid, posing challenges to quantifying ocean-based CDR projects. It is essential to establish scientifically credible methods for quantifying net carbon removal from CDR deployments and codify these as standards for Monitoring, Reporting, and Verification (MRV). [C]worthy is proposing transformative, open-source, technical solutions to confront these challenges.
Project Concept
[C]worthy aims at building the first foundational open-source computational infrastructure for MRV of ocean-based carbon dioxide removal technologies. The [C]worthy team will develop an innovative and first-of-its-kind MRV software infrastructure applicable across the range of ocean-based CDR technologies that are currently under study and testing in the US and internationally. This MRV platform will integrate ocean observations and advanced Earth system modeling tools at regional to global scales; it will incorporate ocean biogeochemistry and ecosystem models and data assimilation capabilities and it will advance techniques for efficient computation and analysis. This platform will be designed as open-source infrastructure with interoperable components—ensuring transparency and broad access to a growing CDR community.
What is a Focused Research Organization?
Focused Research Organizations (FROs) are time-limited mission-focused research teams organized like a startup to tackle a specific mid-scale science or technology challenge. FRO projects seek to produce transformative new tools, technologies, processes, or datasets that serve as public goods, creating new capabilities for the research community with the goal of accelerating scientific and technological progress more broadly. Crucially, FRO projects are those that often fall between the cracks left by existing research funding sources due to conflicting incentives, processes, mission, or culture. There are likely a large range of project concepts for which agencies could leverage FRO-style entities to achieve their mission and advance scientific progress.
This project is suited for a FRO-style approach because the [C]worthy development project requires a level of coordinated computational engineering that is too big for a single academic lab, too complex for a loose multi-lab collaboration, and not directly profitable enough for a venture-backed startup or industrial R&D project. It is aiming to develop an open-source IT infrastructure, which is a product of public benefit for science and technology. This requires tight-knit coordinated effort in a fast-paced start-up-like environment. It is therefore best fit for the FRO model.
Key Contacts
Team lead
- Matt Long, PhD, matt@cworthy.org
Referrers
- Alice Wu, Federation of American Scientists, awu@fas.org
Learn more about FROs, and see our full library of FRO project proposals here.
A Focused Research Organization to Build the Foodome Project for the Future of Nutrition
Our current knowledge of the biochemical compounds in food is incredibly limited, but existing databases of MassSpec scans contain massive amounts of untapped, unannotated information about food ingredients. A project to leverage these databases with the tools of data mining, AI, and high-throughput measurement will systematically unveil the chemical composition of all food ingredients and revolutionize our understanding of food and health.
Diet is the single biggest determinant of health over which we have direct control. An unhealthy diet poses more risk to morbidity than alcohol, tobacco, drug use, and unsafe sex combined. Indeed, our diet exposes us to thousands of food molecules, many of which are known to play an important role in multiple diseases including coronary heart disease, cancer, stroke, and diabetes. Despite the demonstrated and complex role of diet on health, nutrition science remains focused on molecules that serve as energy sources such as sugars, fats, and vitamins, leaving most disease-causing compounds uncatalogued and invisible to researchers and health care professionals. Further, our current understanding of the way food affects health is limited to nutritional guidelines that rely on a panel of 150 essential micro- and macro-nutrients in our diet. This is a tiny fraction of the more than 130,000 compounds known to be present in food, hence limiting our ability to unveil the health implications of our diet.
Project Concept
The Foodome project aims to unveil this “dark matter of nutrition” by creating an open-access high-resolution compendium of food compounds through a strategy that combines Big Data, ML/AI, and experimental techniques, implemented by a focused cross-disciplinary team, motivated to bring transformative change and maximize public benefits.
In the past five years, BarabásiLab has curated the largest library of compounds in food, consisting of more than 135,000 biochemicals linked to 3,500 foods. While the number of biochemicals is exceptional, the coverage is highly uneven, sparse, and largely unquantified. Yet, information about the missing biochemicals is carried by the unannotated MassSpec peaks available for each MassSpec scan of food ingredients. Because chemicals are invisible to the one-chemical-one-peak tools employed today, we have designed a strategy that relies on data mining, AI, and high-throughput measurements to resolve them: we plan to collect the more than 3,000,000 MassSpec scans already available in databases, and mine the full scientific literature to collect knowledge on food composition. We also plan to take advantage of the increasing number of annotated genomes to infer their chemical makeup. These data will serve as input for a ML/AI platform designed to learn associations between biochemical structures and the ingredients’ phylogenetic position, helping us systematically unveil the chemical composition of all food ingredients.
What is a Focused Research Organization?
Focused Research Organizations (FROs) are time-limited mission-focused research teams organized like a startup to tackle a specific mid-scale science or technology challenge. FRO projects seek to produce transformative new tools, technologies, processes, or datasets that serve as public goods, creating new capabilities for the research community with the goal of accelerating scientific and technological progress more broadly. Crucially, FRO projects are those that often fall between the cracks left by existing research funding sources due to conflicting incentives, processes, mission, or culture. There are likely a large range of project concepts for which agencies could leverage FRO-style entities to achieve their mission and advance scientific progress.
This project is suited for a FRO-style approach because the Foodome platform and knowledge base will address problems in health science beyond the competence of any single academic group or start-up. The project started in the academic environment involving groups at Northeastern University, Harvard Medical School, and Tufts Medical School, but typical academic researchers and institutions are motivated by short term publication strategies and unable to devote the years needed to develop a public resource. Federal nutrition research funding exists, but is fragmented, and normal funding channels are generally unable to offer sustained support for a project of this size. With VC funding, we were able to move the project to a startup environment to standardize the toolset and develop key technologies, but company management decided that the Foodome platform’s timeline is too far from the market. Based on these experiences, an FRO appears to be the best framework to accomplish the vision of Foodome. The project enters a field limited by technological stagnation, and will fundamentally change our understanding of health and disease, impacting multiple fields and industries.
How This Project Will Benefit Scientific Progress
A high-resolution knowledgebase on the composition of food will revolutionize our ability to explore the role of each food-borne molecule in human health, impacting multiple fields: 1) It will be transformative for health care, changing our ability to prevent and control disease. 2) It will aid the development of healthier, more nutritious, and biochemically balanced foods. 3) It will facilitate the development of novel pharmaceuticals. 4) By improving MassSpec annotations, it will provide a more accurate biochemical descriptions of any sample, empowering diagnosis, and detection. 5) It will unlock innovations in personalized nutrition and precision medicine, allowing clinicians to offer precision advice to a patient on how to use diet to prevent and manage disease.
Key Contacts
Author
- Albert-László Barabási, Northeastern and Harvard Medical School, alb@neu.edu
Referrers
- Alice Wu, Federation of American Scientists, awu@fas.org
Learn more about FROs, and see our full library of FRO project proposals here.
Flexible Hiring Resources For Federal Managers
From education to clean energy, immigration to wildfire resilience, and national security to fair housing, the American public relies on the federal government to deliver on critical policy priorities.
Federal agencies need to recruit top talent to tackle these challenges quickly and effectively, yet often are limited in their ability to reach a diverse pipeline of talent, especially among expert communities best positioned to accelerate key priorities.
FAS is dedicated to bridging this gap by providing a pathway for diverse scientific and technological experts to participate in an impactful, short-term “tour of service” in federal government. The Talent Hub leverages existing federal hiring mechanisms and authorities to place scientific and technical talent into places of critical need across government.
The federal government has various flexible hiring mechanisms at its disposal that can help federal teams address the complex and dynamic needs they have while tackling ambitious policy agendas and programs. Yet information about how to best utilize these mechanisms can often feel elusive, leading to a lack of uptake.
This resource guide provides an overview of how federal managers can leverage their available hiring mechanisms and the Talent Hub as a strategic asset to onboard the scientific and technical talent they recruit. The accompanying toolkit includes information for federal agencies interested in better understanding the hiring authorities at their disposal to enhance their existing scientific and technical capacities, including how to leverage Intergovernmental Personnel Act Mobility Program and Schedule A(r) fellowship hiring.
FAS Forum: Envisioning the Future of Wildland Fire Policy
In this critical year for reimagining wildland fire policy, the Federation of American Scientists (FAS) hosted a convening that provided stakeholders from the science, technology, and policy communities with an opportunity to exchange forward-looking ideas with the shared goal of improving the federal government’s approach to managing wildland fire.
A total of 43 participants attended the event. Attendee affiliations included universities, federal agencies, state and local agencies, nonprofit organizations, and philanthropies.
This event was designed as an additive opportunity for co-learning and deep dives on topics relevant to the Wildland Fire Mitigation and Management Commission (the Commission) with leading experts in relevant fields (the convening was independent from any formal Commission activities).
In particular, the Forum highlighted, and encouraged iteration on, ideas emerging from leading experts who participated in the Wildland Fire Policy Accelerator. Coordinated by FAS in partnership with COMPASS, the California Council on Science and Technology (CCST), and Conservation X Labs, this accelerator has served as a pathway to source and develop actionable policy recommendations to inform the work of the Commission.
A full list of recommendations from the Accelerator is available on the FAS website.
The above PDF summarizes discussions and key takeaways from the event for participant reference. We look forward to building on the connections made during this event.
AI for science: creating a virtuous circle of discovery and innovation
In this interview, Tom Kalil discusses the opportunities for science agencies and the research community to use AI/ML to accelerate the pace of scientific discovery and technological advancement.
Q. Why do you think that science agencies and the research community should be paying more attention to the intersection between AI/ML and science?
Recently, researchers have used DeepMind’s AlphaFold to predict the structures of more than 200 million proteins from roughly 1 million species, covering almost every known protein on the planet! Although not all of these predictions will be accurate, this is a massive step forward for the field of protein structure prediction.
The question that science agencies and different research communities should be actively exploring is – what were the pre-conditions for this result, and are there steps we can take to create those circumstances in other fields?
One partial answer to that question is that the protein structure community benefited from a large open database (the Protein Data Bank) and what linguist Mark Liberman calls the “Common Task Method.”
Q. What is the Common Task Method (CTM), and why is it so important for AI/ML?
In a CTM, competitors share the common task of training a model on a challenging, standardized dataset with the goal of receiving a better score. One paper noted that common tasks typically have four elements:
- Tasks are formally defined with a clear mathematical interpretation
- Easily accessible gold-standard datasets are publicly available in a ready-to-go standardized format
- One or more quantitative metrics are defined for each task to judge success
- State-of-the-art methods are ranked in a continuously updated leaderboard
Computational physicist and synthetic biologist Erika DeBenedictis has proposed adding a fifth component, which is that “new data can be generated on demand.” Erika, who runs Schmidt Futures-supported competitions such as the 2022 BioAutomation Challenge, argues that creating extensible living datasets has a few advantages. This approach can detect and help prevent overfitting; active learning can be used to improve performance per new datapoint; and datasets can grow organically to a useful size.
Common Task Methods have been critical to progress in AI/ML. As David Donoho noted in 50 Years of Data Science,
The ultimate success of many automatic processes that we now take for granted—Google translate, smartphone touch ID, smartphone voice recognition—derives from the CTF (Common Task Framework) research paradigm, or more specifically its cumulative effect after operating for decades in specific fields. Most importantly for our story: those fields where machine learning has scored successes are essentially those fields where CTF has been applied systematically.
Q. Why do you think that we may be under-investing in the CTM approach?
U.S. agencies have already started to invest in AI for Science. Examples include NSF’s AI Institutes, DARPA’s Accelerated Molecular Discovery, NIH’s Bridge2AI, and DOE’s investments in scientific machine learning. The NeurIPS conference (one of the largest scientific conferences on machine learning and computational neuroscience) now has an entire track devoted to datasets and benchmarks.
However, there are a number of reasons why we are likely to be under-investing in this approach.
- These open datasets, benchmarks and competitions are what economists call “public goods.” They benefit the field as a whole, and often do not disproportionately benefit the team that created the dataset. Also, the CTM requires some level of community buy-in. No one researcher can unilaterally define the metrics that a community will use to measure progress.
- Researchers don’t spend a lot of time coming up with ideas if they don’t see a clear and reliable path to getting them funded. Researchers ask themselves, “what datasets already exist, or what dataset could I create with a $500,000 – $1 million grant?” They don’t ask the question, “what dataset + CTM would have a transformational impact on a given scientific or technological challenge, regardless of the resources that would be required to create it?” If we want more researchers to generate concrete, high-impact ideas, we have to make it worth the time and effort to do so.
- Many key datasets (e.g., in fields such as chemistry) are proprietary, and were designed prior to the era of modern machine learning. Although researchers are supposed to include Data Management Plans in their grant applications, these requirements are not enforced, data is often not shared in a way that is useful, and data can be of variable quality and reliability. In addition, large dataset creation may sometimes not be considered academically novel enough to garner high impact publications for researchers.
- Creation of sufficiently large datasets may be prohibitively expensive. For example, experts estimate that the cost of recreating the Protein Data Bank would be $15 billion! Science agencies may need to also explore the role that innovation in hardware or new techniques can play in reducing the cost and increasing the uniformity of the data, using, for example, automation, massive parallelism, miniaturization, and multiplexing. A good example of this was NIH’s $1,000 Genome project, led by Jeffrey Schloss.
Q. Why is close collaboration between experimental and computational teams necessary to take advantage of the role that AI can play in accelerating science?
According to Michael Frumkin with Google Accelerated Science, what is even more valuable than a static dataset is a data generation capability, with a good balance of latency, throughput, and flexibility. That’s because researchers may not immediately identify the right “objective function” that will result in a useful model with real-world applications, or the most important problem to solve. This requires iteration between experimental and computational teams.
Q. What do you think is the broader opportunity to enable the digital transformation of science?
I think there are different tools and techniques that can be mixed and matched in a variety of ways that will collectively enable the digital transformation of science and engineering. Some examples include:
- Self-driving labs (and eventually, fleets of networked, self-driving labs), where machine learning is not only analyzing the data but informing which experiment to do next.
- Scientific equipment that is high-throughput, low-latency, automated, programmable, and potentially remote (e.g. “cloud labs”).
- Novel assays and sensors.
- The use of “science discovery games” that allow volunteers and citizen scientists to more accurately label training data. For example, the game Mozak trains volunteers to collaboratively reconstruct complex 3D representations of neurons.
- Advances in algorithms (e.g. progress in areas such as causal inference, interpreting high-dimensional data, inverse design, uncertainty quantification, and multi-objective optimization).
- Software for orchestration of experiments, and open hardware and software interfaces to allow more complex scientific workflows.
- Integration of machine learning, prior knowledge, modeling and simulation, and advanced computing.
- New approaches to informatics and knowledge representation – e.g. a machine-readable scientific literature, increasing number of experiments that can be expressed as code and are therefore more replicable.
- Approaches to human-machine teaming that allow the best division of labor between human scientists and autonomous experimentation.
- Funding mechanisms, organizational structures and incentives that enable the team science and community-wide collaboration needed to realize the potential of this approach.
There are many opportunities at the intersection of these different scientific and technical building blocks. For example, use of prior knowledge can sometimes reduce the amount of data that is needed to train a ML model. Innovation in hardware could lower the time and cost of generating training data. ML can predict the answer that a more computationally-intensive simulation might generate. So there are undoubtedly opportunities to create a virtuous circle of innovation.
Q. Are there any risks of the common task method?
Some researchers are pointing to negative sociological impacts associated with “SOTA-chasing” – e.g. a single-minded focus on generating a state-of-the-art result. These include reducing the breadth of the type of research that is regarded as legitimate, too much competition and not enough cooperation, and overhyping AI/ML results with claims of “super-human” levels of performance. Also, a researcher who makes a contribution to increasing the size and usefulness of the dataset may not get the same recognition as the researcher who gets a state-of-the-art result.
Some fields that have become overly dominated by incremental improvements in a metric have had to introduce Wild and Crazy Ideas as a separate track in their conferences to create a space for more speculative research directions.
Q. Which types of science and engineering problems should be prioritized?
One benefit to the digital transformation of science and engineering is that it will accelerate the pace of discovery and technological advances. This argues for picking problems where time is of the essence, including:
- Developing and manufacturing carbon-neutral and carbon-negative technologies we need for power, transportation, buildings, industry, and food and agriculture. Currently, it can take 17-20 years to discover and manufacture a new material. This is too long if we want to meet ambitious 2050 climate goals.
- Improving our response to future pandemics by being able to more rapidly design, develop and evaluate new vaccines, therapies, and diagnostics.
- Addressing new threats to our national security, such as engineered pathogens and the technological dimension of our economic and military competition with peer adversaries.
Obviously, it also has to be a problem where AI and ML can make a difference, e.g. ML’s ability to approximate a function that maps between an input and an output, or to lower the cost of making a prediction.
Q. Why should economic policy-makers care about this as well?
One of the key drivers of the long-run increases in our standard of living is productivity (output per worker), and one source of productivity is what economists call general purpose technologies (GPTs). These are technologies that have a pervasive impact on our economy and our society, such as interchangeable parts, the electric grid, the transistor, and the Internet.
Historically – GPTs have required other complementary changes (e.g. organizational changes, changes in production processes and the nature of work) before their economic and societal benefits can be realized. The introduction of electricity eventually led to massive increases in manufacturing productivity, but not until factories and production lines were reorganized to take advantage of small electric motors. There are similar challenges for fostering the role that AI/ML and complementary technologies will play in accelerating the pace of scientific and technological advances:
- Researchers and science funders need to identify and support the technical infrastructure (e.g. datasets + CTMs, self-driving labs) that will move an entire field forward, or solve a particularly important problem.
- A leading academic researcher involved in protein structure prediction noted that one of the things that allowed DeepMind to make so much progress on the protein folding problem was that “everyone was rowing in the same direction,” “18 co-first authors .. an incentive structure wholly foreign to academia,” and “a fast and focused research paradigm … [which] raises the question of what other problems exist that are ripe for a fast and focused attack.” So capitalizing on the opportunity is likely to require greater experimentation in mechanisms for funding, organizing and incentivizing research, such as Focused Research Organizations.
Q. Why is this an area where it might make sense to “unbundle” idea generation from execution?
Traditional funding mechanisms assume that the same individual or team who has an idea should always be the person who implements the idea. I don’t think this is necessarily the case for datasets and CTMs. A researcher may have a brilliant idea for a dataset, but may not be in a position to liberate the data (if it already exists), rally the community, and raise the funds needed to create the dataset. There is still a value in getting researchers to submit and publish their ideas, because their proposal could be catalytic of a larger-scale effort.
Agencies could sponsor white paper competitions with a cash prize for the best ideas. [A good example of a white paper competition is MIT’s Climate Grand Challenge, which had a number of features which made it catalytic.] Competitions could motivate researchers to answer questions such as:
- What dataset and Common Task would have a significant impact on our ability to answer a key scientific question or make progress on an important use-inspired or technological problem? What preliminary work has been done or should be done prior to making a larger-scale investment in data collection?
- To the extent that industry would also find the data useful, would they be willing to share the cost of collecting it? They could also share existing data, including the results from failed experiments.
- What advance in hardware or experimental techniques would lower the time and cost of generating high-value datasets by one or more orders of magnitude?
- What self-driving lab would significantly accelerate progress in a given field or problem, and why?
The views and opinions expressed in this blog are the author’s own and do not necessarily reflect the view of Schmidt Futures.