Cancer Risk Is Still There, Even If the Data Isn’t
Can you estimate someone’s cancer risk from the air they breathe? Not exactly. Every person has different genetics, lifestyles, occupations, and environmental exposures that change over time. Air pollution varies from day to day, season to season, and even block to block. But scientists can make remarkably reliable estimates of a neighborhood’s cancer risk based on the type and amount of toxic air pollution.
After a year-long delay, this April, EPA released the latest air toxics data, which only included raw air data downloads. This year, for the first time in nearly 25 years, the air toxics data did not include cancer risk estimates.
Without a public explanation or opportunity for input, one of the nation’s most important environmental health datasets has quietly gone dark.
The disappearance of EPA’s cancer risk data continues a broader trend under this administration of environmental, public health, and other government datasets becoming less available, less complete, or more difficult to access.
What’s Changed?
For the 2021 data (this most recent release), only raw air emissions and concentration data are available for download. Previously, the EPA also produced detailed explanations and a mapping tool that made the data easier to discover, understand, and use. Additionally, they included cancer risk estimates to help translate data and complex scientific models into information the public could understand.
Now, if users want to understand cancer risk from air pollutants, they will need to download and analyze detailed air concentrations of nearly 200 toxic air pollutants for more than 8 million individual locations.
Researchers and geospatial data professionals can navigate government databases, process large datasets, and build custom analyses or maps. Everyone else relies on easy-to-use websites, interactive maps, and plain-language explanations to understand the environmental conditions affecting their communities.
Data hidden behind obscure downloads may satisfy a technical definition of transparency, but they do little to promote public understanding, meaningful community engagement, or actions to save lives.
Why Cancer Data Matter
Cancer affects nearly every family in the country, and the United States has one of the highest cancer incidence rates in the world. Understanding where people face higher exposure to cancer-causing pollutants is a key step toward reducing preventable risks.
Most people have no easy way of knowing what hazardous pollutants are being released into the air around their homes, schools, or workplaces. Air toxics are often invisible, odorless, and linked to health effects that can take years or even decades to develop. EPA’s cancer risk estimates gave communities, local governments, researchers, and others a way to identify areas where cancer risks from air pollution may be elevated and where action or additional monitoring might be needed.
History of Cancer Risk Data
The Clean Air Act lists nearly 200 toxic air pollutants known to cause cancer or other serious health effects, and requires reporting of these emissions by facilities. Amendments in 1990 directed the EPA to establish risk standards for any source emitting a cancer-causing pollutant that poses a lifetime risk of cancer of more than one-in-a-million.
EPA began publishing nationwide cancer risk estimates from hazardous air pollutants in 2002 with the first release of the National Air Toxics Assessment (NATA). These cancer risk estimates took hazardous air pollutants information from across the country (using facility emissions, vehicles, and other pollution sources), and combined them together with atmospheric modeling and toxicological research, making the data easier for the public to understand and use. This dataset became a core component of EPA’s mapping applications and was incorporated into numerous federal, state, academic, and community tools. Like any national-scale model, the NATA cancer risk data had limitations. It relied on emissions inventories, many of which were self-reported by industry, and modeled pollution levels for some locations, rather than direct air monitoring. It could not fully capture localized conditions or cumulative exposures, and the results often lagged several years behind current conditions. But public feedback on those limitations also inspired continuous improvements.
Over time, EPA scientists improved the quality of the pollution data, updated what scientists know about the health effects of toxic chemicals, and made the models more accurate. In 2022, the agency rebranded the program as the Air Toxics Screening Assessment or AirToxScreen, making the data more accessible and actionable through a dedicated mapping platform. With EPA having done the heavy lift of creating cancer risk estimates, other organizations can incorporate them into local processes. For example, New Jersey’s Environmental Justice, Mapping, Assessment, and Protection Tool (EJMAP) uses cancer risk data to help evaluate the cumulative impacts of facilities before approving or renewing permits.
Communities Impacted by Air Pollution-Related Cancer
Perhaps nowhere are both the strengths and limitations of these EPA data more apparent than in Louisiana’s Cancer Alley. This 85-mile industrial corridor along the Mississippi River contains one of the highest concentrations of petrochemical facilities in the United States, accounting for roughly one-quarter of the nation’s production. Residents have spent decades raising concerns about pollution and unusually high cancer rates. Here, some communities experience estimated lifetime cancer risks from toxic air pollution up to 50 times higher than the national average, and more than 47 times higher than EPA’s acceptable cancer risk benchmark. These findings have helped draw national attention to environmental justice concerns and have even prompted international human rights scrutiny.
At the same time, Cancer Alley demonstrates why screening and mapping tools alone are not enough. AirToxScreen is based largely on emissions inventories and modeling assumptions. It cannot fully account for cumulative exposures across multiple pathways, capture every emission event, or reflect neighborhood-scale differences that can vary dramatically within a city. Similar concerns have been raised by communities along the Houston Ship Channel, Chicago’s industrial corridors, and other heavily burdened areas across the country.
For these communities, data like AirToxScreen’s cancer risk estimates are an essential first step, but need to be accompanied by direct air monitoring, epidemiological research, or community-based science. And data need to be tied to action, such as strong environmental enforcement, in order to reduce Americans’ exposure to cancer-causing chemicals.
Additional Issues and Concerns
The disappearance of EPA’s cancer risk estimates comes at a time when protections against hazardous air pollutants are also under increasing pressure. The current administration has proposed rolling back or reconsidering several air toxics regulations, including standards for pollutants such as ethylene oxide, one of the largest contributors to cancer risk from industrial emissions.
At the heart of many of these proposals is a broader debate over what level of cancer risk should be considered acceptable and how that risk should be calculated. Loosening these standards does not reduce the amount of pollution in the air or the health risks facing nearby communities. Instead, it changes how those risks are defined and regulated. Combined with the loss of publicly available cancer risk estimates, these policy changes make it more difficult for communities to understand, communicate, and respond to the environmental health risks they face.
What’s the Data Policy Angle?
For more than two decades, EPA scientists steadily improved these estimates despite chronic underfunding, changing regulations, and shifting political priorities. Those challenges have only intensified with the dismantling of EPA’s Office of Research and Development and the erosion of scientific capacity across the agency.
EPA’s cancer risk data, though imperfect, has been an indispensable resource for researchers studying environmental disparities, journalists uncovering pollution hotspots, regulators prioritizing inspections, and communities advocating for cleaner air.
Now, the cancer risk estimate data so many depend on has disappeared. There was no public notice, no opportunity for public comment, and no consultation with one of several terminated scientific federal advisory committees.
The lack of public engagement is a policy problem.
The Open Government Data Act (OGDA), signed by President Trump in his first term, and subsequent OMB Guidance in M-25-05 include explicit requirements that agency data stewards engage with public stakeholders to understand the value of federal data, and ways it could be improved to better serve the needs of the American people.
OGDA and M-25-05 also unambiguously state that agencies:
- shall assist “the public in expanding the use of public data assets”; and
- “must provide adequate notice when initiating, substantially modifying, or terminating significant information dissemination products.”
The disappearance of EPA’s cancer risk data is just the most recent signal that the current patchwork of federal data policies is increasingly unprepared to address the challenges and opportunities facing our nation.
These are exactly the types of issues that the FAS Data Policy Institute is taking on, because cancer risk does not go away when the data disappear.
Revealing the Hidden Data Supporting Alzheimer’s Patients with the Federal Data Field Guide
Federal data play an essential – and mostly invisible – role in supporting the more than 7 million Americans with Alzheimer’s and their nearly 12 million unpaid caregivers.
Too often, we think of federal data as limited to high-profile datasets like jobs, weather, and the census. But beneath that surface is a diverse ecosystem with well over 500,000 datasets – including those tackling Alzheimer’s disease and related dementias (ADRD).
We recently published the Federal Data Field Guide to highlight the different species of federal data that benefit everyday Americans. In this post, we use the Field Guide’s framework of eight categories (Statistical, Administrative, Geospatial, Scientific, Accountability, Evaluation, Navigation, and Reference data) to scout for federal datasets that are improving the lives of Americans affected by ADRD.
Federal Data Improving Our Understanding of ADRD
Here’s a quick look at the valuable federal data underpinning our understanding Alzheimer’s disease and related dementias (ADRD):
How prevalent is Alzheimer’s Disease in the U.S.?
Statistical Data measure population level characteristics. The National Center for Health Statistics’s National Health Interview Survey estimates that about 4% of the non-institutionalized population over 65 has been diagnosed with dementia, while mortality data from the National Vital Statistics System (NVSS) track deaths from Alzheimer’s. NVSS also fits in the Administrative data category from the Field Guide, because it draws information from death certificates processed by state governments.
What are the genetic, social, and environmental determinants of Alzheimer’s disease?
Geospatial Data describe location and environmental information about the world. Evidence links air pollution to increased Alzheimer’s risk. The EPA’s Particulate Matter Pollution data–collected through a network of monitors operated by state, local, and tribal agencies–are vital for enforcing clean air regulations and, by extension, reducing risk for ADRD.
Scientific Data advance knowledge through federal or federally-funded research. The Veterans Administration’s Million Veteran Program (MVP) has identified key variables associated with ADRD in veterans, such as traumatic brain injury, depression, and military environmental exposures, offering critical insights for both prevention and intervention. And, the NIH’s GenBank hosts a trove of genetic data that researchers use to develop screening tools, treatments, and medications for ADRD.
How do federal data help caregivers and patients?
Navigation Data help citizens find and access services. The Centers for Medicare and Medicaid Services’ Nursing Home Care Compare Database helps caregivers assess facilities for their loved ones. This dataset also fits into the Accountability category, as it includes critical quality metrics. NIH’s ClinicalTrials.gov dataset helps families identify the clinical trials that might be a match for their loved one. Another important navigation dataset from NIH is PubMed, which houses over 40 million citations to the biomedical literature, enabling scientists, clinicians, and families to stay up on the latest Alzheimer’s research.
Reference Data provide standardization across systems. How can this type of data help people with ADRD? The Social Security Administration maintains a reference dataset on medical conditions qualifying for Compassionate Allowances. This in turn enables a 45-year-old diagnosed with early-onset Alzheimer’s to fast-track her disability benefits because the condition is officially recognized in the dataset.
Evaluation Data assess how effective different programs are. For example, HHS has funded data collection on millions of telehealth appointments to evaluate access, quality, and utilization. These data are used to improve telehealth services, which can be a game changer for ADRD caregivers – especially those with mobility impairments that make it hard to make in-person medical appointments.
These are just 10 of the many federal datasets related to ADRD that span the eight categories of federal data. Without understanding the different categories of federal data, it is easy to overlook datasets like these that are essential for improving the lives of people impacted by Alzheimer’s disease.
It is more urgent to appreciate the interplays of data brought to bear on just one health issue at a time when administrative actions, budget cuts, and destaffing, threaten the capacity of agencies to collect, maintain, and publish them.
Mapping out a broad range of federal datasets in a specific domain like ADRD is a useful exercise for two reasons. One, most federal data are underutilized or taken for granted. We’ve already paid for these data as taxpayers and we should make good use of them; every reuse of a federal dataset is value-added to its return on investment. And two, the federal data ecosystem will be more resilient when more people know about, care about, and advocate for the datasets on which they depend. Identifying and talking about the value of federal data is the first step to protecting their continued flow.
Take action
Do you want to get a better understanding of the federal data that make your life, or your life’s work, better? And do you want to get the tools to help keep those essential data flowing?
Here are some concrete steps you can take:
- Use the data categories in the federaldatafieldguide.us to map out the federal datasets you might be taking for granted, whether you work on issues around housing, climate, veterans, children, small businesses, agriculture, or whatever.
- Sign up for the newsletter at dataindex.us to be up to speed on opportunities to give feedback to federal agencies on specific datasets or data policies.
- Visit the Data Checkup at dataindex.us to see which of your datasets have been assessed for risk – and let us know what datasets you value and want prioritized.
- Check out essentialdata.us to see if the datasets you’ve identified are represented in our collection of use cases about how federal data benefit everyday Americans. If not, submit a dataset, or reach out to schedule a workshop to create use cases that cover your domain.
The Federal Data Field Guide is a free, plain-language resource developed by Denice Ross and Christopher Marcum as part of an Executive Fellowship in Applied Technology Policy at UC Berkeley. Learn more at federaldatafieldguide.us.
The FAS Data Policy Institute is a catalyst for advancing the field of data policy to build better outcomes for the American people, by building the civic infrastructure to monitor changes to federal datasets; mobilize data stakeholders to engage with government officials; advance policies to protect and improve essential public data; and design America’s future data ecosystem.
Photo: Histopathogic image of senile plaques seen in the cerebral cortex in a patient with Alzheimer disease of presenile onset, CC-BY KGH.