Government Capacity

How Do We Track Terminations of Federal Data?

08.10.26 | 11 min read | Text by Denice Ross & Chris Dick

Federal data benefit American lives and livelihoods in ways most people never see, touching every corner of our lives. This includes a farmer pricing a crop, a county planning a hospital, a business siting a warehouse — all of these decisions use federal data. Other data save patients money by identifying generic drugs that can replace more expensive brand names, help airplanes avoid deadly bird strikes, and warn consumers about recalls of dangerous products. 

Because these data are mostly invisible, their disappearance is invisible too — that is, until we need the data and they aren’t there anymore.

As federal data policy nerds, the question we get asked all the time is How much data has the current administration terminated?

The answer is that it depends on what we count as “data” AND what counts as a termination. That’s not a dodge – working through those two critical nuances is the substance of this piece.

The answer also depends on what the information will be used for. Ours is a data policy question, so we looked for structured, numerical datasets that have been terminated – meaning there will be no collections of those data in the future. Defining terminations that way lets us ask how agencies consulted with the public about a dataset’s value before ending it, and what its loss means for the federal government’s ability to serve the American people.

The purpose of the Federal Data Terminations Tracker is to be the most policy-relevant, verified accounting of federal data terminations available.

To create this Tracker, the dataindex.us team identified dozens of federal datasets and hundreds of data elements that have been terminated – significantly fewer than other reports, but more tailored to informing future data policies needed to run a modern society. These are data that have long underpinned policymaking, journalism, advocacy and research that improve American lives and livelihoods. These figures will change as terminations continue, collections are merged, and as court orders restore data. 

Want more details on why this question about federal data losses is so tricky? Read on. Want to see what data have been terminated? Visit the Federal Data Terminations Tracker at dataindex.us/terminations-tracker.

Why are there different numbers on federal data losses?

The number of datasets in data.gov is not a meaningful metric.

Studies and news reports over the last 18 months have varied widely, with the number “3,000 datasets removed” often being cited. That figure generally traces back to data.gov’s catalog counter, which dropped by roughly 3,000 entries in the administration’s first weeks. The counter is a poor measure of terminations, though. 

In early spring 2025, Harvard’s Library Innovation Lab watched the data.gov collection swing up by 6,000 datasets one month, and then back down by 2,000 the following month; they note that fluctuations like these arise naturally as websites and data stores are reorganized. Also, the definition of a “dataset” is broad in data.gov, including datasets, parts of datasets, and other data assets, meaning it could easily count one American Community Survey as hundreds of “datasets” (if you count each of the many files that are released each year). 

Some definitions of “data” are broader.

Counts from rescue efforts like the Data Rescue Project  or EDGI measure a broader range of information losses, including web pages, files, reports, and historical data taken offline. Additionally, other definitions, taking a broad definition of what constitutes data, focus on content that has been removed from public access, rather than data collections which are terminated going forward.

What should be included when we count data terminations?

Easy choices

Let’s start with the easy part to define. When an agency ends a primary data collection – for instance, the survey stops going out, the forms stop being filed, or the sensor stops collecting data – that is an obvious data termination. The USDA’s CPS Food Security Supplement, SAHMSA’s Drug Abuse Warning Network, and the USDA’s Mink Survey all ended this way. This is clear cut and little judgment is required to declare them “terminated.”

Sometimes a collection is terminated, and technically other entities could or do collect that data, but without the gravitas of the U.S. Government. For example, until last year, many U.S. Embassies collected air quality data and reported those numbers publicly. Of course, others can still collect air quality data for these cities, but the official U.S. data provided an unbiased source of data to compare to numbers from local officials, which research showed resulted in “substantial reductions in fine particulate concentration levels.” 

More difficult choices

The next group – derivative products and modeled data – takes more thought. These include things like composites, indices, and model outputs that are built from data that still exist somewhere in government (or outside of it), so ending them doesn’t destroy an underlying collection of data. 

We include these when they can’t be reproduced outside of government easily, or when they would stop serving their purpose if they were anything other than a government asset. An example is NOAA’s terminated Billion Dollar Weather and Climate Disasters data. Climate Central hired the NOAA researcher behind the data and reconstructed the model. Stewarding this dataset outside of government requires substantial resources and expertise. Being published by Climate Central means that the data may become more accessible and relevant as the nonprofit improves the product, but its effectiveness as a tool for policy change and action might be reduced without the imprimatur of the federal government.

We include data like the Billion Dollar Disasters data in our count, while recognizing that the line is blurry and reasonable people will draw it in different places.


The level of granularity that we count affects the total number of data terminations even more than what we count. For example USAID released multiple files for each country studied in the Demographic and Health Surveys, and each one could be called a “dataset,” but these surveys constitute one data product for our analysis. Counting files is a choice that, in our opinion, inflates the number. Instead our unit of analysis, what we actually count, are products. 

BEA presents the same problem in reverse. Its collections continue, but the published tables and reports, which are sometimes the only place where the public has access to the data, have stopped. An example of this is metro-area GDP, the official measure of whether individual regional economies were growing or shrinking. Is each table or report a termination? Should we count these at all? We landed somewhere in between: we count them as a single discontinued data product for the whole agency and note that the collections behind them survive (as well as providing links to BEA’s own list of terminated reports).

Most difficult choices

Further, “terminated” is not necessarily a forever status. Most of the data taken down to comply with the day-one executive orders came back. The CDC’s Social Vulnerability Index returned by court order, but data are no longer updated. The Household Pulse Survey’s gender identity fields were removed, then back in the historical data by May 2026, but are no longer part of the survey moving forward. Other cases sit in a gray zone: CDC’s PRAMS is still collected but is no longer published at the federal level, and does not receive standardized national weights. The Federal Employee Viewpoint Survey was terminated and is “being reenvisioned.” The Violence Against Children and Youth Survey finished its pilot, lost its team to reductions in force, and ended without any formal notice or publication of results.

How do we measure “data termination”? 

In order to be most relevant for data policy, we count a data termination as:

  1. the documented ending of a primary federal data collection; 
  2. the discontinuation of a derivative or modeled data product that could not easily be reproduced outside government, or would not serve its intended purpose if it were not a government asset; or
  3. the removal of substantive elements from a continuing product (counting at the level of the data product).

Additionally, a data termination must be tied to primary evidence, for instance an agency notice, a press release, or a regulatory filing.

We track but do not count temporary takedowns that have been restored in our current total (though restored data doesn’t equal data that is safe forever). This means our number of terminated datasets is versioned, and updated at least quarterly. 

So how much data has been terminated? (as of July 2026)

With the scope set defined above, we can get to the numbers. Terminated data number roughly two dozen (28 data products as of July 2026). Some of these terminations were relatively easy to find. 

Others ended quietly, or were never announced at all. 

Removed data elements are a larger number. Hundreds of continuing datasets have lost variables, most commonly gender identity, sexual orientation, race, and ethnicity. 

Some of these changes are visible in the paperwork if you know where to look.

Following our granularity rule, BEA’s discontinued tables and reports count as one product, with a link to the agency’s own inventory. 

The 350 discontinued Producer Price Index series counts as one as well. The Billion-Dollar Weather and Climate Disasters dataset and the Future Risk Index each count on their own. These derivative data products aren’t easily recreated outside of government (though some like the Billion-Dollar Weather and Climate Disasters dataset, have been).

Two judgment calls 

Everything above describes what we count. Two harder questions sit underneath the count, and we want to be clear about where our methodology becomes more subjective.

The baseline problem

The first is the baseline problem. Data terminations happen in every administration. Statistical agencies retire collections as industries shrink, methods improve, and budgets tighten. BLS maintains entire pages of discontinued databases and series, and most of its recent CPI discontinuations date to 2024. Judging any single termination means asking whether it is part of that ordinary churn or something above it.

BEA’s metro-area statistics show how ordinary a termination can be. The agency stopped publishing personal income and GDP for metropolitan areas, and at first glance that looks like exactly the kind of loss we track. But BEA documented its reasons: publishing metro aggregates forced more privacy suppressions on its county statistics, and shifting OMB geography definitions made the aggregates unreliable across the full time series. The county-level data continue, and the agency released a geographic aggregator tool so users can build their own metro estimates. We do not consider this termination above baseline, and therefore is not counted in the Terminations Tracker.

The Mink Survey is the harder call. National Agricultural Statistics Service (NASS) ended it in August 2025, saying it was no longer necessary, and the mink industry has in fact been shrinking for years. Read that way, it is routine modernization. But mink farms were a documented site of COVID transmission between animals and people, which makes the survey relevant to public health, not just to agricultural economics. Our assessment is that its public-health value argued for keeping it.

Uniqueness to government

The second judgment call is whether every termination is a loss that only the government can repair. Some data can be produced outside the government. Public Environmental Data Partners and Fulton Ring’s revival of the Homeland Infrastructure Foundation-Level Data (HIFLD) is one example. But that path exists only when the inputs are public and a well-resourced steward steps forward, and neither is guaranteed. No private group can publish misconduct records held inside federal law-enforcement agencies the way the National Law Enforcement Accountability Database did. When data like those end, they simply end.

These two judgment calls are why we publish our reasoning alongside our numbers, which brings us to how we actually did this research.

How we did the research

dataindex.us monitors America’s federal data infrastructure, from dataset availability and new releases to planned and unplanned changes to collections. Our team aggregates the OMB information collection request pipeline, which is otherwise difficult to find and harder to interpret, flag opportunities for public comment, and track selected datasets and their documentation so that when data go missing, get reposted, or fail to release on schedule, there is a record. We mix primary data collection, both manual and automated, with secondary sources to better understand the overall health of federal data.

Finding terminations is harder than it sounds, because there is no single place where they are announced. Federal data policy is fragmented, and some types of data are far more visible than others. Surveys and forms covered by the Paperwork Reduction Act must go through public notice and comment for any substantive change. Curiously, that requirement covers changes – but not terminations – unless the data collection is tied to a regulation. Even so, the paperwork trail on reginfo.gov is often our best evidence. Data collected through sensors, satellites, and other non-paperwork means typically have no statutory notification requirements at all. For those, we depend on agency websites, news reports, and experts who flag changes when they see them.

Some agencies do announce their terminations, and it makes a difference. NESDIS keeps a log of decommissioned products. BEA maintains its list of discontinued or delayed statistics, BLS its pages of discontinued databases and series, NASS its newsroom notices. We wish this were the rule rather than the exception.

For everything else, we rely on partnerships and monitoring: organizations like the Data Rescue Project, the researchers and journalists who publish inventories of lost data, social media and the news, and tips from data users who notice that something they depend on has gone dark. We do this in both manual and automated ways, and our process will continue to improve over time.

And then we verify. Every reported termination gets a line by line review before it enters our count. We deduplicate, because fifty state pages of one product are one product. We separate data products from web pages that describe data. We trace each item to primary evidence (an agency notice, a press release, an ICR, a CRS report, or data actually being no longer available on a website), and we record status changes over time, because as we said above, termination is not necessarily a forever status. 

That standard is also an invitation to you. If a dataset you rely on disappears, loses variables, or quietly stops updating, tell us. The more eyes on federal data, the smaller the chance that something essential vanishes without anyone noticing. 

We are here to verify data terminations at removals@dataindex.us. To learn more about the Federal Data Terminations Tracker, register for our webinar on August 19, 2026.