We Need An Independent AI Evaluation Clearinghouse (IAIEC) to Provide Accreditation, Funding, and Independence in the AI Assurance Ecosystem
Nowhere would consumers simply rely on automobile manufacturers’ own assessments of vehicle safety. Instead, credible crash-test results depend on independent evaluation. Similarly, parents would not give toys to children that toy manufacturers unilaterally certified themselves as safe for children. Aircraft manufacturers are subject to independent testing and certification regimes. Pharmaceutical companies cannot self-certify that a new drug is safe and effective. And public companies have their books audited by third-parties accredited by an independent board with government oversight.
Yet for AI systems, safety claims still rely entirely on company-produced evaluations or company-selected vendors for third-party assessments.
This memo draws on existing models of federal oversight of safety testing and auditing regimes to outline a model for a robust, independent AI auditing, safety evaluation, and accountability system.
It proposes the creation of a federal clearinghouse and an opt-in industry fee structure to finance an ecosystem of accredited third-party auditors. The clearinghouse will also neutrally mediate the assignment of auditors to audits—strengthening independence, trust, and ethical governance across the AI ecosystem.
Our four recommendations, detailed below, are to create an Independent AI Evaluation Clearinghouse (IAIEC) with direct government oversight of AI auditing firms; adopt a formal accreditation process with clear minimum benchmarks and standards for AI auditors; create a sustainable funding environment for new market participants in AI auditing; and use the IAIEC as an neutral intermediary to assign audits to auditors to maximize independence from industry influence and better prevent conflicts-of-interest between model developers and auditing firms.
Challenge and Opportunity
There is a growing consensus across civil society, the research community, and private industry, and across political parties, that we need pre-deployment oversight and testing of AI models to occur independently of AI model developers. Yet at precisely the time when independent federal agencies could step into this role directly, Supreme Court decisions like Loper Bright and Slaughter have introduced new constraints on the independence and expertise of federal agencies. As a result, the function of neutral, technical oversight by independent government agencies will be significantly eroded going forward. In this new regulatory environment with increased partisan influence on agencies, there is a need for Congress to ensure that federal agencies can still play a meaningful role in overseeing private third-party AI governance bodies.
The third-party AI auditor ecosystem represents a convergence of preexisting efforts to advance algorithmic auditing, research interest in the risks and capabilities of frontier models, and the growing consensus that evaluation and assessment will be a core component of frontier AI governance thanks to recent laws like the EU AI Act, California’s SB 53, New York’s RAISE Act, Connecticut’s Senate Bill 5, and Illinois’ AI Safety Measures Act, which often reference and rely on an auditing ecosystem. In its current form, AI auditing is a private ecosystem of corporate self-attestation with voluntary engagement of a preferred auditing firm, allowing AI companies to effectively regulate themselves. This status quo presents four key policy challenges: untrustworthy self-regulation, funding and limited ecosystem development, a lack of robust shared auditing standards, and a narrow, uncompetitive assurance ecosystem.
What does AI auditing and evaluation actually provide?
AI auditing and evaluation is an attempt to answer a simple question: Does an AI system actually perform the way a developer says it does?
In current form, AI evaluators and auditing firms test a model against their internally defined benchmarks, probe for specific defined risks and failure points, and provide an outside assessment of a product’s potential capabilities, limitations, and safeguards. When done well, this can give companies, governments, and the public information they cannot get from a model developer’s claims alone.
However, an audit is not a guarantee that a model is “safe” or even accurate. The value of these evaluations depends on a myriad of factors: what was tested, when a model was tested, how it was tested, what information the auditor could access, and whether the auditor was genuinely independent. AI audits should thus not be seen as a proxy for safety, nor should they be used in lieu of proper regulatory oversight and enforcement action.
Self-regulation in a high-risk market creates doubt.
Currently, developers hire their own auditors, and thus can choose auditing methods and firms that may present developers performance and product safety in the most favorable light rather than the most accurate. This casts doubt on the reliability and rigor of AI auditing outcomes, and leaves policymakers uncertain about supporting auditing as a mechanism for AI risk assurance. The self-regulatory environment risks reproducing the failures seen in the advertising technology ecosystem, where industry self-regulation has often provided insufficient accountability and limited recourse for harmful or irresponsible actors. As such, there is an urgent need for independent assessment of AI products that is overseen by federal regulators, not just self-attestation of safety by industry-adjacent auditors. Historically regulated industries, and domains such as financial services and bank supervision, provide a model of independence to guide this design.
Existing funding models compromise the independence of investigations.
The existing market for third party auditing currently suffers from serious risks of industry capture and inadequate mechanisms to prevent conflicts of interest. The issuer-pays model used by credit rating agencies in the leadup to the global financial crisis is already widely understood to breed negative market incentives for participating firms in financial markets. The proliferation of issuer-pays models in AI auditing and evaluation closely mirrors these dynamics. Therefore, there is a real need to move away from the current “developer-pays” model, towards a more holistic approach to assurance in the AI industry. A stronger system of verification, such as Independent AI Evaluation Clearinghouse, would prevent model developers from directly retaining their own third-party auditor, and would instead build meaningful firewalls between model developers and the AI evaluator industry.
A lack of common benchmarks, standards, and evaluation processes.
Another persistent challenge within the ecosystem of third party AI auditing is the lack of universally accepted signals of empirical rigor and clear standards against which auditors can validate AI systems or developer practices. This issue, combined with a lack of standardized process for model evaluation, results in justifiable concerns about the rigor of both individual audit results and the discipline as a whole. As far back as 2024, a report from the NTIA highlighted the need for increased standardization of auditing methods coupled with routine public disclosures to improve the integrity and validity of AI audits. Work towards this goal has continued, with auditors zeroing in on important minimum standards for embedded audits. This is a longstanding challenge within the field, and there is a role for the federal government to play in establishing common benchmarks, standards, and evaluation processes for AI evaluations.
The federal entity that is currently best equipped to centralize the level of technical expertise needed for industry-level standards setting is CAISI, and this proposal would provide them with both the mandate and funding to craft consistent methodologies that can be applied at scale. This would help to address questions about the validity, reliability, and repeatability of many evaluation frameworks, benchmarks, and testing methods and would additionally provide an improved regulatory standard to hold companies accountable for the distribution of unsafe digital products.
Minimal competition exists within the auditing ecosystem.
Strengthening competition between firms within the AI auditing and verification ecosystem should be another priority. A consistent federal funding source administered through CAISI could lower barriers to entry and support new market participants, including organizations capable of bringing greater scientific and technical rigor to the field. A larger and independently resourced pool of potential auditors would build a stronger professionalized community of practice, in which firms and regulators alike can scrutinize methodologies, have replicable findings, and independently verify one another’s results. This is particularly important as the field evolves from a nascent ecosystem into an established professional industry in its own right. Without intentional efforts to support competition and prevent the premature narrowing of the field now, the market risks consolidating too early around a small number of industry-aligned firms, concentrating both market power and influence over emerging AI auditing and evaluation practices. Encouraging a diverse and competitive marketplace of independent assurance providers through intentional market design decisions can reduce the likelihood of industry capture, strengthen auditor independence, and improve the credibility and long-term resilience of the AI auditing ecosystem.
AI auditing is reaching an inflection point for policy intervention.
In January 2026, a group of leading researchers published a paper that describes: a vision for frontier AI auditing as a discipline, four AI Assurance Levels (AALs) for audits, a set of four policy goals necessary to realize the paper’s vision, and a variety of recommendations and research questions—many of which point directly at the challenges discussed above. In addition, organizations in the AI evaluation, auditing, and assessment ecosystem have become increasingly organized in the last year, with highlights including the creation of the AI Evaluator Forum in December 2025 and PACT AI in August 2026.
Legislators seem to have taken notice. There are numerous enacted state laws, and federal proposals for AI governance that require or encourage AI auditing in some form. This growing alignment inside the community of practice, along with an increase in demand driven by a steady drumbeat of high-profile AI safety failures by developers building frontier models, presents an opportunity to establish a federal system that addresses a variety of the existing challenges.
Existing examples and analogies.
Numerous existing models of regulation and private industry oversight leverage federal oversight in collaboration with private industry to advance safety and good governance in other domains. These existing models provide both cautionary tales and success stories to adapt to the particular needs of AI evaluation and assessment. One such useful precedent for independent product-safety verification comes from vehicle safety testing and the model established by the Insurance Institute for Highway Safety (IIHS). The IIHS is a private organization funded by insurers to ensure greater insights into vehicle safety in addition to testing and attestations required by federal regulators such as the National Highway Traffic Safety Administration. Similarly, under the Prescription Drug User Fee Act, the U.S. Food and Drug Administration collects application fees from pharmaceutical companies to support the prescription drug review process, helping ensure that medical products are thoroughly evaluated for safety and effectiveness before they are approved for the market.
The securities industry potentially offers the closest structural parallel. In the wake of accounting scandals in the early 2000s, the Sarbanes-Oxley Act created the Public Company Accounting Oversight Board (PCAOB), a private nonprofit that registers and inspects the firms auditing public companies, with the SEC holding approval authority over its rules, standards, and budget. Its funding does not come from Congressional appropriations but from an annual fee assessed on the issuers whose financials are audited. The Financial Industry Regulatory Authority (FINRA) operates on a similar logic for broker-dealers in the financial services industry. In both cases, Congress determined that a technically specialized oversight function could be housed outside a federal agency so long as the agency retained final authority over what that body does.
Securities law contains the closest existing precedent for the assignment mechanism proposed here. Section 939F of the Dodd-Frank Act (commonly called the Franken Amendment) directed the SEC to study a system in which a board, rather than the issuer, would assign credit rating agencies to rate structured finance products, and to implement such a system unless it identified a better alternative. The provision responded to precisely the conflict at issue in AI auditing today: when the rated party chooses and pays the rater, the rater’s commercial interest runs against candor. The SEC studied the question and declined to build the board but support for random assignment remains.
Additional benefits of the clearinghouse model.
Federal clearinghouses are a key component of advancing evidence-based policy. The clearinghouse model is typically utilized by federal agencies seeking to develop a centralized mechanism to collect, store, and share data from external sources. In this instance, the use of a clearinghouse for AI audits would provide a streamlined way for model developers and auditing firms to route audit results to a single centralized federal clearinghouse. A public repository of audit results (with sensitive or proprietary information redacted) would enable for greater transparency, and would allow for open-source fact checking of audit results. The cross-agency search interoperability that is typically afforded to federal data clearinghouses means that this data could then be shared with other federal agencies as needed in alignment with other federal information sharing apparatuses.
Plan of Action
Recommendation 1: Congress should statutorily authorize and fund CAISI and an Independent AI Evaluation Clearinghouse (IAIEC).
Congress should pass a bill to statutorily authorize and formalize CAISI while directing substantial funding to ensure it can carry out its already significant mandates.
That same bill should also authorize and fund an IAIEC. The IAIEC should be a non-profit organization, similar to FINRA or the PCAOB, where CAISI has ultimate oversight and approval authority for all rules and standards.
Congress should ensure the statutory mission of the IAIEC has four foundational parts:
(1) oversee the accreditation of independent AI assurance organizations;
(2) act as a clearinghouse for user fees that sustain an auditing ecosystem;
(3) serve as a neutral intermediary that pairs auditors with AI developers seeking audits in an independent fashion;
(4) facilitate the collection and publication of redacted audit and evaluation results in a centralized repository for the public, and for other federal agencies.
Functionally, IAIEC will serve as an “auditor of auditors” to devise, operationalize, and oversee a system of accreditation for AI assurance organizations, with ultimate approval vested in CAISI. Individual AI assurance organizations or Independent Verification Organizations (IVOs) can interact with this entity to seek accreditation and thereby access the federal auditing ecosystem.
AI developers seeking auditing from this ecosystem will pay a fee for access. The fee schedule will be set in advance and scaled based on the size of the developer, level of assurance sought, and other factors necessary to ensure the sustainability of the overall ecosystem and serve the public interest.
The IAIEC will assign accredited auditors, approved for the requested assurance level, to complete requested audits as an intermediary using a lottery, rota, or other non-discretionary system. This removes any decision-making from the clearinghouse, and also prevents AI companies from selecting their own auditors.
This assurance ecosystem proposal does not include any mechanism for legal mandates or enforcement, nor does it contemplate what the appropriate scope of any such regulation ought to be. This simply describes the bureaucratic and institutional machinery necessary to develop a more sustainable, robust, and independent AI assurance ecosystem.
Developers could voluntarily seek accredited evaluators; be incentivized by investors, directors, insurers, or others; or be required by future legal mandates. However, to further incentivize AI companies participation in this auditing process, the federal government should condition access to procurement processes and the FedRAMP marketplace on securing audits from this federal ecosystem.
Table 1. Components of the Independent AI Evaluation Clearinghouse (IAIEC)
| Function | Proposed mechanism |
|---|---|
| Standards | CAISI researches and approves final standards for evaluation protocols and minimum methodological requirements for auditors. |
| Accreditation | IAIEC oversees an accreditation process and registry for AI assurance organizations, including a tiered assurance levels. |
| Funding | Developers pay scaled user fees into a pooled fund managed by IAIEC. |
| Assignment | IAIEC assigns accredited auditors to AI developers that request audits through a non-discretionary system. |
| Payment | IAIEC pays auditors from the pooled fund once work is complete. |
| Oversight & Monitoring | IAIEC audits assurance organization performance, investigates conflicts, assists in information sharing with relevant regulators, maintains auditor conduct standards, and can suspend accreditation, subject to review by CAISI. |
| Transparency | IAIEC publicly publishes standards, accreditation criteria, aggregate anonymized results, and provides enforcement information. |
Recommendation 2: Use CAISI and the IAIEC to adopt a formal accreditation process for AI auditors, including tiered and specialized certification to accommodate broad participation.
The need for robust, reliable, and rigorous AI auditing standards that are widely accepted is one of the central challenges of the growing AI assurance field. An accreditation system with federal oversight is the best way to rapidly focus efforts on research and standard setting, and ensure broad adoption. A formal accreditation process overseen in part by a federal agency would establish a higher baseline of standards for all AI auditing firms. This would include requirements governing auditors’ access to information and the level of access or integration within AI developers necessary to conduct meaningful evaluations.
This memo’s CAISI-IAIEC proposal operationalizes Recommendations 2 and 3 from the Frontier AI Auditing paper. It will create a centralized point for industry advocates to engage with, and the clearinghouse model allows for the easy integration of information from state-level auditing regulators as they emerge. This process is also compatible with, and can directly leverage, models like an Independent Oversight Marketplace that will push good certifications to the top of the field. Ultimate oversight and approval resting with CAISI is recommended because of its placement within NIST, and NIST’s expertise in convening and navigating this kind of standard setting, as well as its fluency in engagement with the research community, especially in this nascent period of the development of the field. However, CAISI and NIST do not represent the same strong, independent oversight culture or function as the SEC does in the context of FINRA or the PCAOB. Ultimately, the oversight role of CAISI could, and really should, ultimately be moved to an independent digital regulator. Research communities like the EvalEval Coalition will be necessary contributors to ensure that the accreditation process is not merely a rubber stamp for the current best practices, but that accreditation pushes the rigor and reliability of auditing standards forward.
Accreditation should also be divided into tiers and should account for specialization. At minimum, the tiers should be aligned with both AALs and the bandwidth and capabilities of a given auditor. Creating tiers of approved auditors would help ensure opportunities for participation by organizations of differing sizes, capabilities, and areas of specialization.
In addition, public sector oversight of accreditation and the associated standards provides a surface area to advance assurance in areas besides frontier safety risks. While much of the current conversation currently focuses on existential and catastrophic safety risks, the AI auditing ecosystem has its roots in testing for bias, and should be expanded to include emerging harms like consumer manipulation, privacy risks, and other categories of concern.
Recommendation 3: Ensure sufficient start-up funding for CAISI and the IAIEC and then leverage scaled user fees to create a sustainable funding environment for AI auditors.
First, it should be noted that additional direct appropriation funding for this work will be necessary up front. In order to properly carry out this accreditation function, both CAISI and the newly-formed clearinghouse will need considerable technical staff resources. The oversight of the accreditation standards in CAISI and the management of the accreditation process in the IAIEC will require additional staff with AI technical expertise. The financing should come from a combination of congressional funds to kickstart and independently sustain the establishment of standards, with some operating revenue coming from the fees paid to access the auditing process. This could resemble the governance structure of the CDC Foundation while being modified to support the mission of CAISI.
But the real goal is to develop a thriving AI assurance field for the AI auditors and developers. A sustainable funding mechanism for the third-party AI assurance ecosystem that does not compromise independence is necessary for the maturation of the field. Currently, existing auditor organizations are reliant on philanthropic funding, research grants, or revenue directly from industry clients. Many AI-focused auditing organizations are relatively small, and they face significant salary pressures when hiring technical staff due to competition from AI companies. This proposal for a clearinghouse-managed fund, with scaled fees to allow broad developer access, is necessary for tapping into AI industry revenues to sustain the auditing ecosystem. Ideally, payments issued to participating auditors will be sufficient to more sustainably support organizations that choose to undertake this work.
The core mechanism is that parties seeking audits will pay a fee into a fund managed by the IAIEC, and auditors will in turn be paid from this fund for completing audits. The goal of the fee and payment schedule, which will be set by the IAIEC with final approval from CAISI, should be to maximize size and participation in the ecosystem on both sides.
On the developer side, the audits will be provided at no direct cost to the developers, and because participation is voluntary and involves receiving a defined regulatory service (“structured safety auditing”), the payment is properly characterized as a user fee rather than a tax. The model is operationally analogous to the FDA’s user fee programs, in which regulated entities fund part of the oversight infrastructure in exchange for access to regulatory review services. Similar to the FDA’s Prescription Drug User Fee Act (PDUFA) program, the program would operate on a sliding fee scale and could even provide fee waivers to allow participation by public interest model developers. Fee rates will be determined by information such as research budget, revenue, model size, and assurance level sought.
On the auditor side, accredited organizations can be assigned audits and will be compensated per completed audit from the fund, with payment amount dependent on the services rendered, rather than the fee paid by the organization being audited. The IAIEC will publish transparent and concrete requirements for receiving payment, along with schedules to guarantee timely reimbursement.
Recommendation 4: Use the IAIEC as an intermediary to assign audits to auditors in a neutral, independent, capture-resistant manner designed to maximize auditor independence and prevent conflicts-of-interest.
The last component of this proposal is that the IAIEC be used as an intermediary to assign auditors to audits through a non-discretionary process that is backed by clear conflict-of-interest rules and prohibitions. Managing accreditation and funding through the same clearinghouse allows for random or rotational assignment of auditors while maintaining assurance that matched parties will be able to both render the appropriate services and be compensated fairly. This feature has not been present in previous legislative proposals for AI auditing, and other policy proposals for addressing industry self-regulatory concerns tend to focus on the possibility of funding a full government auditing regime instead. The hybrid model proposed here of mediated assignment leverages the speed and flexibility of a third-party auditing regime while removing the choice of auditor from industry hands.
It is important that the mechanism of assignment be non-discretionary, in other words, the IAIEC staff should not be in the position to exercise judgement about how to assign audits. Discretionary assignment or a bidding process would only reintroduce avenues for favoritism, conflicts of interest, and regulatory capture. Instead a lottery or rotation through a list of use-case dependent accredited and qualified auditors should be used. Depending on the size of the available pool, to prevent gamification or shopping, even a rotation system should be randomized to an extent, using a system of random draw without replacement to ensure work is spread as evenly as possible. Non-discretionary assignment of auditors, however, is not sufficient on its own to guarantee independence of auditing and evaluation work. The IAIEC should also establish and enforce clear conflict-of-interest standards governing which auditing and evaluation firms are eligible for a particular assignment. This would work to rebuild public trust in the validity of their outputs.
There are several concrete advantages to this system: it will eliminate industry’s ability to shop for favorable audits, free auditor organizations from the need to maintain comfortable relationships with their subjects for access and revenue, and allow new auditors who can meet accreditation requirements to enter the market on the same footing as established auditor organizations. In addition, the non-discretionary nature of the assignment protocol combined with brightline rules to restrict conflicts-of-interest makes it highly resistant to subversion or capture—by both AI developers and auditors—at the clearinghouse level. This is a mechanism for true independence for third-party auditors, and AI developers will have assurance that the auditing ecosystem will not be captured either.
Conclusion
Independent AI auditing will only be as credible as the institutions, incentives, and standards that underpin it. As the assurance ecosystem matures, policymakers have an opportunity to utilize existing government institutions to engage in marketcraft to build meaningful independence into the AI auditing and evaluation market from the outset. This would require removing the issuer-pays incentive structure, creating brightline restrictions on conflicts-of-interest seen in other regulated industries, and establishing unified federal standards for AI auditing and evaluation benchmarks, methodologies, and processes to build public trust.
The proposed CAISI-IAIEC framework would operationalize this approach, introducing the oversight and institutional safeguards necessary to strengthen confidence in the rigor and reliability of AI audits and evaluations. The result would be a system that preserves meaningful third-party independence while maintaining necessary federal oversight without requiring the government itself to conduct every audit.
This does not mean that CAISI will need to remain the permanent home for this essential oversight function. In the event that a new digital regulatory agency is established, the final approval of rules and oversight for the IAIEC entity could, and likely should, be placed within that newly established agency.
For purposes of this proposal, “assurance” is the broadest term and refers to the larger ecosystem of activities and organizations that provide independent information about AI system performance, risks, and safeguards. An “audit” is a structured assessment of an AI system, developer practice, or claim against defined criteria. “Evaluations” and “verification” can constitute components or forms of an audit: evaluations measure particular capabilities, behaviors, or risks, while verification seeks to independently substantiate specific claims of results made by a model developer.
Depending on the standards approved by CAISI, auditing and evaluation organizations accredited under this federal framework could conduct pre-deployment testing, measure performance against standardized benchmarks set by CAISI, independently verify safety claims made by developers, or evaluate systems for specific risks. Those risks could include but are not limited to: frontier concerns such as cybersecurity, biosecurity, or other national-security capabilities, biased or unfair outcomes, disparate access to credit or other services, privacy risks, deceptive or manipulative design, and other potentially harmful interactions with users.
In contrast, the public clearinghouse function addresses the problems posed by the issuer-pays model and also provides the public and other government actors with centralized repositories of audit information. Appropriately redacted results could be made publicly available to improve transparency, increase public trust, and enable outside scrutiny from the broader research community. In addition, authorized federal and state regulators could use more detailed information for enforcement, comparison, and cross-referencing as needed.
This could resemble existing federal repositories of information such as the Federal Trade Commission’s Consumer Sentinel Network or the Consumer Financial Protection Bureau’s Consumer Complaint Database. These systems share information across jurisdictions, and are a key component in the enforcement of existing consumer protection laws.
The CAISI-IAIEC model does not supplant market-based standards solutions, and is likely to benefit from consultation with IVO, academic, and industry standards setting processes that are underway. The IAIEC would serve as a centralized point for consensus-building, provide government oversight analogous to what is provided for other credit ratings agencies, and can serve as an important baseline and backstop as the assurance field continues to innovate and adapt.
The proposed IAIEC could also establish fee waivers or reductions for qualifying academic institutions, public-interest research organizations, or other resource-constrained public-interest institutions.
We estimate an initial outlay of approximately $180 million over three years, with the bulk being a first year outlay of $100 million. This should be paired with a statutorily authorized CAISI baseline in the range of $85 million annually which CAISI needs to robustly fulfill its existing mandates, based on assessments from independent analysts associated with Federation of American Scientists and Institute for Progress.
The first-year figure includes roughly $40 million for IAIEC operations (about half the PCAOB’s first-year budget, reflecting a narrower regulated population), roughly $12 million for new CAISI standards and oversight work, and a one-time $50 million seed fund. The seed fund would cover initial fee waivers, accreditation support for new entrants, and general operation needs. Following the model of the SEC’s startup advances to the PCAOB, the seed funding should be structured as a repayable advance recovered through future user fees.
Because participation is voluntary, the pace at which fees replace appropriations will depend heavily on developer and deployer participation. Congress should expect to sustain IAIEC operations through appropriations until fee revenue is demonstrated.
The clearinghouse will also neutrally mediate the assignment of auditors to audits—strengthening independence, trust, and ethical governance across the AI ecosystem.
Soft law was never meant to be a permanent solution. Treating it as one, and letting the sandcastle stand in for the skyscraper indefinitely, is how we end up with a decade of voluntary commitments and no enforceable accountability to show for it.
Jobseekers, employees at smaller organizations, and local public workers need more than a weekend crash course in AI. They need practice choosing useful tasks, protecting data, testing outputs, and explaining where human judgment remains necessary.
“What excites me is that it’s very tempting to be very discouraged, and say, ‘Oh, we’ve got these archaic institutions that are calcified and you could never change them.’ But I think we’re in the middle of a technological revolution that will upend lots of things, and does provide a window.”