Strengthening EO 14409’s Approach to Frontier AI Through Independent Assessment and Published Criteria for Trusted Partners
The Executive Order on Promoting Advanced Artificial Intelligence Innovation and Security (EO 14409) marks a significant shift in the federal government’s approach to frontier artificial intelligence (AI). While keeping with previous efforts to outline voluntary regulation and prioritize American innovation, the EO establishes a new classified government assessment framework for the cyber capabilities of AI models. It also seeks to work with private industry to select “trusted partners” eligible for early access to frontier models.
Both the assessment and trusted-partner designation processes are internal processes of government agencies. Section 3 tasks a small group of national security actors with building a classified benchmarking process for measuring advanced cyber capabilities and identifying “covered frontier models” that meet the classified threshold. While the EO provides no public definition of the term “covered frontier models”, the designation of the title is significant because it determines which models participating developers are then asked to make available to the government for up to 30 days before they are first released to trusted partners.
However, the framework designates overlapping authority to the Director of NSA to both set the standards and make the final determination as to whether models should be designated as a covered frontier model under those standards. Grouping the roles of standards-setter and verifier creates a concentration of authority with no outside validation, which weakens the framework’s credibility because there are no parties equipped with relevant technical expertise to catch errors. The trusted-partner title does not have any established criteria, and therefore an entity’s eligibility cannot be properly assessed.
While preserving its original goals, the EO 14409’s classified benchmarking and trusted-partner selection processes can be strengthened. First, by the Department of Defense (DOD) designating qualified Federally Funded Research and Development Centers (FFRDCs) as third-party assessors of frontier AI evaluation standards. Second, by designating the National Institute of Science and Technology (NIST), in consultation with the Directors of National Security Agency (NSA) and Cybersecurity and Infrastructure Security Agency (CISA), to publish a set of function-based criteria defining qualifications for the “trusted partner” status.
FFRDCs such as MITRE National Security Engineering Center (NSEC), The Institute for Defense Analyses (IDA) Center for Communications and Computing, Massachusetts Institute of Technology (MIT) Lincoln Laboratory, and Carnegie Mellon University (CMU)’s Software Engineering Institute (SEI) already have past AI and cybersecurity project portfolios and existing working relationships with the very agencies that EO 14409 tasks with benchmarking. Because FFRDCs’ sponsoring agreements restrict outside work, these centers do not rely on funding provided by developers whose model they would assess, and therefore hold no stake in whether a given model is designated as frontier AI. Final authorization would remain decided by cleared federal agencies, mirroring the Federal Risk and Authorization Management Program (FedRAMP) procedure.
Inserting an assessor between developers and the deciding agencies adds required time and resources. That cost is manageable because the named FFRDCs already operate under contracts with the DOD and maintain secure facilities and technical staff that evaluation requires.
The EO’s “trusted partners” designation, which grants access to the most technologically advanced AI models to critical infrastructure firms, has no public and explicit criteria for qualification. Implementing a transparent trusted-partner standard is necessary. It would not only mitigate criticisms of opaque membership criteria that the CISA’s Joint Cyber Defense Collaborative (JCDC) faced, but also would encourage private industry to strive towards meeting standards and fostering a collaborative frontier AI industry.
AI’s high-stakes outputs directly shape American economic, industrial, and technological outcomes. They must not be accepted entirely on the basis of institutional faith rather than verifiable evidence.
Challenge and Opportunity
Whether a model is designated as a covered frontier model determines whether its developers are asked to give the government up to 30 days of pre-release access and to release it to trusted partners first. Under Section 3 of EO 14409, the final authority to make this designation rests with the Director of NSA alone, with only consultation with other government actors. The classified benchmarking process that establishes the threshold the models are checked against is developed by the Secretaries of Treasury, War (through the Director of NSA), and Homeland Security (through the Director of CISA), in consultation with the National Cyber Director, the Assistant to the President for Science and Technology, and NIST. NSA is therefore given both a role in setting the classified benchmarks and the authority to make the final call on which models meet them.
Grouping the roles of standards-setter and verifier together creates a concentration of authority without independent validation, which weakens the credibility of the framework as there are no outside parties equipped with the relevant technical expertise to correct errors. A verification process that is not well supported with technical knowledge and opinion that holds no stake in the final designation outcome could cause significant national security risks in two areas. First, a model may be designated as a covered frontier AI model when its capabilities do not warrant it, causing significant cost for developers in delayed timelines and constraints upon release, and for government actors in spent resources with no corresponding security benefit. Second, a frontier AI model capable of creating significant national security impacts may pass without designation, leaving it to be deployed without the needed safeguards.
The absence of external technical review also weakens the incentive for developers to participate. A voluntary framework succeeds when participants have something to gain from engaging it. As the framework currently stands, developers cannot examine what is being tested, as the benchmarks are classified, and cannot verify whether the benchmarks are being applied consistently across firms, because verification is internal. The benefits that come from engaging the framework, such as reputational benefit or a working relationship with government agencies, may not outweigh the time and resources it needs. As a result, developers may choose not to participate in the voluntary framework, which will become ineffective when government cannot evaluate the models it is most concerned with.
DOD has the opportunity to outsource verification of classified AI benchmarks to FFRDCs. Federally Funded Research and Development Centers (FFRDCs) are research institutions typically administered by universities or non-profits that are funded by federal support to provide additional R&D capabilities and independent scientific perspectives. In the past, when the government needed to consolidate voluntary agreements into audited, credible requirements, it has collaborated with FFRDCs for technical capacity it couldn’t support by itself.
For instance, the Cybersecurity Maturity Model Certification (CMMC) came from the collaboration between DOD and the CMU Software Engineering Institute (SEI). The CMU SEI FFRDC drew on its decades of cybersecurity expertise to co-develop a certification structure, assessment standards, and train over 160,000 contracting workers in the defense acquisition workforce.
There are multiple existing FFRDCs conducting highly classified work under the major actors (NSA, NCD, APST, CISA, and DOD) named by the EO that would serve as reliable third-party assessors. (See Table 1.) Together, these institutions have conducted classified cyber and machine learning work in areas the EO identified as major priorities of frontier AI risk in Sec. 2(d), such as detecting and patching software vulnerabilities. All four FFRDCs are sponsored by agencies the EO has tasked with implementation.
Delegating frontier AI assessment to FFRDCs increases government technical capacity, since these institutions already maintain secure infrastructure and dedicated personnel with relevant expertise. By directing major frontier AI projects to FFRDCs, the government also reinvests funding and time back into American scientific personnel, which in turn spurs further technological innovation and produces reusable capacity for AI tasks that come after EO 14409.
FFRDCs’ unique status as a non-profit with access to private-sector resources required to “operate in the public interest with objectivity and independence” allows them to sit outside of government while working with it. Because FFRDCs operate under sponsoring agreements that restrict the amount of outside work they can accept, and because frontier AI vendors cannot control FFRDCs’ revenue as their funding is directly sourced from their sponsoring agencies, FFRDCs do not hold a stake in the outcome of the benchmarking process. Furthermore, since FFRDCs are required to “fully disclose their activities to their sponsoring agency,” security agencies will still be able to maintain close monitoring of activities with sensitive data and preserve safeguards against dangerous actors that the EO seeks to establish.
The effectiveness of differentiating the roles of standards-setter and assessors has been proven effective under previous federal practice with FedRAMP’s accredited third-party assessment organizations (3PAOs), or assessors, as the current FedRAMP Consolidated Rules for 2026 refer to them. Under this system, assessors conduct independent security assessments of commercial cloud service providers (CSPs) and document Security Assessment Reports (SARs) detailing how well services meet federal standards. However, authorizing officials (AO) from federal agencies get the final say in the risk decision. Therefore, the current administration can designate FFRDCs a similar role and separate verification structurally from the final authorization of frontier AI models. Doing so guarantees that FFRDCs provide additional technical verification against established benchmarks while retaining the final say in the hands of trusted federal agencies.
Additionally, as EO 14409’s Sec. 3(b)(ii)-(iii) currently lays out, industry and the federal government will collaborate to select “trusted partners” who gain early access to the most powerful frontier models. The EO explicitly states that selecting trusted partners serves to “promote secure innovation and strengthen the cybersecurity of critical infrastructure.” Early access is therefore treated both as a risk to be managed and a benefit to be allocated. The emphasis on “secure innovation” suggests that powerful models must be placed in the hands of entities trusted to use them responsibly, since each additional user of early-access models adds more risk of the model being leaked. At the same time, these tools must reach the frontline, critical infrastructure defenders who would otherwise lack access to tools of similar capability.
Although the general purpose of selecting trusted partners is stated, there is no definition of who could serve as trusted partners or what specific qualifications make an entity eligible for this title. The absence of published, concrete criteria creates conditions that allow regulatory capture, whereas the very regulatory rules made to protect public interest may become a tool to push industry interests. The government will effectively be picking the private firms that “win” the market by selecting which firms receive the commercial advantage of the most powerful frontier models. Without explicit, transparent rules, the allocation of this title may become based on preexisting relationships with government or industry actors rather than real need and qualification, leaving smaller critical-infrastructure firms that could benefit from having early access at a disadvantage.
It is imperative that change is enacted promptly to minimize the resource and time cost of amending the verification process once it has been embedded into agencies’ operations. NIST can publish a publicly accessible set of qualifications for “trusted partners” before the first round of assessments is made and before trusted partners are selected.
Plan of Action
Recommendation 1. Establish FFRDCs as the third-party assessors of AI models against classified benchmarks
DOD should task FFRDCs as third-party assessors who will check submitted AI models against classified frontier AI benchmarks through existing sponsoring agreements under the FAR 35.017. Because their sponsoring contracts restrict outside work and competition with private industry, FFRDCs cannot be funded by the private companies whose models they may be asked to assess. Additionally, existing contracts with the most central agencies identified in the EO mark the opportunity for a low-burden process to map tasks onto an existing structure between FFRDCs and DOD.
In reference to FedRAMP’s assessment model, FFRDCs would not be the final decision-makers on whether a frontier AI model exceeds the threshold set by the established benchmarks. Their role would be limited to providing strictly technical evaluation and reporting recommended actions back to agencies the EO already designates as assessors, such as the NSA and CISA. Ultimately, authority would remain where the order previously placed it with cleared national security actors, but the addition of FFRDCs will ensure that the technical work behind the most influential decisions in AI and cybersecurity today gains more credibility and rigor. While outsourcing the verification process would likely lead to additional time for interagency coordination, the benefits of increasing credibility, as well as the act of reinvesting in fostering the scientific talent around AI technology in America, far outweigh these considerations.
A complete pipeline connecting FFRDCs to the sponsoring agency and NSA could be established as follows:
- Benchmark development. The same group of actors identified in EO 14409 Sect. 3—including the Secretaries of the Treasury, War (through the Director of NSA), and Homeland Security (through the Director of CISA), as well as corresponding consultants—develops a set of classified benchmarks on what qualifies as a “covered frontier model.”
- Task Assignment. AI developers voluntarily engage the Federal Government in evaluation of their model. DOD issues a new task order to the appropriate FFRDC while developers provide model access directly to the FFRDC under confidentiality and cybersecurity protections.
- Evaluation Execution. The FFRDCs assess model capabilities against the classified benchmarks within secure computing environments. Projected areas of testing may include cyber-specific capabilities that frontier models present, such as the ability to discover and exploit cybersecurity weaknesses, and whether safeguards hold against repeated jailbreaking. However, the precise areas of interest would be tailored to the benchmarks being measured against. Evaluation is confined within isolated environments, and full evaluation logs and code versions are preserved to be combined into the final report.
- Reporting. FFRDCs provide a classified report to the sponsor agency, outlining whether the model “exceeds threshold,” “undetermined,” or “does not exceed threshold,” to each of the set standards. The report should be completed with methodology, results, as well as model limitations.
- Final Decision. NSA Director, referring to the FFRDC evaluation record and set benchmarks, makes the final decision as to whether the AI model should be designated as a “covered frontier model.”
To expand upon existing technical capacity and map new AI verification responsibilities to these FFRDCs, further funding should be secured. Given that this memo’s proposed FFRDC assessors are under DOD sponsorship, DOD would be the ideal actor to provide additional funding. Establishing a firm estimate would require a formal cost estimate by DOD and the corresponding FFRDCs. As a starting point, policymakers should plan for roughly $30-75 million in the first year for facilities, initial hiring, and compute, and $20-50 million per year onwards for ongoing costs such as salaries and maintenance. If full launch funding cannot be identified, DOD may instead opt for a pilot year to identify a smaller, one-year program. After the full run-time of the pilot year, agencies may then choose to renew the program and continue having FFRDCs’ technical expertise on board.
Ultimately, this recommendation aligns with the national priority to accelerate American AI innovation. By reinvesting in America’s scientific talent through delegating funding and more expansive AI projects to our federally funded laboratories, we can create a positive feedback loop that propels America further as a global technological leader.
Recommendation 2. Develop and publish a set of publicly available function-based criteria to qualify as “trusted partners”
To ensure verifiable eligibility standards and minimize government control on the market for frontier AI, NIST, in consultation with NSA and CISA, should develop and publish a set of function-based criteria that outlines what will qualify as a “trusted partner” eligible for early access to covered frontier models. This set of standards should be published after a period for the public and private industry to submit comments to fully assess current industry needs and capabilities.
Undefined criteria for membership have been proven to be ineffective in previous government collaborations with the private industry. CISA’s Joint Cyber Defense Collaborative (JCDC) was established in 2021 to secure cyberspace by bringing collaborators from the government, industry, and international organizations, but it never published a set of clear qualifying criteria for membership. Its first cohort was drawn largely from the largest cloud and security vendors, including Amazon Web Services, AT&T, CrowdStrike, FireEye Mandiant, Google Cloud, Lumen, Microsoft, Palo Alto Networks and Verizon. But the ambiguity around what qualifies an entity for membership led to CISA’s own Cybersecurity Advisory Committee (CSAC) to advise CISA to clarify and provide more transparency on membership requirements and joining processes. Moving forward with the implementation of EO 14409, it’s imperative to learn from this past example that published requirements are necessary.
NIST should be the leading actor for this effort. First, its Center for AI Standards and Innovation (CAISI) is already charged under the AI Action Plan with developing evaluation guidelines for federal agencies and has meaningfully coordinated efforts on pre-deployment standards with major companies such as OpenAI and Anthropic. NIST has also previously constructed and published the AI Risk Management Framework that established publicly accessible and uniform technical vocabulary as well as voluntary guidelines for AI systems. NSA and CISA should act as consulting agencies and contribute guidance on security and clearance requirements.
This set of function-based criteria might include:
- Demonstrated and Significant Mission. The entity has demonstrated significant and positive contributions to a critical infrastructure sector and can articulate a specific function that early access to frontier AI models would serve.
- Security Capacity. The entity is equipped with sufficient security capacity to protect access to models.
- Matching Public Benefit to Commercial Advantage. The entity’s commercial advantage derived from being designated as a trusted partner should match the critical-infrastructure security benefit it would produce.
- Accountability and Auditability. The entity has accepted responsibility to oversee and log its usage of frontier AI models.
The drafted criteria should explicitly state that a commercial relationship with the developer, while not disqualifying, cannot serve as the basis for the designation as a trusted partner. Additionally, designated trusted partners should be re-evaluated periodically to ensure that the entity continues to align with published criteria. The estimated cost for this recommendation is a one-time $2.5 million to support a NIST drafting team (of five to ten staff), public comment processing, and consultation with NSA and CISA. Ongoing future monitoring and revision of the criteria may cost under $1 million per year.
All critical infrastructure firms stand to benefit from this approach. A set of defined and functional standards, rather than an omitted list of names, will allow any existing firm to assess its own eligibility and work toward meeting it. As a result, this recommendation would contribute to building a collaborative frontier AI industry that benefits from its working relationship with the government.
Conclusion
While EO 14409 is an important first step toward creating a national response to frontier AI development, further measures are needed to reinforce the credibility of the AI evaluation process and define the currently unspecified “trusted partner” criteria.
Under the recommendations of this policy memo:
- DOD would outsource the 30-day pre-deployment assessment process to FFRDCs, separating evaluation from the office that authored the benchmarks, while incorporating technical capacity and leaving standards-setting and decision-making responsibilities to the original actors under EO 14409.
- NIST, in consultation with the NSA, would write and publish function-based criteria for what qualifies an entity as a “trusted partner” and obtain access to the most technologically advanced frontier AI models.
Implementing these changes would strengthen national technological capacity, reinvest in American scientific talent, and foster a collaborative frontier AI field.
This proposal orients itself around improving accountability and transparency. These recommendations address two pressing American AI concerns: one, reinvesting in American scientific talent by equipping FFRDCs with relevant AI experience and funding, and two, promoting an equitable frontier AI market without additional regulations that may increase burden and slow innovation.
Private industry may express concern over the concentration of authority and the absence of external technical review. They may believe this leads to an ad-hoc process of AI evaluation and designation that is biased. However, having clear and published trusted partner criteria and an evaluation process promotes fair partnership and spurs healthy competition. That opens participation to smaller and newer firms (i.e. startups and middle tech).
Implementation would be built upon an existing contract and relationships between DOD and its sponsored FFRDCs rather than requiring a new group of actors to execute. Accountability would be measured by DOD, which acts as the sponsoring agency, receives all direct reports, and oversees task order. The Director of NSA then receives the final report and uses it to make a documented and well-supported designation decision. The publication of trusted partner criteria will be verifiable via the draft released during the public comment period by the projected timeline and the final, publicly accessible document. .
Although both the Director of NSA and the broader DOD will remain central actors in the policy memo’s recommendations, the conflict of interest is reduced because the evaluation records are produced outside of the office making the final decision. Since FFRDCs are not the final decision makers and only serve to provide technical evaluation reports, their role remains as an independent actor rather than an entity that could potentially gain from swaying frontier AI delegation decisions.
DOD budgets FFRDC work through assigning staff years of technical effort (STE) (see: GAO-20-31, FEDERAL RESEARCH: DOD’s Use of Study and Analysis) that represents the amount of resources required by one employee for one year. Congress caps the total number of STEs that DOD may allocate each fiscal year, and the current ceiling is set at 6,053 STEs, and Defense STE during the FY may be reallocated as needed. There is no publicly accessible record of the amount of STEs assigned to each DOD FFRDCs as well as the precise amount of STEs required to support respective research projects. A very general and hypothetical estimate of how many STEs can support a large-scale remapping of AI verification tasks onto the recommended FFRDCs would likely take an estimated 38-58 STEs.
Per the policy memo recommendations, the larger FFRDC evaluation process may be broken down to three areas that will require personnel and resource support. First, approx 20-30 STEs may be assigned to support conducting model evaluations, with approximately 15-30 evaluations per year starting off and expanded as needed. Second, a portion or approx 12-18 STEs may be needed to build secure evaluation infrastructure with appropriate GPUs and equipment, building isolated networks for red-team testing, etc. Finally, a portion of approx 6-10 STEs could be used to support efforts of writing and delivering evaluation reports to sponsoring agencies who will hold the power to make final audit decisions.
This amount of STE will likely support the number of needed technical evaluators at FFRDC rates, secure infrastructure, and increasing computation dedicated solely to supporting verifying frontierAI. However, as is not disclosed to the public, further input from involved FFRDC and DOD personnel is needed to precisely identify the amount of needed resources.
The verification role requires an assessor who is independent from developers and holds no stakes in the outcome. FFRDCs are the best fit given their access to government clearance, status as a non-profit, and secure facilities which ordinary private contractors and academic experts do not have. GAO’s concerns are based on oversight failures, noncompetitive contracts, mission creep, as well as contract management. These are governance and cost problems that do not significantly reduce FFRDCs’ technical capacity for such work or the quality of outputs.
FFRDCs are not perfectly insulated, but against the current layout of the EO 14409 where the verifier is also acting as the standards-setter, outsourcing to FFRDCs is still a significant improvement.
First, NIST CAISI operates on a very limited budget. There has been a lot of effort to advocate for an expanded budget, but it has not yet been granted. Due to constrained staffing resources, NIST CAISI may not be able to meaningfully produce the volume of model evaluations at the speed (30 days) required by the EO’s process.
Second, all major actors named in the EO are located from federal agencies operating behind very classified standards and processes. Therefore, it is unlikely that any responsibilities can be re-directed to public-facing organizations. Ultimately, it seeks to keep the authority within national security actors, which the NIST CAISI does not qualify for. Therefore, to find a meaningful compromise between national security and the need for transparency, FFRDCs are the most viable actors that allow ultimate authority to reside with central actors while adding technical capacity and credibility.