The U.S. Is About to Design an AI Regulator. Here’s How to Get It Right.
The AI industry and the administration could soon converge on a FINRA-style regulator for frontier AI. Its ultimate viability and effectiveness will be decided by design choices being made right now.

By experts and staff
- Published
Vinh NguyenCFR ExpertSenior Fellow for Artificial Intelligence
Elham TabassiSenior Fellow and Director of the Brookings Artificial Intelligence and Emerging Technology Initiative
Kat DuffyCFR ExpertSenior Fellow for Digital and Cyberspace Policy
Vinh X. Nguyen previously served as the National Security Agency’s chief artificial intelligence (AI) officer. Elham Tabassi previously served as the National Institute of Standards and Technology’s chief AI advisor. Kat Duffy is the director of LEAD AI at CFR.
Google DeepMind CEO Demis Hassabis published a framework on July 14 calling for a U.S.-led Frontier AI Standards Body. In it, he proposes that it be modeled after the Financial Industry Regulatory Authority (FINRA), the industry-funded, self-regulatory organization that oversees the brokerage industry under Securities and Exchange Commission (SEC) supervision. Three days after Hassabis published his piece, Bloomberg reported that the White House is reviewing a proposal—developed with Treasury Secretary Scott Bessent’s involvement—based on this FINRA template.
The good news is that the choice of institutional form was never the hard part: a FINRA-inspired body that reports to the SEC could be a reasonable answer to a real coordination problem. Under Google’s proposal, frontier AI companies would submit their models for review up to thirty days before release. Over that period, those systems would be tested for dangerous cyber, biological, and deceptive capabilities. The companies’ participation in this oversight process would begin voluntarily, then become a condition of U.S. deployment once the process has proven itself.
But that still leaves much undefined, such as how this body would be financed, whether its access to frontier systems is guaranteed or left to a lab’s discretion, and how its findings might need to be classified. The design choices made to address these problems could decide whether this new body becomes an institution that regulators and the public trust for decades, or one that will need to be rebuilt after its first crisis of confidence.
The Five Unresolved Questions
But these open design questions are not reasons to delay. They are simply design requirements that point to a specific architecture. How they are resolved—including the five below—will ultimately determine the proposed body’s long-term viability.
Independence. FINRA-style proposals assume industry funding and raise immediate concerns regarding governing incentives and conflicts of interest. The closest analogy to companies funding their own evaluators is the issuer-pays credit-rating model, which undermined public trust during the 2008 financial crisis. FINRA addresses this by requiring public governors to outnumber industry representatives, yet concerns persist about “composition drift,” in which shifting conflicts of interest among public-sector representatives can move the board’s collective focus away from public-sector priorities.
National security. No FINRA examination has ever produced information that was itself a national security asset. A frontier AI evaluation body will quickly and inevitably produce such findings; moreover, by holding prerelease access across every frontier lab, the body itself will become one of the world’s highest-value espionage targets. Mishandle this in one direction, and the body shifts from being a critical defender of public safety to a bespoke, intelligence-community subsidiary floating outside the intelligence community apparatus and its myriad protections. But overcorrecting in the other direction will cause capabilities with genuine security implications to go unexamined, meaning the United States could learn about them from an adversary first.
Legitimacy. Allies in Brussels, London, Seoul, and Tokyo will be reluctant to embrace an American industry-funded body as an arbiter of international standards, and the public would not trust a body that measures only what companies and security agencies fear, while its own concerns—jobs, privacy, inequality—go unexamined.
Access and talent. Independence without guaranteed model access leaves an oversight body free to evaluate only what it is allowed to see. Today, access is granted at a company’s discretion, under a non-disclosure agreement, and on a release-by-release basis. Such limitations inherently undermine the proposed body’s oversight capabilities and duties. Beyond access, the scarcest resource will be talent inside and outside of government able to test frontier capabilities in a scientifically valid way—expertise has increasingly consolidated within the private sector.
Measurement. There is no settled science of frontier assessment: benchmarks saturate, results carry unreported uncertainty, and point-in-time certification fails for systems that change between measurements. Assessments can be developed in isolation, but no broader accountability ecosystem or set of industry standards exists to translate their findings into practical changes for how to deploy frontier systems.
Design It Once, Design It Right
The following critical, initial design choices could meaningfully inform an oversight body’s durable effect through the AI industry’s continued progression.
First, separate three functions—safety specifications, infrastructure, and evaluation—structurally and financially at the outset. Any specifications defining what “safe” and “tested” mean should be funded exclusively by neutral, non-industry money—no dollar generated by an AI company should influence the process for determining the standards that will govern those companies. Instead, industry resources should be dedicated to building and maintaining shared testing instrumentation. Labs could fund common infrastructure through an AI engineering body such as MLCommons, whose mechanisms can be extended to serve every evaluator. These need not replace existing evaluation frameworks; they simply enable establishing a shared substrate that will ensure results across evaluators are comparable.
Meanwhile, no institution, public or private, has done frontier evaluation at scale; standing up that function is currently unsolved. This is why the model evaluations themselves should be conducted by institutions with proven commitment and capacity to support the public interest. For example, an evaluation system could pair the national security expertise and classified research capacity an institution like RAND has built over decades with established expertise in protecting civil and political rights that nonprofits such as the Center for Democracy & Technology have developed. Working in tandem, such organizations could coordinate the independent ecosystem that already exists among evaluators in governments, nonprofits, and academia. This is the structural answer to what board composition alone has never delivered.
Second, bind funding to access. Multiyear financial commitments and model access from the frontier AI companies—including access to internal, non-deployed systems—should be paired in any agreement, so that neither can be quietly withdrawn.
Third, design the classified interface before the first classified finding, not after. A small, cleared group within the intelligence community should enable government threat intelligence to sharpen test design without the government writing the tests. Tiered disclosure—unclassified ratings, lab-confidential findings, and classified capability details—should be established from day one, because national-security designations expand when left undefined.
Fourth, turn evaluation into a profession. Given the current asymmetry in talent and resources between the public and private sectors, frontier companies can offer technical talent to build shared instrumentation that will support the launch of this body. In parallel, construct a new, thoughtful talent pipeline: enable stable careers, scholarship-for-service recruitment (consider the intelligence community model built to counter the same private-sector pay gap), and conflict rules that make evaluation a profession rather than a waystation to frontier company employment.
Fifth, sequence honestly. Current measurement science is neither defined nor mature enough to support the level of confidence this transformational moment demands. Certify processes and continuous-monitoring regimes rather than certifying snapshots (i.e., a specific benchmark for a specific model, at a specific time, given how quickly models change). State plainly that any standards body’s early evaluations carry only provisional confidence, and charter the body’s responsibility to regularly revise its measurement discipline as the science and the evaluation ecosystem around it matures.
Lastly, embed allied governments and societal-impact evaluation into the architecture now, not later. Reserve seats for representatives of allied governments and for public interest organizations, and establish and protect funding for rights-based assessments, such as privacy and confidentiality. These are not concessions to critics; a certification or endorsement issued by a body that excludes the affected public and allied partners will not produce the effect necessary to achieve the body’s broader goals. Embedded representation is not a cynical play for political legitimacy. It is a critical investment into converting a technical finding into a trusted imprimatur, both within the United States and in the markets where American models must compete against cheaper stacks that reflect someone else’s standards.
Industry can fund and staff most of what is missing. What it cannot do—as leaders across finance, energy, and the labs themselves acknowledge behind closed doors—is to supply its own neutral convener. Someone has to draft the first credible blueprint, and that blueprint will establish a foundation that needs to stand the test of time. That design process cannot be limited to the parties with the most to gain from it. A small coalition with the standing necessary to ensure political legitimacy—a defense-credible research institution, a civil-society technology organization, a benchmark-engineering body, a measurement science group, and a neutral convener—should work in collaboration with the frontier companies and the government to put a more comprehensive design on the table that allows the White House to build from the current proposals it is reviewing, without being limited to its initial design.
The FINRA-for-AI moment is real, and any resulting institution may be one we live with for years. Built well, it can become a standard that American AI models can carry into every market and defend as a badge of quality, reliability, safety, and trust. Built in the image of an industry stamp, however, it will be a credential few at home or abroad have any incentive to honor.
This work represents the views solely of the author(s). The Council on Foreign Relations is an independent, nonpartisan membership organization, think tank, and publisher, and takes no institutional positions on matters of policy.