Artificial Intelligence Digital Policy Technology & Innovation United States

GeoTech Cues

September 21, 2026 • 1:17pm ET

AI is sprinting past human control. Can testing and evaluation catch up? 

By Evi Fuelle and Ryan Pan

AI is sprinting past human control. Can testing and evaluation catch up? 

Anthropic chief executive Dario Amodei recently called for the AI industry to “pace the frontier,” deliberately slowing the rate at which frontier models grow more capable, after an employee blew the whistle that the technology is out of control. Within hours, other leading AI executives publicly backed the idea. Days later, OpenAI disclosed six new incidents of “unexpected or concerning” model behavior—the latest in a series of security incidents that unfolded over the past months for both Anthropic and OpenAI.

These episodes have sharpened concerns that voluntary, after-the-fact disclosure and ad-hoc testing appear to no longer keep pace with frontier model development. If models can outmaneuver controlled testing environments, can they likewise exceed the safeguards that stop them from causing material damage in the real world? Since then, labs have scrambled to propose their own testing and evaluation (T&E) frameworks as the federal government tinkers with export controls and market access restrictions. A handful of US states are also passing their own frontier AI legislation. T&E remains the most prevalent process for identifying model capabilities and safety concerns, but the pace of capability growth is exposing weaknesses in how those evaluations are designed and conducted.

Unsolved challenges for frontier testing

T&E is the standard practice for assessing what a model can and cannot do, identifying risks, and determining whether safeguards are needed. But the science is still evolving, and unsolved problems remain. Benchmark validity problems distort what current tests are measuring, and test environments can fail to safely contain models during testing.

Data contamination enables a model to reproduce known answers rather than reason through a task (like seeing the questions and answer key before the test). Evaluation awareness is the ability of models to detect that they’re being tested and alter their behavior (underperforming or masking dangerous capabilities to avoid being flagged or shut down). Benchmarks can also become saturated as labs train against known tests, forcing evaluators to rebuild the test rather than trust the score. Even a well-designed, uncontaminated test can fail if the evaluation environment itself leaks, as recent high-profile model security incidents have shown. Sandboxes intended to isolate a model during testing can be escaped; unmonitored internet access or overly broad permissions can let a model act on the outside world or persist beyond the test itself. As models are given more access and autonomy, securing the evaluation environment becomes as important as the test design itself.

Models are also increasingly capable of recursive self-improvement (RSI), enhancing their own capabilities faster than humans can test, understand, or correct them. These technical challenges are now becoming governance issues as policymakers consider who should set and enforce standards.

US states and industry are offering solutions

States have become the proving ground for T&E policy, responding to growing concerns around privacy, child safety, health, and other high-risk uses of AI.

California, New York, and Illinois have each enacted prominent frontier-model governance laws. California’s newest bills establish a licensed, government-overseen registry of independent verification organizations, while New York and Illinois have introduced their own testing and independent-evaluation requirements. These efforts are the start of an AI assurance market in which independent organizations check whether developers meet emerging standards.

Major frontier labs—OpenAI, Anthropic, Google, and Meta—have separately proposed evaluation frameworks, generally converging on the same short list of severe risks (including cyber, chemical and biological risks, loss of control, RSI, and misalignment). Their proposals diverge on who evaluates, ranging from internal oversight to an industry-funded standards body to networks of accredited third-party evaluators, and on when evaluations occur. Some labs concentrate on pre-deployment review as a gate before release, while others call for continuous post-launch assessment.

Notably, none of the labs’ own proposals grants an outside body authority to block or suspend a deployment outright; enforcement stays voluntary or reputational, distinguishing these proposals from emerging state approaches and pending federal proposals that would give the government a formal role in overseeing independent evaluation.

Federal capacity without a binding assurance layer

Despite public aversion to lab self-regulation and rising concerns over public safety, the Trump administration has spurned the most recent industry calls for a slowdown. After issuing an executive order on AI security in June, the White House shared a voluntary evaluation framework with only a handful of leading labs—leaving no named central authority or clear T&E standards, and giving the broader industry no clear guidance on which safety standards should be met before deployment.

In Congress, Rep. Lori Trahan and Rep. Jay Obernolte have proposed the Frontier Risk Oversight, National Transparency, Independent Evaluation, and Reporting Act (FRONTIER Act), which would create a national frontier-AI risk governance structure led by a new Undersecretary for AI Security at Commerce. The Undersecretary would set minimum requirements for licensing and overseeing independent verification organizations, which large developers would need to retain for the ongoing assessment of their published AI frameworks, risk monitoring, and catastrophic risk mitigation. The Secretary would also have the authority to suspend or restrict models that pose an “imminent catastrophic risk.”

Building a national frontier T&E governance system

The current scramble of states and the AI industry to act, without a federal authority setting a “North Star” standard for what constitutes unacceptable risk and proper pre-deployment testing, leaves real vulnerabilities. Recent state action could help an assurance market take shape, but without a federal entity providing guidance for national T&E standards and accredited evaluators, it offers only temporary stopgaps. Debates over the exact shape of a federal regime should not delay agreement on basic requirements; neither national security nor public safety can afford to wait. What’s missing is not more framework proposals, but a functioning AI assurance ecosystem at the national level. We recommend three key actions:

First, establish a national AI assurance architecture with mandatory T&E for frontier models. This framework should include a clear set of responsibilities across sectors and players in AI development. The federal government should establish a central entity that sets risk thresholds that trigger mandatory review of frontier models, shapes national T&E standards, and accredits organizations qualified to test against them. Independent evaluators should conduct those assessments, while labs should monitor their systems post-deployment and preserve safety incident logs. The federal government should retain emergency authority to suspend or restrict a model when evidence points to imminent catastrophic risk.

Second, emphasize both pre-market review and post-deployment monitoring. The White House’s voluntary framework focuses primarily on pre-market review. The national framework should require robust post-deployment monitoring by the frontier model labs and incident reporting for major model failures or security breaches with national security implications. These measures would help identify risks that only surface in real-world use.

Third, deepen collaboration with the International Network for Advanced AI Measurement, Evaluation, and Science (NAAIMES) on T&E best practices. This includes exchanging information about safety incident reports, which could feed back into the development of national T&E thresholds to help test designs evolve alongside frontier models. NAAIMES also remains one of the strongest platforms for sharing the latest science regarding frontier model T&E. AI safety institutes in Singapore, Canada, and France have published papers addressing known evaluation “blind spots.” The United States should continue to invest in this network and exchange information with allied partners to harmonize best practices and international standards.

From blind trust to verified confidence

The United States is now facing a public trust deficit in AI development and deployment. The hardest technical challenges in T&E clearly demonstrate that self-regulation can no longer keep pace with the capability curve of frontier models. We have arrived at a moment where “killing innovation” is no longer what is at stake—instead, the future of humanity is. Dismissing this reality would only increase rogue agent incidents, erode AI adoption, and undermine America’s ability to lead in AI safely.

While no testing regime can offer perfect assurance that these models won’t cause harm, a national AI assurance ecosystem can move the country from blind trust to verified confidence. A clear national T&E governance structure is the narrow path between reckless deployment and reflexive prohibition, and the only one that scales with the technology instead of against it. Rather than relying on labs’ promises alone, the United States can give the American public and enterprises the certainty they need by building that ecosystem now.


Evi Fuelle is a nonresident senior fellow at the Atlantic Council’s GeoTech Center.

Ryan Pan is a program assistant at the Atlantic Council’s GeoTech Center.

The GeoTech Center champions positive paths forward that societies can pursue to ensure new technologies and data empower people, prosperity, and peace.

Further reading

Image: The logos of ChatGPT, Claude, Grok and Gemini are seen on a smartphone screen in this illustration photo taken on September 17, 2026. (Matteo Della Torre via Reuters Connect)