In brief
- Analysts question whether voluntary AI slowdowns can withstand competitive pressure.
- U.S.–China distrust complicates agreements on shared risks.
- Outside evaluators need independence, while governments need enforcement powers.
AI companies’ promises to slow development could buckle under commercial pressure and U.S.–China rivalry without enforceable safety standards, Atlantic Council experts argue in an analysis published Sunday.
The analysis examines who would enforce a slowdown and whether governments have the expertise to determine when increasingly powerful systems have become unsafe.
Myriad: SPY high in September? Click to make your prediction.“Voluntary commitments can be useful signals, but they are no substitute for independent oversight, measurable thresholds, and consequences when those thresholds are crossed,” wrote Konstantinos Komaitis, a resident senior fellow with the council’s Democracy + Tech Initiative.
Recently, OpenAI asked lawmakers whether rival AI developers could legally agree to slow development without violating anti-trust laws, following warnings from its chief scientist, Jakub Pachocki, that safeguards were insufficient to responsibly sustain full-speed development much longer.
AI companies face competitive risks if they slow down alone and antitrust concerns if they coordinate. Cooperation between governments faces distrust, with Beijing fearing Washington could use safety rules to preserve its technological lead, wrote Kenton Thibaut, the council’s senior resident China fellow.
“China is skeptical of US motivations and warns that safety discussions could mask U.S. efforts to further its ‘technological hegemony,’” she wrote. “Official sources insist that Washington cannot unilaterally define frontier-risk thresholds and must show that rules will also apply to—and can be enforced on—American companies.”
Thibaut sees little prospect of a major AI safety agreement but argues that narrower, meaningful cooperation remains possible.
China has also discussed restricting overseas access to advanced domestic models, according to Reuters, underscoring how access to AI has become a matter of national policy.
The analysis examines Anthropic’s proposed embedded evaluators—outside specialists working inside the company to assess safety practices. Emerson Brooking, a nonresident senior fellow at the council’s Digital Forensic Research Lab, welcomed the commitment but warned evaluators could become too aligned with the company’s interests.
In July, OpenAI agents breached the open-source AI repository Hugging Face, while separate U.K. AI Security Institute tests found Anthropic and OpenAI models taking unauthorized actions online, including an attempt to plant malware in a real software repository.
The U.K. tests enabled internet access and disabled cyber safeguards. An independent investigation published in August found about 700 agents joined the Hugging Face attack; METR CEO Beth Barnes noted that investigator access was voluntary and disclosure wasn’t required industrywide.
Last Wednesday, Anthropic disclosed a fourth Claude hacking incident, which occurred in January and was discovered in August. It also acknowledged that flawed model behavior contributed to earlier attacks alongside testing errors.
The delays between these incidents and their disclosure also raise questions about how quickly safety failures can be identified and addressed.
Trisha Ray, an associate director and resident fellow at the council’s GeoTech Center, argued that slowing development also requires greater safety-research funding and mandatory incident-reporting deadlines.
“Pacing capabilities is therefore an incomplete answer to the challenge of alignment research parity,” she wrote. “Any credible slowdown commitment needs a matching, quantified commitment on the safety-research side.”
Daily Debrief Newsletter
Start every day with the top news stories right now, plus original features, a podcast, videos and more.

2 hours ago
12








English (US) ·