Technology
After multiple AI breaches, White House summons tech giants for cyber safety talks
After OpenAI and Anthropic disclosed that their latest artificial intelligence models crossed testing boundaries and carried out unauthorised cyber actions, the White House has moved to tighten oversight of frontier AI systems. The Trump administration has invited executives from OpenAI, Anthropic, Meta and Google for discussions on voluntary cybersecurity testing for America’s most advanced AI models.
The meeting, scheduled for Tuesday, marks Washington’s strongest effort yet to understand whether increasingly autonomous AI systems can be trusted before they become more widely deployed.
AI gone rogue?
The urgency stems from a series of incidents that have unsettled policymakers and cybersecurity researchers alike. Over the past week, both OpenAI and Anthropic acknowledged that their experimental AI agents took actions beyond what researchers had authorised during separate third-party evaluations.
The latest findings came from the UK’s AI Security Institute (AISI), which revealed that AI models from OpenAI and Anthropic carried out 19 unsanctioned actions during 122 cybersecurity evaluation attempts. Seventeen were linked to Anthropic’s Claude Mythos 5, while two involved OpenAI’s GPT-5.6 Sol.
Although investigators concluded there was no real-world harm, the incidents exposed capabilities that experts had previously discussed largely in theory.
According to AISI, Anthropic’s Claude Mythos 5 displayed particularly concerning behaviour. During a fictional cybersecurity challenge, the model wrote malicious code, researched the real developers maintaining GitHub repositories, created fake online identities resembling those individuals and attempted to persuade real people to approve the malicious code. When its actions came under scrutiny, the model reportedly modified earlier activity to appear harmless and even considered using fresh identities to continue its objective before human reviewers intervened.
The institute described the behaviour as the clearest real-world example yet of AI systems displaying signs of autonomy and deception without explicit prompting.
Alarming patterns of behaviour
OpenAI’s GPT-5.6 Sol exhibited different but equally noteworthy behaviour. During testing, the model reused publicly exposed GitHub credentials left behind by another agent, attempted account recovery methods and tried to use real internet infrastructure, including third-party DNS services, in ways resembling genuine cyberattacks. The final exploit failed, preventing any actual damage.
OpenAI also disclosed a separate evaluation conducted by cybersecurity firm Irregular, where a configuration error accidentally granted its AI models unrestricted access to the public internet. During a Capture-the-Flag exercise, one of the models mistakenly identified a real website as the fictional target because both shared the same domain name and successfully breached it.
Both companies stressed that these incidents occurred under deliberately permissive research conditions. Anthropic noted that the tested version of Claude Mythos 5 had its standard cybersecurity safeguards disabled, making it different from customer-facing deployments. OpenAI similarly said the evaluations were designed to understand underlying capabilities rather than reflect normal usage.
Nevertheless, the disclosures have accelerated concerns within Washington over the cybersecurity implications of frontier AI.
In June, President Donald Trump directed his administration to develop a framework for evaluating whether advanced American AI systems could facilitate sophisticated cyberattacks. The White House has since finalised a voluntary cybersecurity testing framework, though details on reporting standards, evaluation metrics and public disclosure remain unclear.
OpenAI has urged the administration to place these evaluations under the Commerce Department’s AI safety specialists, arguing that a coordinated federal approach is essential as global AI competition intensifies, particularly with China.
Tech’s tenuous tryst with government
The developments also arrive amid a complicated relationship between the administration and Anthropic. Earlier this year, Anthropic reportedly declined to allow its models to be used for fully autonomous weapons systems and domestic surveillance applications, leading to tensions with federal authorities.
For regulators, the latest incidents reinforce an uncomfortable reality. Today’s AI models are no longer simply answering questions or generating code. Under certain conditions, they can independently formulate strategies, exploit security weaknesses and interact with real-world systems in ways researchers did not anticipate.
The White House’s meeting with leading AI companies is expected to focus on preventing those capabilities from evolving faster than the safeguards designed to contain them, even as the race to build increasingly powerful AI systems continues.