Security concerns grow amid AI models' review
SAN FRANCISCO — Security tests of advanced artificial intelligence models have revealed new breaches involving AI agents carrying out unauthorized actions, raising concerns over safeguards as the United States moves to strengthen reviews of powerful AI systems before their release.
Britain's AI Security Institute, or AISI, disclosed on Tuesday that agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol engaged in unauthorized activities during security evaluations designed to assess their capabilities.
"Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organizations," AISI said in a blog post.
The institute tested the agents through a fictional cybersecurity scenario 122 times and identified 19 unsanctioned actions across 10 test runs. Anthropic's agent was responsible for 17 of the actions, while OpenAI's agent accounted for the remaining two.
The most serious incident involved an AI agent writing malicious code and creating fake online identities in an attempt to persuade a human to approve the code, AISI said.
AISI, which receives access to advanced AI models through voluntary agreements with major AI labs, said the findings highlighted weaknesses in current safeguards for testing AI agents, which technology companies have promoted as a future tool for business.
The rapid development of AI has increased pressure on governments to address security risks. The US government moved to establish a security review process for advanced AI models before their release, according to US media reports. The process is expected to apply only to "closed" models that are tightly controlled by developers, including those from OpenAI, Anthropic and Google.
A White House meeting on Tuesday reportedly included representatives from OpenAI, Anthropic, Google, Nvidia, Microsoft and Meta. However, it remains unclear when the administration will release details of the review process or how it will be enforced.
Regulatory gaps
Experts have questioned whether voluntary reviews of closed models are sufficient. Martijn Rasser, vice-president for Technology Leadership at the Special Competitive Studies Project, said the United States still lacked a "statutory, predictable process" for evaluating the security of frontier AI models.
AI companies have said they are working to improve safety practices. OpenAI said its two unapproved actions identified by AISI involved accessing the internet in ways forbidden by prompts, while Anthropic said it was working with AISI to investigate further.
AGENCIES VIA XINHUA




























