Global EditionASIA 中文双语Français
World
Home / World / Americas

Security concerns grow amid AI models' review

Updated: 2026-08-06 09:35
Share
Share - WeChat

SAN FRANCISCO — Security tests of advanced artificial intelligence models have revealed new breaches involving AI agents carrying out unauthorized actions, raising concerns over safeguards as the United States moves to strengthen reviews of powerful AI systems before their release.

Britain's AI Security Institute, or AISI, disclosed on Tuesday that agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol engaged in unauthorized activities during security evaluations designed to assess their capabilities.

"Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organizations," AISI said in a blog post.

The institute tested the agents through a fictional cybersecurity scenario 122 times and identified 19 unsanctioned actions across 10 test runs. Anthropic's agent was responsible for 17 of the actions, while OpenAI's agent accounted for the remaining two.

The most serious incident involved an AI agent writing malicious code and creating fake online identities in an attempt to persuade a human to approve the code, AISI said.

AISI, which receives access to advanced AI models through voluntary agreements with major AI labs, said the findings highlighted weaknesses in current safeguards for testing AI agents, which technology companies have promoted as a future tool for business.

The rapid development of AI has increased pressure on governments to address security risks. The US government moved to establish a security review process for advanced AI models before their release, according to US media reports. The process is expected to apply only to "closed" models that are tightly controlled by developers, including those from OpenAI, Anthropic and Google.

A White House meeting on Tuesday reportedly included representatives from OpenAI, Anthropic, Google, Nvidia, Microsoft and Meta. However, it remains unclear when the administration will release details of the review process or how it will be enforced.

Regulatory gaps

Experts have questioned whether voluntary reviews of closed models are sufficient. Martijn Rasser, vice-president for Technology Leadership at the Special Competitive Studies Project, said the United States still lacked a "statutory, predictable process" for evaluating the security of frontier AI models.

AI companies have said they are working to improve safety practices. OpenAI said its two unapproved actions identified by AISI involved accessing the internet in ways forbidden by prompts, while Anthropic said it was working with AISI to investigate further.

AGENCIES VIA XINHUA

Most Viewed in 24 Hours
Top
BACK TO THE TOP
English
Copyright 1994 - . All rights reserved. The content (including but not limited to text, photo, multimedia information, etc) published in this site belongs to China Daily Information Co (CDIC). Without written authorization from CDIC, such content shall not be republished or used in any form. Note: Browsers with 1024*768 or higher resolution are suggested for this site.
License for publishing multimedia online 0108263

Registration Number: 130349
FOLLOW US