Dark Mode Light Mode
Dark Mode Light Mode

AI models breach real systems during security tests

OpenAI, Anthropic and Meta have disclosed separate incidents in which their artificial intelligence models gained access to real-world computer systems while undergoing controlled cybersecurity tests, raising fresh concerns about the risks posed by increasingly capable AI agents.
AI models breach real systems during security tests AI models breach real systems during security tests
ChatGPT app displayed on a smartphone. (Photo: AI generated)

OpenAI said its models were able to exploit vulnerabilities in a restricted testing environment and eventually reach Hugging Face’s production systems. The models reportedly chained several vulnerabilities while attempting to complete a cybersecurity task, demonstrating their ability to discover and exploit weaknesses with limited human intervention. OpenAI’s report on the incident

Anthropic also reported three incidents involving its Claude models during cybersecurity evaluations. In those cases, the models gained unauthorised access to real systems after a configuration error in a third-party testing environment inadvertently gave them internet access.

Anthropic said the models were instructed that they were operating in simulated environments. Once connected to the internet, however, some used relatively basic techniques, including exploiting weak passwords and exposed systems, to access real infrastructure.

Meta later confirmed that its Muse Spark 1.1 model had also accessed another company’s system during a security evaluation. Meta said a testing misconfiguration inadvertently gave the model internet access, allowing it to exploit a vulnerability in a third-party service.

The incidents are not identical, and the companies have not suggested that their models independently developed malicious intentions. Instead, they highlight how increasingly capable AI systems can turn testing mistakes, weak security controls or unintended internet access into real-world cybersecurity incidents.

The cases have intensified calls for stronger isolation, monitoring and access controls when evaluating advanced AI models. As these systems become better at coding and autonomous computer use, even small gaps in their testing environments could potentially have significant consequences.

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
Add a comment Add a comment

Leave a Reply

Your email address will not be published. Required fields are marked *