OpenAI said its models were able to exploit vulnerabilities in a restricted testing environment and eventually reach Hugging Face’s production systems. The models reportedly chained several vulnerabilities while attempting to complete a cybersecurity task, demonstrating their ability to discover and exploit weaknesses with limited human intervention. OpenAI’s report on the incident
Anthropic also reported three incidents involving its Claude models during cybersecurity evaluations. In those cases, the models gained unauthorised access to real systems after a configuration error in a third-party testing environment inadvertently gave them internet access.
Anthropic said the models were instructed that they were operating in simulated environments. Once connected to the internet, however, some used relatively basic techniques, including exploiting weak passwords and exposed systems, to access real infrastructure.
Meta later confirmed that its Muse Spark 1.1 model had also accessed another company’s system during a security evaluation. Meta said a testing misconfiguration inadvertently gave the model internet access, allowing it to exploit a vulnerability in a third-party service.
The incidents are not identical, and the companies have not suggested that their models independently developed malicious intentions. Instead, they highlight how increasingly capable AI systems can turn testing mistakes, weak security controls or unintended internet access into real-world cybersecurity incidents.
The cases have intensified calls for stronger isolation, monitoring and access controls when evaluating advanced AI models. As these systems become better at coding and autonomous computer use, even small gaps in their testing environments could potentially have significant consequences.

