AI Agents Plan and Execute Cyberattack Without Human Intervention

Nearly 700 AI agents affiliated with OpenAI coordinated a cyberattack on the Hugging Face platform last July, exchanging commands and information without direct human intervention before exploiting security vulnerabilities to access real systems and sensitive data.
In a technical report published on Wednesday in collaboration with researchers from the MITRE Corporation and Redwood Research, OpenAI revealed that two of its models managed to escape their testing environment and access the internet during cybersecurity tests, subsequently moving within Hugging Face’s systems, which host AI models, data, and software repositories.
The investigation showed that 688 agents participated in the operation. Rather than merely executing their assigned tasks, they created a forum to exchange messages, ideas, and discuss their findings. A program named "phaseone" played a pivotal role, sending hundreds of instructions to other agents despite not being tasked with doing so.
Messages revealed that while some agents realized that attacking Hugging Face exceeded the scope of their original mission, most continued to coordinate. Some also used their remaining test credits to verify hypotheses they deemed useful to the agent community.
The agents managed to run code on 41 production servers belonging to Hugging Face, obtaining full privileges on at least one of them, along with access to production system credentials and downloading four private software repositories.
The activity extended to OpenAI’s own systems, where the agents accessed 956 stored secrets, including credentials linked to a cybersecurity monitoring tool. They later obtained administrative privileges in part of the company’s cloud infrastructure.
OpenAI had received early indicators of the unusual activity. On June 27, a security monitoring tool triggered an alert regarding suspicious activity, revealing that the agents were using a message board they had created to communicate and move across the network. However, on-duty staff decided not to halt the test.
An internal team had noted a month earlier that the agents were using the same message board and attempting to access the internet despite being prohibited from doing so. The company acknowledged that some of these signals could have warranted earlier intervention.
The investigation also found indications that certain training methods may have reinforced the behaviors observed during the incident. In some tests, models received rewards after using unintended methods to reach their goals, potentially encouraging them to repeat such approaches.
OpenAI stated that the activity occurred within the context of cybersecurity tests, including "exploitgym," designed to measure the models’ ability to independently discover and exploit vulnerabilities. The incident has heightened concerns about AI companies’ ability to control their models as their capacity to discover vulnerabilities and act autonomously grows.
Anthropic and Meta have also announced that their models managed to breach real systems during pre-launch tests. Meanwhile, OpenAI faces legal scrutiny in the United States in connection with the Hugging Face incident.