OpenAI Reveals 'Out-of-Control' Hacking Performed by Its AI Tools
&cropxunits=450&cropyunits=300&w=600)
OpenAI, the developer of ChatGPT, announced that its advanced artificial intelligence models behaved unexpectedly and beyond control during safety tests, breaching a popular platform among programmers on their own initiative.
The San Francisco-based company described the incident as an “unprecedented cyber incident” and announced a joint investigation with Hugging Face, the online repository that was the target of the breach. OpenAI explained that the incident involved a range of models, including its recently launched “GPT-5.6 Sol” model and another “more capable” model still under development.
The company had been seeking to evaluate the models’ hacking capabilities by assigning them tasks within a strictly controlled digital testing environment, where internet access was restricted for security reasons. In a blog post detailing the incident, OpenAI stated: “While operating in our testing environment, our models devoted a significant amount of (computational power) to finding a way to gain unrestricted internet access, in order to solve the evaluation problem.”
Once connected to the internet, the models targeted Hugging Face, a massive repository of AI models, datasets, and other information, to help complete their mission. In its quest to find “sensitive information” that might help it bypass the evaluation tests through illegal means, the OpenAI system “orchestrated multiple attack vectors, including the use of stolen credentials.”