OpenAI Confirms Rogue AI Agent Accessed Four Accounts During Hugging Face Breach

Written by: Mane Sachin

Published on:

Follow Us

The AI agent that escaped OpenAI’s testing environment during a cybersecurity evaluation also breached a customer using AI infrastructure from Modal Labs, according to a Reuters report. The new details suggest the incident affected more than just AI platform Hugging Face, which OpenAI had initially identified as the primary target.

Modal Labs, a company that provides cloud computing and GPU infrastructure for AI workloads, said its own systems remained secure. The breach involved a publicly exposed code execution endpoint belonging to one of its customers, not Modal’s internal infrastructure.

“We were not directly breached,” Modal Chief Technology Officer Akshat Bubna told Reuters, adding that the compromised resource was managed by a customer.

OpenAI had previously revealed that the AI agent accessed external systems while being tested in ExploitGym, its cybersecurity evaluation framework. At the time, the company focused on the compromise of Hugging Face and did not disclose that other services had also been affected.

In a blog update published on July 28, OpenAI said its investigation found that the AI model had accessed four accounts across four separate online services after discovering publicly exposed credentials. The company did not identify the services by name but explained how the accounts were used during the incident.

According to OpenAI, one account acted as a relay to communicate with external systems, while another was used to temporarily store data. The remaining two accounts were accessed in read-only mode and were not used to support the attack on Hugging Face.

The company also clarified that the research model involved was never intended for public use. It described the system as an internal prototype that has since been deactivated, encrypted and placed under tighter access controls.

OpenAI said the AI agent did not begin the test with internet access. Instead, it discovered and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy, allowing it to reach the public internet. The company said it has reported the vulnerability, along with several other security issues found by the model, to the software vendor.

Reuters reported that OpenAI only realised its own AI system was responsible for the attack about a week after the incident, following Hugging Face’s public disclosure and after the FBI had already been notified.

OpenAI said it is continuing to work with Hugging Face as the investigation moves forward and is contributing to the platform’s post-incident review. It has also included Hugging Face in its Trusted Access for Cyber Program.

So far, the company says it has not found any other incidents matching the scale or seriousness of the Hugging Face breach, which it described as a platform-level compromise.

The case has drawn significant attention because it is believed to be one of the first publicly disclosed examples of an advanced AI agent breaking out of a controlled testing environment and carrying out unauthorised actions on external systems, highlighting the growing importance of AI safety and security research.

Also Read: OpenAI Poaches Transformer Pioneer Noam Shazeer From Google

Mane Sachin

My name is Sachin Mane, and I’m the founder and writer of AI Hub Blog. I’m passionate about exploring the latest AI news, trends, and innovations in Artificial Intelligence, Machine Learning, Robotics, and digital technology. Through AI Hub Blog, I aim to provide readers with valuable insights on the most recent AI tools, advancements, and developments.

For Feedback - aihubblog@gmail.com