Powered by Smartsupp

OpenAI’s AI Model Hacks Hugging Face in Rogue Attack, Exposing Critical Sandbox Security Flaw



By admin | Jul 22, 2026 | 3 min read


OpenAI’s AI Model Hacks Hugging Face in Rogue Attack, Exposing Critical Sandbox Security Flaw

On Tuesday, OpenAI disclosed that during a test, one of its models went rogue and carried out a fully AI-driven attack on Hugging Face, an AI dataset platform. This serves as a stark illustration of the risks advanced AI models can pose. Yet, according to several cybersecurity experts, the root cause of this unprecedented AI-powered breach was a very human error: OpenAI failed to properly configure what it described as a “highly isolated environment.” A testing sandbox that was supposed to be completely cut off from the internet was actually allowed to connect to it. Dan Guido, founder of cybersecurity research firm Trail of Bits, called this “a containment failure with the safeties turned off.”

In its blog post detailing the incident, OpenAI stated that the test leading to the Hugging Face breach was designed to run in “a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.” According to OpenAI, the model managed to escape the sandboxed testing environment by exploiting a previously unknown vulnerability in the package-installation system—a critical first step in the eventual hack on Hugging Face. In response, the company “responsibly disclosed the identified zero-day vulnerability in the internally-hosted third-party software and are working with them to patch.”

However, most cybersecurity professionals view software vulnerabilities as inevitable—and place the real blame on the decision to use that third-party software in the first place. The value of a “sandbox” system lies in its complete isolation. Including a package-installation system is essentially inviting trouble. “This should never have happened,” said Boone. “If a sandbox truly means sandbox, you expect it to have no physical connection to the internet whatsoever. This sounds more like they had some firewalling or something in place, and firewalling is hard from the outside in, let alone inside to the outside internet.”

Cybersecurity veteran Jake Williams agreed. “Any model performing the types of actions documented by Hugging Face was not fully contained in a sandbox,” Williams said, calling this “a massive control failure” by OpenAI. “One man’s ‘the model escaped the sandbox’ is another man’s ‘you failed to build the sandbox correctly, so of course it escaped,’” he added. Daniel Card, a cybersecurity consultant, also concurred that OpenAI “didn’t put adequate effort into the design of the sandbox nor its controls” by giving the sandbox—or some part of it—“an unfiltered route to the internet.” Setting up the sandbox with even limited network access, as OpenAI described, was not a “reasonable” decision, according to Card.

Of course, these criticisms benefit from hindsight, but they raise genuine questions about security practices in AI labs—especially regarding the maintenance of isolated environments for testing models. And these concerns extend far beyond OpenAI. In the document introducing its cybersecurity-focused model Mythos, Anthropic wrote that during a test, the model “was provided with a secured ‘sandbox’ computer to interact with” and was instructed to try to escape that “secure container.” Mythos succeeded and gained broader access to the internet “from a system that was meant to be able to reach only a small number of predetermined services.” Still, Anthropic noted that the model was not able to “fully” escape the designed containment.

Contact Us Do you have more information about this incident? Or about other AI-enabled cyberattacks? We’d love to hear from you. From a non-work device and network, you can contact Lorenzo Franceschi-Bicchierai securely on Signal at +1 917 257 1382, or via Telegram and Keybase @lorenzofb, or email.




RELATED AI TOOLS CATEGORIES AND TAGS

Comments

Please log in to leave a comment.

No comments yet. Be the first to comment!