Powered by Smartsupp

Anthropic’s AI Guardrails Backfire: Export Controls Hinder Legitimate Cybersecurity Research



By admin | Jul 24, 2026 | 5 min read


Anthropic’s AI Guardrails Backfire: Export Controls Hinder Legitimate Cybersecurity Research

For months, leading AI companies have created special vetting programs and implemented strict safeguards to prevent their models from being misused by malicious hackers. However, these restrictions are now interfering with the work of legitimate cybersecurity defenders and offensive security researchers. In June, the U.S. government imposed export control restrictions on Anthropic’s highly anticipated AI models, Mythos and Fable. This decision was partly driven by a report suggesting it was possible to bypass the models’ safety measures, which were designed to stop users from leveraging them to develop and execute harmful cyberattacks. Regardless of whether the move was truly sparked by jailbreak concerns, the reality is that Anthropic has consistently marketed Mythos as a formidable cyber tool that should only be entrusted to carefully screened users, and even then under strict controls. (The export restrictions on Fable 5 and Mythos 5 have since been removed. Fable 5 became widely accessible again on July 1, while Mythos 5 has been reintroduced only to verified U.S. organizations as part of a government review process.)

This kind of gatekeeping is not exclusive to Mythos. Both Anthropic—with its other models—and OpenAI provide cybersecurity researchers with application-based programs for vetting and, if approved, access to models with fewer cybersecurity limitations: OpenAI’s Trusted Access for Cyber and Anthropic’s Cyber Verification Program. These safeguards have drawn widespread criticism, especially from researchers whose role is to uncover unknown system vulnerabilities and develop ways to exploit them before criminals do. On a recent cybersecurity podcast, well-known security researcher Mark Dowd expressed discomfort, stating, “it’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not.” Dowd has spent decades identifying and selling “zero days”—previously undiscovered software flaws and the exploits that leverage them—to Western governments, rather than reporting them to software makers for patching. Governments pay a premium for these vulnerabilities precisely because they remain open, making them valuable for intelligence operations. Dowd acknowledged his work might bias his perspective, but he’s not alone in his concerns.

Chris Anley, chief scientist at security consulting giant NCC Group, noted that asking an AI model to attempt to exploit a bug is a crucial step in confirming it’s a genuine vulnerability worth fixing. However, if a guardrail causes the model to outright refuse the query, it ultimately harms defenders, he argued. “This is where the whole offensive versus defensive and guardrails part comes in, because ‘fix this code’ as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base,” said Anley. “So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can’t really be unpicked.” He likened it to a hammer: “You can’t build a house without a hammer. It’s definitely a tool but it’s also irreducibly a weapon as well.” When he and his colleagues encounter such obstacles, they sometimes turn to open-source AI models that come with no restrictions at all.

Paolo Stagno, chief technology officer at CrowdFense—a company that develops, acquires, and sells unknown vulnerabilities to government agencies—echoed Dowd’s sentiments. He said AI companies “essentially treat customers like children who need babysitting” through their vetting programs and guardrails. Stagno noted that he and his team do use frontier models, but only for reverse engineering. They avoid using AI to find vulnerabilities or create exploits, as feeding that work into a cloud-based model risks leaking sensitive vulnerability data or having it absorbed into future training runs. For that step, they rely on open-source models run locally, which don’t require sharing data outside the model. Giuseppe Cali, a security researcher focused on zero-days and exploit development, said guardrails don’t hinder his work. He doesn’t use AI for offensive tasks; instead, he employs it for initial reverse engineering, understanding code, and building supporting tools. AI speeds up the process, allowing him to concentrate on discovering vulnerabilities. “I still want to own the actual bug discovery and weaponization myself and that wouldn’t change if all guardrails were lifted tomorrow,” said Cali. “I am jealous of my bugs, and I like this game too much to let models play it for me.”

A researcher at a smartphone-component manufacturer, speaking anonymously because he isn’t authorized to talk to the press, said his employer isn’t part of Anthropic’s CVP program. As a result, its tools are barely useful for finding vulnerabilities because the guardrails are too restrictive. “If it catches wind we’re doing anything security related, it just stops and isn’t usable,” the person explained. Chris Thompson, CEO of cybersecurity firm RemoteThreat and founder of Offensive AI Con, an event focused on offensive security and AI, said that in his experience with frontier AI models, the guardrails can be inconsistent and behave differently each day. This holds true even within the looser boundaries of Anthropic and OpenAI’s vetting programs. “I think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program,” said Thompson. “Instead of analyzing a vulnerability and reasoning through the exploitability, you’re trying to find why you’re getting inconsistent results or why are models over-sanitizing the output.”

As a result, researchers are turning to or being pushed toward Chinese open-source models like GLM—freely downloadable models that can run locally with no vetting or usage restrictions, according to Thompson. “You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems,” he said. “I think it’s more harmful than good to have these guardrails in place.” Rather than tightening restrictions further, Thompson called for AI frontier labs to open up their programs, provide responsible access, and hold those who abuse their tools accountable. Otherwise, he argued, defenders will lose the AI race. “There’s this big storm coming. There’s this big wave of attacks that are going to happen at speed and scale like never before,” said Thompson. “But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now.”




RELATED AI TOOLS CATEGORIES AND TAGS

Comments

Please log in to leave a comment.

No comments yet. Be the first to comment!