How AI guardrails are hindering the work of offensive cyber security researchers

For many months, the AI giants have developed special audited programs and strict rules to limit the use of their models by malicious hackers. But these restrictions now hinder the work of legitimate network defenders, as well as those of pessimistic cyber security researchers.
In June, the US government imposed export control restrictions on the Anthropic AI models most associated with Mythos and Fable. The move was prompted at least in part by a report that said it was possible to bypass the security lines of models designed to prevent users from using them to build and execute malicious cyber attacks.
Regardless of whether the incident was really inspired by the fear of a prison break, the fact is that Anthropic has repeatedly marketed Mythos as some kind of cyber doomsday device that can only be given to carefully vetted users, and even then with strong monitoring devices. (Export controls on Fable 5 and Mythos 5 have since been lifted. Fable 5 returned to general access on July 1; Mythos 5 was also introduced to US organizations that were vetted as part of the government’s review process.)
That kind of gatekeeping isn’t unique to Mythos. Both Anthropic, and its other models, and OpenAI offer cybersecurity researchers programs that they can use to test and – if approved – access models with fewer cyber security restrictions: OpenAI’s Trusted Cyber Access and Anthropic’s Cyber Verification program.
These security measures have been widely criticized, especially by researchers whose job it is to find unknown vulnerabilities in systems and develop ways to exploit them before criminals do.
During a recent appearance on a cybersecurity podcast, Mark Dowd, a well-known security researcher, said, “It’s very uncomfortable for me that these big random companies are making arbitrary decisions about what’s safe in terms of security and what’s not.”
Dowd has spent decades finding and selling “zero days” – previously unknown software flaws and exploits that use them – to Western governments, rather than reporting them to software developers for patches. Governments pay for vulnerabilities precisely because they remain open, which is useful for intelligence operations.
Dowd admitted that his job may make him biased, but he’s not alone. Several people who work in offensive cybersecurity — proactively probing systems for vulnerabilities — explained to TechCrunch how they’re using AI tools and dealing with their monitoring tools.
Chris Anley, senior scientist at security giant NCC Group, said asking an AI model to try to exploit a bug is an important step to ensure it’s a real vulnerability that needs to be fixed. But if the guardrail encourages the model to refuse to answer the question directly, the guardrail hurts the defenders, he said.
“This is where the whole offensive vs. defensive and defensive aspect comes in, because ‘fixing this code’ as information is both an important form of defense but also a guide to finding significant vulnerabilities in the code base,” Anley said. Therefore, at the same time, the same tool is an offensive tool and a defensive tool, and the two cannot really be removed.
“It’s like a hammer,” he continued. “You can’t build a house without a hammer. It’s a tool, but it’s also an unstoppable weapon.”
When he and his colleagues encounter such a roadblock, they sometimes fall back on open-source AI models that come with almost no oversight.
Paolo Stagno, chief technology officer at CrowdFense, a well-known company that builds, acquires and sells unknown vulnerabilities to government agencies, agreed with Dowd, saying that AI companies “actually treat customers like children who need to be watched” with their vetted programs and surveillance lines.
Stagno said he and his colleagues use frontier models — but only in reverse engineering. They avoid using AI to help detect vulnerabilities or build helpers, he said, because feeding that work into a cloud-based model risks leaking sensitive vulnerability data or getting it into future training. With that step, he said, they use open source models that work locally, as they don’t rely on sharing data outside of the model.
Giuseppe Cali, a security researcher who finds zero days and improves, said the guardrails do not hinder his work. That’s because he doesn’t use AI for offensive work; instead, he uses it for initial reverse engineering, understanding the code he’s analyzing, and building support tools. For that, he said, AI tools can speed up the process and allow him to focus on detecting vulnerabilities.
“I still want to be the real owner of the pest control and weapons myself and that will not change if all the tools are removed tomorrow,” said Cali. “I’m jealous of my bugs, and I love this game too much to let models play it for me.”
Another researcher at a smartphone parts company, who spoke on the condition of anonymity because he is not authorized to speak to the media, said his employer is not part of Anthropic’s CVP program and as such, its tools are not very helpful in detecting vulnerabilities because the guardrails are so tight.
“When it catches air and we do anything related to security, it just stops and doesn’t work,” said the person.
Chris Thompson – CEO of the cybersecurity firm RemoteThreat and founder of Offensive AI Con, an offensive and AI-focused security event – said that in his experience using border AI models, guardrails can change and work differently every day. That’s true even within the loose boundaries of Anthropic and OpenAI’s tested systems.
“I think the practical impact is spending a lot of time negotiating the model instead of working on the basic security system,” Thompson said. “Instead of doing risk analysis and thinking through exploits, you’re trying to figure out why you’re getting inconsistent results or why the models are polluting the output so much.”
As a result, researchers have relied on or pushed for Chinese open-source models such as GLM – freely downloadable models that can be run locally without testing or usage restrictions – said Thompson.
“You have these responsible researchers being pushed out of US-dominated programs into foreign programs,” he said. “I think it does more harm than good to have these precautions.”
Rather than further tighten restrictions, Thompson called for AI frontier labs to open up their systems, provide responsible access, and hold those who abuse their tools accountable. Otherwise, he argued, defenders will lose the AI race.
“There’s this big storm coming, there’s this big storm surge that’s going to happen at a speed that’s never been seen before,” said Thompson. “But the same security companies and legitimate researchers who are trying to make a difference are being blocked right now.”
If you shop through links in our articles, we may earn a small commission. This does not affect our editorial independence.



