AI companies have implemented rigorous programs and protective measures to prevent misuse of their models by cybercriminals. However, these restrictions are now creating challenges for legitimate network defenders and offensive cybersecurity researchers.Â
In June, the U.S. government imposed export controls on Anthropic’s AI models, Mythos and Fable, following a report suggesting that these models could be manipulated to conduct cyberattacks. The guardrails intended to prevent such misuse were reportedly bypassable.
Despite the reasons behind this action, Anthropic has consistently marketed Mythos as a highly advanced cyber tool accessible only to thoroughly vetted users, with strict safeguards. The export controls on Fable 5 and Mythos 5 have since been lifted, with Fable 5 becoming generally available on July 1 and Mythos 5 being reintroduced to select U.S. organizations as part of a government review.
This selective access is not exclusive to Mythos. Both Anthropic and OpenAI provide cybersecurity researchers with programs to gain vetted access to models with fewer restrictions: OpenAI’s Trusted Access for Cyber program and Anthropic’s Cyber Verification Program.Â
These guardrails have faced criticism, particularly from researchers focused on identifying system vulnerabilities and devising strategies to address them before they are exploited by criminals.
During a cybersecurity podcast, Mark Dowd, a security researcher, expressed discomfort with large companies making unilateral decisions about security standards. Dowd has a history of discovering and selling “zero-days” — previously unknown software vulnerabilities — to Western governments, instead of reporting them to be fixed. These vulnerabilities are highly valued by governments for intelligence operations.
Though Dowd acknowledges potential bias due to his line of work, he is not alone in his views. Several offensive cybersecurity professionals, who actively search for system weaknesses, shared with JS their experiences of using AI tools and navigating their limitations.Â
Chris Anley, the chief scientist at NCC Group, emphasized that asking an AI model to exploit a bug is crucial for verifying if it’s a vulnerability worth addressing. If the model refuses due to guardrails, it hampers defenders.
Anley explained that the prompt “fix this code” serves both defensive and offensive purposes, making it difficult to separate the two functions. “It’s like a hammer,” he said. “You can’t build a house without a hammer. It’s definitely a tool but also inherently a weapon.”
To overcome these limitations, Anley and his team sometimes use open source AI models without any guardrails.
Paolo Stagno, the chief technology officer at Crowdfense, supported Dowd’s views, arguing that AI companies treat customers like children with their restrictive measures. Stagno and his team use frontier models for reverse engineering but avoid using AI for vulnerability discovery to prevent sensitive data leakage. Instead, they rely on locally run open source models.
Giuseppe Cali, another security researcher, stated that guardrails do not hinder his work. He uses AI for reverse engineering and developing support tools, allowing him to concentrate on finding vulnerabilities. “I still want to own the actual bug discovery and weaponization myself,” he said, adding that he enjoys the challenge too much to let AI take over.
An anonymous researcher at a smartphone-component manufacturer revealed that his company, not part of Anthropic’s CVP program, finds the tools barely useful due to overly strict guardrails that halt operations on any security-related activity.
Chris Thompson, CEO of cybersecurity firm RemoteThreat and founder of Offensive AI Con, noted that the inconsistency of guardrails in frontier AI models complicates work, even within vetted programs. Researchers often turn to Chinese open source models like GLM, which lack such restrictions, he said.
Thompson warned that the guardrails push responsible researchers toward foreign systems, potentially causing more harm than benefit. He urged AI labs to provide responsible access and hold abusers accountable to prevent losing the AI race. “There’s a big storm coming,” he said. “Security consulting firms and legitimate researchers trying to make a difference are being stifled.”
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

