The Tightening Grip of AI Guardrails is Choking Cybersecurity Defenders

For months, leading artificial intelligence developers have implemented stringent vetting processes and robust safety protocols, ostensibly to prevent malicious actors from weaponizing their powerful models. However, these carefully constructed barriers, designed to safeguard against cyber threats, are now inadvertently impeding the critical work of legitimate network defenders and offensive cybersecurity researchers. This unintended consequence is creating a significant bottleneck in the race to identify and neutralize digital vulnerabilities before they can be exploited by adversaries.
The Anthropic Export Control Incident: A Catalyst for Concern
A pivotal moment in this evolving dynamic occurred in June when the U.S. government imposed export control restrictions on Anthropic’s highly anticipated AI models, Mythos and Fable. The official rationale, at least in part, stemmed from a report indicating that these models’ built-in safeguards, intended to prevent their misuse for malicious cyber activities, could be circumvented. This development, regardless of the precise motivations behind the government’s action—whether a genuine concern over AI "jailbreaks" or other strategic considerations—underscored the growing tension between AI’s potential for defense and its inherent risks.
Anthropic had, prior to the restrictions, actively promoted Mythos as a sophisticated tool capable of complex cyber operations, a narrative that positioned it as a "doomsday cybermachine" requiring exceptionally careful deployment and rigorous oversight. The subsequent export controls, which targeted the most powerful versions of these models, highlighted the government’s acute awareness of the dual-use nature of advanced AI. While restrictions on Fable 5 and Mythos 5 have since been lifted, with Fable 5 returning to general access and Mythos 5 being reintroduced under strict government review to vetted U.S. organizations, the initial clampdown sent a clear signal about the perceived dangers and the regulatory landscape.
The Dual-Edged Sword: AI as Both Shield and Sword
The gatekeeping exemplified by Anthropic’s approach is not an isolated phenomenon. Major AI players like OpenAI and Anthropic themselves offer specialized programs designed to grant cybersecurity professionals access to models with relaxed restrictions. OpenAI’s "Trusted Access for Cyber" program and Anthropic’s "Cyber Verification Program" aim to provide vetted researchers with tools they can leverage for defensive purposes.
However, these very guardrails, while intended to promote responsible use, have become a significant point of contention within the cybersecurity community. Researchers whose livelihoods depend on uncovering novel vulnerabilities before they are exploited by criminals find these limitations to be a substantial hindrance.
Mark Dowd, a highly respected security researcher with decades of experience in identifying and selling "zero-day" vulnerabilities to Western governments, has voiced his concerns. He articulated a sentiment shared by many: "it’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not." Dowd’s work, which involves acquiring previously unknown software flaws and the exploits that leverage them, often for intelligence operations rather than immediate patching, highlights a critical aspect of the cybersecurity ecosystem. Governments pay a premium for such vulnerabilities precisely because they remain undisclosed, offering strategic advantages.
The Researcher’s Dilemma: Guardrails Hampering Discovery
The friction between AI developers’ safety mandates and the practical needs of offensive cybersecurity professionals is palpable. Chris Anley, chief scientist at the prominent security consulting firm NCC Group, explained the paradoxical nature of AI in cybersecurity: "asking an AI model to try to exploit a bug is a key step in confirming it’s a real vulnerability worth fixing." Yet, if an AI model, due to its guardrails, refuses to engage with such a request, it directly undermines the defensive effort.
Anley eloquently described this paradox: "’fix this code’ as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base. So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can’t really be unpicked." He likened advanced AI models to a hammer—an indispensable tool for construction, but also inherently a weapon. This analogy underscores the difficulty of bifurcating AI’s capabilities into purely benevolent or malevolent applications.
When faced with these AI-imposed roadblocks, many researchers resort to open-source AI models that come without any restrictions, allowing for unfettered exploration of potential vulnerabilities.
Paolo Stagno, CTO of Crowdfense, a company specializing in the acquisition and sale of unknown vulnerabilities to government agencies, echoed Dowd’s frustration. He characterized the vetted programs and guardrails as an infantilizing approach, where "AI companies essentially treat customers like children who need babysitting." Stagno and his colleagues navigate this landscape by using frontier models primarily for reverse engineering, but they deliberately avoid employing AI for vulnerability discovery or exploit development. The risk of leaking sensitive data or having proprietary vulnerability information absorbed into future training runs of cloud-based models is too significant. For these critical tasks, they rely on locally hosted open-source models, ensuring data remains within their control.
Giuseppe Cali, another security researcher focused on discovering zero-days and developing exploits, presents a slightly different perspective. While guardrails don’t directly impede his work, it’s because he strategically employs AI. He uses these tools not for direct offensive operations, but for initial reverse engineering, to gain a deeper understanding of complex codebases, and to build auxiliary tools that streamline his analysis. In this capacity, AI significantly accelerates his workflow, allowing him to dedicate more time to the nuanced task of vulnerability discovery. Cali firmly believes in retaining ownership of the "actual bug discovery and weaponization," stating, "I am jealous of my bugs, and I like this game too much to let models play it for me."
The Unintended Consequences: Pushing Researchers Towards Foreign Systems
The restrictive nature of AI guardrails is not merely an inconvenience; it has tangible consequences for the efficacy of cybersecurity research. A researcher at a major smartphone-component manufacturer, who requested anonymity, revealed that their employer’s inability to participate in Anthropic’s Cyber Verification Program has rendered their AI tools "barely useful for finding vulnerabilities" due to overly strict guardrails. "If it catches wind we’re doing anything security related, it just stops and isn’t usable," the researcher stated.
Chris Thompson, CEO of cybersecurity firm RemoteThreat and founder of Offensive AI Con, a prominent event focused on offensive security and AI, has observed firsthand the inconsistencies and unpredictability of these guardrails, even within ostensibly more permissive vetted programs. "I think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program," Thompson explained. "Instead of analyzing a vulnerability and reasoning through the exploitability, you’re trying to find why you’re getting inconsistent results or why are models over-sanitizing the output."
This frustration is leading researchers to seek alternatives, often turning to Chinese open-source models like GLM. These models are freely downloadable, can be run locally, and impose no vetting or usage restrictions, offering an unrestricted environment for security research.
"You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems," Thompson observed. He posits that the current approach to guardrails is "more harmful than good."
A Call for Openness and Accountability
Thompson advocates for a fundamental shift in strategy. Instead of intensifying restrictions, he urges AI frontier labs to expand access to their programs, implement mechanisms for responsible usage, and crucially, establish clear lines of accountability for those who abuse their AI tools. Failure to do so, he argues, risks ceding the advantage in the rapidly escalating AI-driven arms race.
"There’s this big storm coming. There’s this big wave of attacks that are going to happen at speed and scale like never before," Thompson warned. "But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now." The current trajectory, he believes, is counterproductive, hindering the very entities tasked with defending against the anticipated surge in sophisticated cyber threats. The implications are far-reaching, potentially impacting national security and the global digital infrastructure if defensive capabilities are not adequately empowered. The ongoing debate underscores a critical juncture where the innovation of AI must be balanced with the practical realities of cybersecurity defense, ensuring that safety measures do not inadvertently become impediments to security itself.







