Hustler Words – The artificial intelligence revolution, while promising unprecedented advancements, is inadvertently creating a complex dilemma for the cybersecurity community. Major AI developers have implemented stringent "guardrails" and vetted access programs, primarily designed to prevent their powerful models from being exploited by malicious actors. However, these very safeguards are now facing criticism for significantly impeding the critical work of legitimate offensive cybersecurity researchers and network defenders, raising questions about the balance between preventing misuse and fostering innovation in digital security.
The core issue, as highlighted by experts, is that the tools essential for proactive defense often mirror those used for offense. Researchers whose mission is to uncover unknown system vulnerabilities and develop exploitation methods before criminals do, find themselves constrained by these restrictions.
Mark Dowd, a veteran security researcher known for discovering and selling "zero-days" – previously unknown software flaws and their exploits – to Western governments, voiced his discomfort on a recent cybersecurity podcast. He expressed concern over "random large companies making arbitrary decisions about what is safe in security and what’s not." Dowd, whose work involves identifying these critical weaknesses, acknowledges a potential bias but is far from alone in his sentiment. Indeed, as professionals in offensive cybersecurity shared with Hustler Words, navigating these AI tools and their inherent guardrails presents a constant challenge.

Related Post
Chris Anley, chief scientist at the security consulting firm NCC Group, likens AI models to a "hammer." He explains that trying to exploit a bug with an AI is a crucial step in confirming a vulnerability’s existence and severity. "Fix this code," a seemingly defensive prompt, also serves as a "roadmap for finding critical vulnerabilities." Anley argues that the offensive and defensive capabilities of such tools are inextricably linked, and guardrails that prevent exploration ultimately harm defenders. When faced with such AI roadblocks, Anley and his team often resort to open-source AI models that lack these restrictions.
This gatekeeping isn’t isolated. Companies like Anthropic, with models such as Mythos and Fable, and OpenAI, offer specialized programs like OpenAI’s Trusted Access for Cyber and Anthropic’s Cyber Verification Program. These initiatives allow vetted researchers access to models with fewer cybersecurity restrictions. Yet, even within these programs, challenges persist.
The U.S. government’s temporary export control restrictions on Anthropic’s Mythos and Fable models in June, reportedly spurred by claims of guardrail bypasses, underscored the high stakes involved. While these controls have since been largely lifted, Anthropic’s prior marketing of Mythos as a "doomsday cybermachine" accessible only to carefully vetted users with strict guardrails, illustrates the industry’s cautious approach.
Paolo Stagno, CTO at CrowdFense, a company specializing in acquiring and selling vulnerabilities to government agencies, echoes Dowd’s frustration. He believes AI companies "essentially treat customers like children who need babysitting" with their restrictive programs. While Stagno’s team utilizes frontier models for reverse engineering, they deliberately avoid using cloud-based AI for vulnerability discovery or exploit building. This precaution is to prevent sensitive data leaks or its absorption into future AI training datasets, opting instead for locally run open-source models for such critical tasks.
Not all researchers find guardrails to be an impediment. Giuseppe Cali, a security researcher focused on zero-days, primarily uses AI for initial reverse engineering and building supporting tools, not for the core offensive work itself. He states, "I still want to own the actual bug discovery and weaponization myself… I am jealous of my bugs, and I like this game too much to let models play it for me." For Cali, AI speeds up the analytical process, allowing him to concentrate on the human element of discovery.
However, for others, the impact is significant. An anonymous researcher at a smartphone-component manufacturer, whose company isn’t part of Anthropic’s CVP, described their experience: "If it catches wind we’re doing anything security related, it just stops and isn’t usable."
Chris Thompson, CEO of RemoteThreat and founder of Offensive AI Con, points out the inconsistency of these guardrails, even within vetted programs. He notes that researchers often spend "a lot of time negotiating with the model instead of working on the core security program," struggling with inconsistent results or over-sanitized outputs.
This frustration, Thompson warns, is pushing responsible researchers away from U.S.-governed systems towards foreign-owned, open-source models like GLM, which offer no vetting or usage restrictions. He argues that these guardrails are "more harmful than good," creating a scenario where "responsible researchers… are being pushed away from U.S.-governed systems to foreign-owned systems."
Thompson advocates for AI frontier labs to open their programs further, provide responsible access, and hold abusers accountable, rather than tightening restrictions. He cautions that without this shift, defenders risk losing the "AI race" against a coming "big wave of attacks that are going to happen at speed and scale like never before," while legitimate researchers are "stifled." The cybersecurity community grapples with this critical juncture, seeking a path that leverages AI’s power for defense without inadvertently disarming its most skilled practitioners.






Leave a Comment