How AI constraints are hindering offensive cybersecurity researchers’ effectiveness.
Image Credits:Nikolas Kokovlis/NurPhoto / Getty Images
The Impact of AI Guardrails on Cybersecurity Research
In recent months, major AI companies have created strict programs and regulations to limit the potential misuse of their models by malicious hackers. However, these restrictions now pose significant challenges for legitimate network defenders and offensive cybersecurity researchers.
Export Controls on AI Models
In June, the U.S. government implemented export control restrictions on Anthropic’s AI models, Mythos and Fable. This decision was reportedly influenced by concerns that the models’ built-in guardrails could be circumvented, allowing users to execute malicious cyberattacks.
Despite the uncertainty surrounding the motivations for these controls, Anthropic has consistently marketed Mythos as a powerful tool for cybersecurity, accessible only to thoroughly vetted users and equipped with stringent guardrails. The export restrictions on Fable 5 and Mythos 5 have since been lifted, with Fable returning to general public access on July 1. Meanwhile, Mythos 5 is now available only to carefully vetted U.S. organizations as part of a government review.
Limitations of Current Gatekeeping
The gatekeeping seen with Mythos is not an isolated incident; both Anthropic and OpenAI provide programs that allow cybersecurity researchers to apply for access to models with fewer restrictions, namely OpenAI’s Trusted Access for Cyber program and Anthropic’s Cyber Verification Program.
However, these guardrails have drawn widespread condemnation. Researchers who focus on identifying unpatched vulnerabilities in systems have voiced frustrations about the limitations imposed by these AI models.
Concerns Over Arbitrary Regulations
Mark Dowd, a veteran security researcher, recently expressed discomfort with the notion that large companies are arbitrarily determining what constitutes safe practices in cybersecurity. Dowd’s expertise lies in discovering and selling “zero-days”—vulnerabilities that have not been disclosed to software vendors—primarily to government agencies. The ongoing nature of these vulnerabilities makes them useful for intelligence-gathering operations, and Dowd appreciates the complexities involved in evaluating safety standards.
While Dowd acknowledges that his experience may create a bias, many professionals in offensive cybersecurity share his views. They often elaborate on how they utilize AI tools while navigating the challenges presented by existing guardrails.
Offensive vs. Defensive: A Delicate Balance
Chris Anley, chief scientist at the security consultancy NCC Group, explained that requesting AI models to attempt exploiting a bug is essential for confirming whether a vulnerability warrants remediation. However, if guardrails prevent the model from responding, it ultimately obstructs defenders’ efforts.
“This illustrates the intricate balancing act between offensive and defensive cybersecurity,” Anley stated. He compared AI tools to hammers: “You need a hammer to build a house, but the same tool can also be a weapon.”
Faced with such limitations, Anley and his colleagues sometimes revert to open-source AI models that lack restrictive guardrails.
Views from Cybersecurity Experts
Paolo Stagno, CTO at Crowdfense, voiced agreement with Dowd, accusing AI firms of treating clients like children who require constant supervision. While he and his team utilize advanced models for reverse engineering, they avoid employing AI for discovering vulnerabilities or crafting exploits. This is primarily due to concerns over data leakage when working with cloud-based models.
Giuseppe Cali, another security researcher, claimed that guardrails do not hinder his efforts, as he primarily uses AI for initial reverse engineering. He prefers to maintain full control over the discovery and weaponization of vulnerabilities, asserting, “I am jealous of my bugs, and I enjoy this game too much to let models play it for me.”
Struggles with Inconsistent Models
One researcher at a smartphone-component manufacturer, wishing to remain anonymous, found that their employer’s lack of participation in Anthropic’s Cyber Verification Program rendered its tools ineffective for vulnerability discovery. “If it detects any security-related activity, it becomes unusable,” they explained.
Chris Thompson, CEO of cybersecurity firm RemoteThreat, provided additional insights from his experience with frontier AI models. He noted that the guardrails are often inconsistent, sometimes yielding different results even within Anthropic’s and OpenAI’s vetted programs.
Thompson emphasized that the practical ramifications include spending extensive time negotiating with the model rather than focusing on core security tasks, such as analyzing vulnerabilities. “You end up figuring out why the models are providing inconsistent outputs,” he indicated.
The Shift to Open Source
As a result, many researchers are increasingly turning to foreign open-source models like GLM, which allow for local deployment without vetting or usage restrictions. “Responsible researchers are being steered away from U.S.-regulated systems towards foreign models,” Thompson noted, underscoring the detrimental effects of existing guardrails.
A Call for Change
Thompson advocates for a reevaluation of the current landscape, urging AI frontier labs to broaden their access policies and maintain accountability for misuse. He warns that failing to address these issues may jeopardize defenders in the cybersecurity race.
“There’s a major wave of cyberattacks on the horizon, unprecedented in speed and scale,” Thompson warned. “The very security consulting firms and legitimate researchers striving to mitigate these threats are being hampered by current restrictions.”
Conclusion
The ongoing tension between the need for cybersecurity measures and the limitations imposed by AI guardrails presents a complex challenge. While the intention behind these restrictions is to prevent misuse, they are inadvertently stifling legitimate cybersecurity efforts. For the industry to effectively combat emerging threats, a reevaluation of access policies could be essential in empowering researchers and defenders alike to remain vigilant and proactive in the face of evolving cyber threats.
Thanks for reading. Please let us know your thoughts and ideas in the comment section down below.
Source link
#guardrails #impeding #work #offensive #cybersecurity #researchers
