AI Guardrails: A Double-Edged Sword in Cybersecurity
In the world of cybersecurity, the use of AI models has become increasingly prevalent. These models are designed to help researchers and defenders find vulnerabilities in systems and devise ways to exploit them before criminals do. However, the introduction of AI guardrails has raised concerns among the cybersecurity community.
AI guardrails are designed to limit the use of AI models by malicious hackers. However, these limits are now hindering the work of legitimate network defenders, as well as that of offensive cybersecurity researchers. The U.S. government has slapped export control restrictions on Anthropic’s AI models, Mythos and Fable, due to concerns that they could be used to build and execute malicious cyberattacks.
Anthropic has repeatedly marketed Mythos as a powerful AI model that can only be given to carefully vetted users, with strict guardrails in place. However, this approach has been criticized by researchers, who argue that it is not up to large companies to decide what is safe in security and what’s not.
Cybersecurity Researchers Speak Out Against AI Guardrails
Cybersecurity researchers, including Mark Dowd, a well-known security researcher, have spoken out against AI guardrails. Dowd argued that it’s not comfortable to have random large companies making arbitrary decisions about what is safe in security and what’s not. He also pointed out that governments pay a premium for vulnerabilities precisely because they stay open, which is useful for intelligence operations.
Chris Anley, the chief scientist at security consulting giant NCC Group, also spoke out against AI guardrails. He argued that asking an AI model to try to exploit a bug is a key step in confirming it’s a real vulnerability worth fixing. However, if a guardrail prompts the model to refuse to answer the question outright, the guardrail hurts defenders.
The Impact of AI Guardrails on Cybersecurity Research
The impact of AI guardrails on cybersecurity research is significant. Researchers rely on AI models to find vulnerabilities and devise ways to exploit them. However, the introduction of AI guardrails has made it difficult for them to do their job effectively. Some researchers have even resorted to using open-source AI models that come with no guardrails at all.
Paolo Stagno, the chief technology officer at CrowdFense, a well-known company that develops, acquires, and sells unknown vulnerabilities to government agencies, argued that AI companies essentially treat customers like children who need babysitting. He also pointed out that feeding AI work into a cloud-based model risks leaking sensitive vulnerability data or having it absorbed into future training runs.
The Future of AI in Cybersecurity
The future of AI in cybersecurity is uncertain. While AI guardrails are designed to limit the use of AI models by malicious hackers, they are also hindering the work of legitimate network defenders and offensive cybersecurity researchers. The U.S. government has slapped export control restrictions on Anthropic’s AI models, Mythos and Fable, due to concerns that they could be used to build and execute malicious cyberattacks.
Cybersecurity researchers are calling for the AI frontier labs to open up their programs, provide responsible access, and hold those who abuse their tools accountable. Otherwise, they argue that defenders will lose the AI race.
As the cybersecurity landscape continues to evolve, it’s clear that AI guardrails are a double-edged sword. While they are designed to limit the use of AI models by malicious hackers, they are also hindering the work of legitimate network defenders and offensive cybersecurity researchers.
The impact of AI guardrails on cybersecurity research is significant, and it’s clear that the future of AI in cybersecurity is uncertain. As the cybersecurity landscape continues to evolve, it’s essential to strike a balance between security and innovation.