Open-weight AI models are advancing, but significant safety concerns persist.
Image Credits:Z.ai
The Evolving Landscape of AI Governance
As the debate continues among policymakers regarding the governance of increasingly sophisticated AI systems like OpenAI’s GPT-5.6 Sol and Anthropic’s Mythos, a prominent open-weight AI model from China, GLM-5.2, is rapidly closing the gap with industry leaders. According to a recent report by the AI safety nonprofit SaferAI, GLM-5.2, developed by Z.ai, lags just a few months behind OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7 in cyber and biological capabilities. However, the growing disparity between frontier capabilities and safety practices raises important concerns.
The Safety Gap in AI Models
SaferAI’s evaluation revealed that GLM-5.2 did not refuse any offensive cyber or dual-use biology tasks presented to it. In stark contrast, Claude Opus 4.7 consistently declined such tasks to the point that SaferAI could not complete the CyberGym assessment, a benchmark designed to evaluate cybersecurity skills. This discrepancy highlights ongoing warnings from critics about the risks associated with open-weight AI models, which could empower potential attackers without any oversight once the weights are downloaded.
“The frontier of capability is not the frontier of risk,” said Henry Papadatos, Executive Director of SaferAI, emphasizing that understanding the state of risk mitigations is essential for effective governance.
Challenges of Open-Weight Models
While Z.ai is able to enforce safety measures on its hosted API, these protections become unenforceable if the model is run on personal hardware. Users can remove or alter any existing safeguards, fine-tune the models, or modify system prompts at will. Meanwhile, industry leaders like OpenAI and Anthropic implement safeguards such as classifiers and API-level controls to curb hazardous cyber and biological assistance.
Despite these measures, vulnerabilities remain. Unauthorized jailbreaks are increasingly common, allowing attackers to bypass the protections of deployed models. Research from Far.ai has identified numerous universal jailbreaks affecting frontier models like xAI’s Grok 4.5 and Google DeepMind’s Gemini 3.1 Pro. Attackers often exploit weak points in a model’s defenses through various manipulation techniques—such as role-playing or authority impersonation.
In contrast, open-weight models lack such safeguards and can operate on any infrastructure, which makes oversight nearly impossible.
Toward Safer AI Practices
Papadatos argues that the goal should be to make only safe capabilities widely available while minimizing the risks associated with dangerous functions. A potential technique he notes is “pre-training data filtering,” where offensive cybersecurity information is excluded from training datasets to lower hazardous biological understanding. While research indicates that this method could enhance safety without sacrificing overall model performance, its application in cybersecurity remains complicated.
Creating a general-purpose model that excels at coding without enabling hacking capabilities presents a significant challenge. Given the financial incentives tied to coding advancements, developers are under pressure to strike a balance between enhancing capabilities and mitigating misuse.
Restrictive Approaches to Cybersecurity Assistance
To counteract potential threats, some frontier companies have begun to selectively limit the types of cybersecurity assistance their models provide. For example, Anthropic’s Opus 5 can identify vulnerabilities in uncompiled code but not in compiled software, thereby reducing its utility for offensive applications. Other strategies include rigorous pre-deployment safety evaluations and withholding model weights perceived as too dangerous.
SaferAI highlighted that Z.ai has not published a safety framework, pre-deployment testing commitments, or risk assessments for GLM-5.2. Inquiries by TechCrunch regarding internal or third-party safety evaluations prior to the model’s release went unanswered.
Chinese AI Regulations: A Double-Edged Sword
Chinese authorities have become increasingly aware of the risks posed by advanced AI technologies. At the recent World AI Conference, President Xi Jinping underscored the significance of open-weight models while insisting that AI must remain under strict human control. Despite China’s rigorous regulatory framework, it historically has focused more on issues like politically sensitive content rather than catastrophic risks related to offensive cyber capabilities or biological misuse.
Graham Webster, a China AI policy researcher from the Stanford Cyber Policy Center, noted that U.S. AI experts generally express more concern about these existential risks compared to their Chinese counterparts. Many in the Chinese regulatory community believe that American companies are more likely to encounter significant risks.
“The Chinese system has confidence that they control the use of these technologies within China,” Webster explained, elaborating that online behavior is often traceable to real identities, which enhances accountability.
The Case for Open-Weight AI
Proponents of open-weight AI models assert that releasing weights is vital for cybersecurity. Such transparency allows companies to bolster defenses against attacks. For instance, Hugging Face leveraged GLM-5.2 to counteract breaches caused by OpenAI. Clem Delangue, CEO of Hugging Face, recently tweeted that the systems that defend against AI-driven cyberattacks can also help identify and fix vulnerabilities.
Even so, Papadatos cautions that the benefits of open-source development should not lead to unrestricted access to dangerous capabilities.
“We shouldn’t simply accept that dangerous capabilities are easily accessible to anyone,” he said, advocating for the prioritization of safe capabilities that can be accessed responsibly. The speed with which attackers adapt and deploy new tools continues to outpace defenders. A ransomware group can change its methods almost overnight, while institutions like hospitals often cannot keep up.
Conclusion
As the capabilities of AI systems evolve, the conversation about governance and safety must also progress. The emergence of models like GLM-5.2 underscores the urgency of addressing the risks posed by open-weight AI. By focusing on responsible governance, rigorous safety evaluations, and selective access, stakeholders can better navigate the challenges associated with these powerful technologies while maximizing their benefits for society.
Thanks for reading. Please let us know your thoughts and ideas in the comment section down below.
Source link
#Openweight #models #catching #frontier #safety #gap #remains
