Open-Weight AI Models Gain Ground, but Safety Measures Lag Behind
A recent report from the AI safety nonprofit SaferAI has highlighted the rapid progress made by open-weight AI models in narrowing the gap with industry leaders. The Chinese open-weight model GLM-5.2, developed by Z.ai, has demonstrated capabilities comparable to those of OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7 in cyber and bio capabilities, according to the report.
However, despite the advancements in capabilities, the divide between frontier capabilities and safety practices remains a significant concern. The report notes that GLM-5.2 refused none of the offensive cyber or dual-use biology tasks it was given, whereas Claude Opus 4.7 consistently refused to complete such tasks, rendering SaferAI’s evaluation of its cybersecurity capabilities incomplete.
The stark reminder of the risks associated with open-weight AI models has sparked debate among policymakers and industry experts. Critics have long warned that open-weight AI models could put highly capable AI into the hands of potential attackers, with no way to police how they use the technology once the weights are downloaded.
The Frontier of Capability vs. the Frontier of Risk
Henry Papadatos, executive director of SaferAI, emphasizes that the frontier of capability is not the same as the frontier of risk. He stresses that to assess the risk properly, one must take into account the state of mitigations as well. The report highlights the importance of considering both the capabilities and safety measures of AI models.
Frontier developers like OpenAI and Anthropic rely on safeguards such as classifiers, refusal training, and API-level controls to limit dangerous cyber and biological assistance. However, these measures are far from foolproof, and jailbreaks can routinely bypass protections on deployed models.
Far.ai, an AI safety nonprofit, has found hundreds of universal jailbreaks in frontier models like xAI’s Grok 4.5 and Google DeepMind’s Gemini 3.1 Pro. These jailbreaks can succeed when attackers combine multiple manipulation techniques, such as roleplaying, authority impersonation, and follow-up prompts, to amplify weak points in a model’s defenses.
The safeguards in place for closed models don’t work at all on open-weight models, which are designed to run on any infrastructure with any set of safeguards or lack thereof.
Pre-Training Data Filtering: A Potential Solution?
One technique that could help is called pre-training data filtering, where an AI company removes offensive cybersecurity information from their training data and then trains the model on the curated dataset. Research suggests that this can reduce hazardous biological knowledge without harming overall model performance.
However, for cybersecurity, data filtering is much less practical. It’s difficult to train a general model that excels at coding without also being a good hacker, as coding has become AI’s biggest moneymaker. Developers face pressure to keep improving those capabilities even as they search for ways to limit misuse.
Frontier developers have increasingly relied on other mitigations, such as selectively restricting the kinds of cybersecurity assistance models will provide. Anthropic’s Opus 5, for example, can search for vulnerabilities in uncompiled source code but not compiled software, per the model’s system card.
Z.ai, the developer of GLM-5.2, did not publish a safety framework, pre-deployment testing commitments, or risk assessment for the model. TechCrunch has asked Z.ai whether it conducted internal or third-party frontier safety evaluations before release, but did not receive a response.
Chinese leaders have increasingly acknowledged the risks of advanced AI. At the World AI Conference last month, Chinese President Xi Jinping emphasized the importance of open-weight models while also stressing the necessity of ensuring AI remains a tool under strict human control.
Graham Webster, who studies Chinese AI policy at the Stanford Cyber Policy Center, notes that China has robust regulations governing AI but those rules have historically focused on politically sensitive content, misinformation, and social stability rather than catastrophic AI risks like offensive cyber capabilities and biological misuse.