China's GLM-5.2, an open-weight AI model from Z.ai, now trails the industry's most advanced systems by only a few months in cyber and biological capabilities, according to a new report from AI safety nonprofit SaferAI published this week. The evaluation reveals that while open-weight models are rapidly closing the performance gap with leaders like OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7, the distance between what these systems can do and how well they're protected from misuse is widening. The report underscores a longstanding concern: open-weight models may deliver powerful AI tools to adversaries who can download the weights and remove any restrictions, with no mechanism to control how the technology gets used once it's released.
SaferAI's evaluation, conducted through Z.ai's public API, found that GLM-5.2 declined none of the offensive cybersecurity or biology tasks it was assigned. In contrast, Claude Opus 4.7 refused requests so reliably that SaferAI couldn't finish running CyberGym—a benchmark measuring cybersecurity abilities that OpenAI used before last month's Hugging Face breach—on the system at all. The report notes that Z.ai did not publish a safety framework, pre-deployment testing commitments, or risk assessment for GLM-5.2. While Z.ai could implement safety measures on its hosted API, those controls become impossible to enforce once someone operates the weights on their own infrastructure, where they can strip out or alter any protections, fine-tune the models, or modify system prompts.
"The frontier of capability is not the frontier of risk, and so we do have to take into account the state of the mitigations as well to assess the risk properly," Henry Papadatos, executive director of SaferAI, told the outlet. The report explains that frontier developers such as OpenAI and Anthropic typically depend on safeguards including classifiers, refusal training, and API-level controls to restrict dangerous cyber and biological assistance, though these measures remain far from perfect—jailbreaks routinely circumvent protections on deployed models. According to the report, jailbreaks succeed when attackers combine multiple manipulation techniques—including roleplaying, authority impersonation, fake conversation history, and follow-up prompts—to amplify weak points in a model's defenses. But the safeguards that exist for closed models don't function at all on open-weight models, which are built to run on any infrastructure with any set of protections—or none.
One technique that could help is pre-training data filtering, where an AI company scrubs offensive cybersecurity information from training data before building the model, Papadatos noted. Some research suggests this approach can cut hazardous biological knowledge without damaging overall model performance. However, for cybersecurity the method is much less practical—it's difficult to train a general model that excels at coding but isn't also skilled at hacking. Because coding has become AI's biggest revenue driver, developers face pressure to keep improving those capabilities even while searching for ways to limit misuse. Because of that dynamic, frontier developers have increasingly turned to other mitigations instead, including selectively restricting the types of cybersecurity help models will provide. Anthropic's Opus 5, for instance, can search for vulnerabilities in uncompiled source code but not compiled software, per the model's system card, with the reasoning that this makes it harder to use Opus 5 for offensive purposes.
The report highlights that Chinese leaders have increasingly acknowledged the risks of advanced AI, with Chinese President Xi Jinping emphasizing the importance of open-weight models at last month's World AI Conference while also stressing the necessity of ensuring AI remains a tool under strict human control. Graham Webster, who studies Chinese AI policy at the Stanford Cyber Policy Center, told the outlet that China has robust regulations governing AI, but those rules have historically focused on politically sensitive content, misinformation, and social stability rather than catastrophic AI risks like offensive cyber capabilities and biological misuse. Webster noted that the Chinese system has confidence in controlling the use of these technologies inside China, where being online is attributed to real names and companies and users can be held accountable. Advocates of open-weight AI argue that releasing the weights matters for cybersecurity because it lets companies defend themselves against attacks—Hugging Face relied on GLM-5.2 to defend itself against OpenAI's breach—and because it helps them better prepare for future threats if they know what's coming. But Papadatos said that benefit is often overstated and doesn't mean "we should open-source dangerous capabilities," stressing that the industry should strive for making only the "good capabilities" easily accessible. "By default attackers adopt new tools faster than defenders do," he said. "For example, a ransomware group can change its methods in a week. A hospital cannot." The arms race between capability and control will likely define the next phase of AI governance, as regulators confront the challenge of managing powerful systems that can be downloaded, modified, and deployed beyond any centralized oversight. Companies releasing open-weight models will face mounting pressure to demonstrate that performance gains don't come at the cost of public safety.

