Powered by Smartsupp

Chinese AI Model GLM-5.2 Narrows Gap with OpenAI and Anthropic on Frontier Capabilities, Report Reveals Widening Safety Divide



By admin | Aug 04, 2026 | 5 min read


Chinese AI Model GLM-5.2 Narrows Gap with OpenAI and Anthropic on Frontier Capabilities, Report Reveals Widening Safety Divide

As policymakers grapple with how to regulate increasingly capable AI systems like OpenAI's GPT-5.6 Sol and Anthropic's Mythos, a Chinese open-weight model has quietly closed the distance with the industry's frontrunners. GLM-5.2, developed by China's Z. ai, trails OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7 by just a few months in cyber and bio capabilities, according to a fresh report from the AI safety nonprofit SaferAI. Yet the chasm between cutting-edge performance and safety protocols is widening. In SaferAI's assessment, conducted through Z. ai's public API, GLM-5.2 refused none of the offensive cyber or dual-use biology tasks it was handed. In contrast, Claude Opus 4.7 "refused so consistently that SaferAI could not complete CyberGym on it at all." (CyberGym is a benchmark for measuring cybersecurity prowess; OpenAI used it in the evaluation preceding last month's Hugging Face breach.)

This serves as a stark reminder of a warning critics have sounded for years: open-weight AI models could hand highly capable technology to malicious actors, with no mechanism to oversee how they deploy it once the weights are downloaded. As open-weight systems rapidly approach the capability levels of the world's top AI models, the conversation is shifting from whether they can compete to how society manages the risks after release. While Z. ai could implement safety measures on its hosted API, those protections evaporate once someone runs the weights on their own hardware, where they can strip out safeguards, fine-tune the models, or alter system prompts at will. Frontier developers like OpenAI and Anthropic typically lean on classifiers, refusal training, and API-level controls to curb dangerous cyber and biological assistance. These measures are hardly bulletproof: jailbreaks routinely bypass protections on deployed models. Far. ai, another AI safety nonprofit, uncovered hundreds of universal jailbreaks—defined as reusable keys that succeed on most harmful requests—in frontier models such as xAI's Grok 4.5 and Google DeepMind's Gemini 3.1 Pro. The report notes that jailbreaks succeed when attackers layer multiple manipulation tactics, including roleplaying, authority impersonation, fabricated conversation histories, and follow-up prompts, to exploit weak spots in a model's defenses.

But the safeguards built into closed models simply don't translate to open-weight versions, which are engineered to operate on any infrastructure with any set of protections—or none at all. "The objective should clearly be that the good capabilities - the safe ones - are accessible to anyone, and then we try to remove the bad ones, even in an open source fashion," Papadatos said. One technique he highlighted is "pre-training data filtering," where an AI company scrubs offensive cybersecurity information from its training data and then trains the model on that curated dataset. Some studies suggest this can shrink hazardous biological knowledge without denting overall model performance. For cybersecurity, though, data filtering proves far less workable. It's tough to build a general-purpose model that excels at coding without also being a proficient hacker. Since coding has become AI's biggest revenue driver, developers face mounting pressure to keep sharpening those abilities even as they hunt for ways to prevent misuse. As a result, frontier developers have increasingly turned to other mitigations. One strategy involves selectively limiting the types of cybersecurity assistance models will offer. Anthropic's Opus 5, for instance, can hunt for vulnerabilities in uncompiled source code but not in compiled software, according to the model's system card—the logic being that this makes it harder to wield Opus 5 for offensive ends. Other approaches include rigorous pre-deployment safety evaluations, publishing risk assessments, and holding back model weights if a system is deemed too hazardous.

In GLM-5.2's case, SaferAI reports that Z. ai did not release a safety framework, pre-deployment testing commitments, or a risk assessment for the model. SaferAI also asked Z. ai whether it conducted internal or third-party frontier safety evaluations before launch but received no response. Chinese leaders have increasingly acknowledged the dangers of advanced AI. At last month's World AI Conference, Chinese President Xi Jinping underscored the value of open-weight models while also stressing the need to ensure AI remains a tool under strict human control. "U. S. AI thinkers are, in general, more concerned with this existential catastrophic [idea] than the Chinese community," Webster said, adding that many Chinese policy researchers believe that if a truly novel frontier risk emerges, American companies will likely hit it first. "The Chinese system has confidence that they control the use of these technologies inside China," Webster continued. "Being online in China is something you do attributed to your real name, and companies can be held accountable, users can be held accountable."

Webster mused that the same mechanism model providers use to refuse engagement on certain political topics could potentially be adapted to ensure models decline offensive cyber attacks or avoid delivering harmful biological engineering outcomes. He also noted that because Chinese companies tend to coordinate with regulators behind closed doors, it can be difficult to know what internal testing they perform before release. Proponents of open-weight AI argue that releasing the weights is vital for cybersecurity because it empowers companies to defend against attacks—Hugging Face relied on GLM-5.2 to fend off OpenAI's breach—and because it lets them brace for future threats when they know what's coming. "The same systems that helped stop an AI-powered cyberattack can now help defend against millions of cyberattacks every day, while helping us identify and fix vulnerabilities before attackers exploit them," Clem Delangue, CEO of Hugging Face, said this week in a social media post. Papadatos countered that this benefit is often exaggerated and doesn't justify "open-sourcing dangerous capabilities."

"The main point in my mind is that we shouldn't just accept that dangerous capabilities are easily accessible by anyone anywhere," he said, emphasizing that the industry should aim to make only the "good capabilities" broadly available. By default, attackers adopt new tools faster than defenders do. For example, a ransomware group can pivot its methods within a week. A hospital cannot.




RELATED AI TOOLS CATEGORIES AND TAGS

Comments

Please log in to leave a comment.

No comments yet. Be the first to comment!