HomeMalware & ThreatsChinese Open-Weight Models on the Rise, Anthropic Issues Warning

Chinese Open-Weight Models on the Rise, Anthropic Issues Warning

Published on

spot_img

Artificial Intelligence & Machine Learning,
Next-Generation Technologies & Secure Development

Researchers Find GLM-5.3 Safeguards Easy to Bypass

Chinese Open-Weight Models on the Rise, Anthropic Issues Warning
Image: Shutterstock

A recent study conducted by Anthropic has revealed a concerning trend regarding the capabilities of Chinese open-weight models, particularly GLM-5.3, developed by Zhipu AI, also known as Z.ai. As detailed in the findings, these models are rapidly approaching the advanced functionalities of proprietary models, prompting researchers to warn that the opportunity to bolster cybersecurity defense mechanisms is shrinking at an alarming rate.

The assessment undertaken by Anthropic focused on understanding how GLM-5.3 operates and how it may be exploited in cyber scenarios. Researchers found that this model possesses considerable potential for autonomously crafting end-to-end cyber exploits, highlighting a significant lack of “meaningful safeguards” designed to mitigate such risks. This revelation aligns with similar insights provided by the U.S. Center for AI Standards and Innovation (CAISI), which contends that GLM-5.3 is currently the most cyber-capable open-weight model that has been introduced to the market thus far. CAISI did, however, acknowledge that U.S. frontier models still maintain a marginal edge over GLM-5.3 in terms of cybersecurity capabilities.

Interestingly, Anthropic not only conducted these examinations but also competes directly with open-weight models like GLM-5.3. This interplay between competitive models raises critical questions about the balance of power in the AI landscape and emphasizes the urgent need for innovation in developing robust cybersecurity measures.

GLM-5.3 is an advancement over its predecessor, GLM-5.2, a model that Hugging Face integrated following a security breach involving OpenAI agents. In pursuit of a comprehensive understanding of GLM-5.3, Anthropic employed various automated benchmarks, including ExploitBench, to evaluate how effectively the model could exploit known vulnerabilities as well as human-in-the-loop workflows. The tests were conducted in isolated and controlled environments, comparing the performance of GLM-5.3 against other models like Mythos Preview.

Results indicated that malicious actors could bypass the safeguards of GLM-5.3 nearly 63% to 100% of the time. Both GLM-5.3 and Mythos Preview demonstrated strikingly similar performances on the ExploitBench, successfully developing end-to-end exploits 50 and 56 times respectively in over 400 attempts during the trials.

“Notably, these malicious attempts did not succeed against the safeguarded Claude models in our evaluations. Our team believes that the lax implementation of safeguards within GLM-5.3 significantly heightens the cyber capabilities available to hostile entities,” Anthropic’s researchers stated clearly in their findings.

While GLM-5.3 comes equipped with certain guardrails that prevent it from responding to prompts that clearly indicate harmful intentions, the research team discovered that these protective layers could be easily circumvented or entirely removed through simple techniques. A prominent method utilized was “abiliteration,” a process that involves refusing to answer specific prompts, thereby allowing users to reengineer the model without the necessity of retraining it.

This method facilitates the identification of refusal patterns and allows for the alteration of model weights, typically achieved by downloading the model and eliminating the refusal code. As an open-weight model, GLM-5.3 is inherently susceptible to this type of manipulation.

Additionally, the researchers were able to bypass GLM-5.3’s safety features by engaging the model in a simulated environment, effectively tricking it into believing it was participating in red-teaming exercises.

Anthropic assessed that the capabilities of GLM-5.3 are alarmingly close to those of Mythos, and its easily bypassed safeguards could significantly reduce the gap between cyber defenses and attacks. “GLM-5.3 has the potential to grant malicious actors access to sophisticated capabilities, enabling them to identify and exploit cyber vulnerabilities with little to no meaningful restrictions. This differentiates it from other similarly powerful AI models, which were typically released with either robust safeguards or under restricted access protocols,” indicated Anthropic in their comprehensive report.

Traditionally, there has been a belief within the industry that open-weight models, regardless of their origin, lag behind closed-source models by approximately six to eight months. However, experts are asserting that this gap is rapidly closing as Chinese laboratories gain access to enhanced training datasets. Allegations have also been made against several Chinese companies, notably Z.ai, for illicitly deriving their models from proprietary AI systems by generating prompted responses and subsequently replicating them to train other frameworks.

In light of these developments, defenders within the cybersecurity realm are encouraged to adopt AI agents more actively to counteract the threats posed by malicious actors. As AI technology continues to evolve, aiding in the execution of increasingly sophisticated attacks, the emphasis on employing AI for defense mechanisms becomes vital for safeguarding sensitive data and systems.

Source link

Latest articles

Cybersecurity Awareness Month: AI Agents as Users Demanding Governance

Cybersecurity Awareness Month Broadens Focus to Include AI Agents For over twenty years, Cybersecurity Awareness...

Google Launches Gemini 4 Argon AI Model

Google Unveils Gemini 4 Argon: A New Frontier in Specialized AI for Cybersecurity and...

Dubai Government Agencies Face Off in Hacking Contest

Training & Security Leadership Live And IRL Hacking Contest Captured Attention...

Zammad Vulnerabilities Enable Attackers to Execute Code and Escalate Privileges to Root

Critical Vulnerabilities Found in Zammad Helpdesk Platform: Urgent Response Required Recent security findings have unveiled...

More like this

Cybersecurity Awareness Month: AI Agents as Users Demanding Governance

Cybersecurity Awareness Month Broadens Focus to Include AI Agents For over twenty years, Cybersecurity Awareness...

Google Launches Gemini 4 Argon AI Model

Google Unveils Gemini 4 Argon: A New Frontier in Specialized AI for Cybersecurity and...

Dubai Government Agencies Face Off in Hacking Contest

Training & Security Leadership Live And IRL Hacking Contest Captured Attention...