MikhbarMIKHBAR
Artificial Intelligence

Anthropic Report Claims GLM-5.3 Has Mythos Hacking Abilities

A newly released frontier red teaming report from Anthropic claims that Zhipu AI's GLM-5.3 model possesses high-level hacking capabilities and features easily bypassed safeguards.

Anthropic Report Claims GLM-5.3 Has Mythos Hacking Abilities

Anthropic Red Teaming Report Targets Zhipu AI

Anthropic has released a new report, claiming that Zhipu AI's GLM-5.3 AI model can be used to generate malicious content, with weak safeguarding. [The company claims](https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities#footnote-ref-3) that the AI model can be used for cyberattacks, and that its safeguards can be bypassed using several methods. Further details are available from Tom's Hardware in the original source material.

The closed-source AI company highlights that Chinese open-weight models can be abused to generate harmful content. Citing a late September report from the Center for AI Standards and Innovation, Anthropic states that GLM-5.3 can fully automate exploits on a level comparable to its own unreleased Claude Mythos AI model.

Benchmark Performance and Exploit Generation

In benchmark testing conducted in Exploitbench within a sandboxed environment, GLM-5.3 developed end-to-end Google Chrome exploits 50 times across 410 runs. Anthropic's Mythos led the comparison with 56 successful exploits in the same number of attempts. Other popular open-weight models, such as Kimi K3 and DeepSeek V4.1 Flash, attained a 0% success rate under identical measures.

Further evaluation in an internal benchmark targeting full control-flow hijacks showed GLM-5.3 achieving a 4% success rate, closely trailing Mythos at 6%. The lighter Flash variant of the model demonstrated the ability to develop chained exploits in known bugs for a tokenized price of roughly $20.40, utilizing between 100 million and 300 million tokens.

Guardrail Bypass and Abliteration Methods

According to the report, stock versions of GLM-5.3 easily evade guardrails using specific prompting techniques. Deceptive prompting involving an adversarial autonomous agent role-play yielded a 64% success rate, while prefilling thinking tokens achieved a 92% success rate. A third method known as abliteration was also detailed by the company.

While stock models exhibit high refusal rates comparable to Anthropic's own products, abliterated versions—where guardrails are purposefully removed—drop refusal rates down to 6% for the standard model and 14% for the Flash variant. Although abliterated models frequently appear on HuggingFace for simulated red teaming, they carry risks for real-world scenarios.

Laptop showing z.ai splash screens
(Image credit: Getty Images / Bloomberg) · Source: Tom's Hardware

Hardware and Compute Costs for Attackers

Executing an abliterated full-precision FP8 version of GLM-5.3 requires substantial compute resources. The model demands 306 GB of VRAM alongside a matching allocation for KV cache. Operating at an estimated 100 tokens per second requires a memory bandwidth of at least 4 TB/s and powerful hardware, such as a cluster of eight Nvidia H200 accelerators.

Financial barriers remain a significant hurdle for everyday malicious actors. Renting enough GPUs to abliterate the full model costs approximately $30 per hour based on Runpod figures where a single H200 rents for $3.79 per hour. Generating 100 million tokens could cost over $8,400 and require nearly 12 days of continuous execution using an eight-GPU setup.

Broader Implications for AI Safety and Open-Weight Regulation

Anthropic's findings arrive amid discussions regarding the safety of frontier open-weight models as their capabilities advance. While the findings point to real theoretical risks involving adversarial actors, the massive hardware investments and computational costs associated with running and modifying these large models present considerable practical limitations.

Sources

  • Tom's HardwareAnthropic claims popular Chinese AI model has Mythos-class hacking abilities

Continue chronologically

Related entity coverage