Safety#AI safety

Anthropic threat report flags GLM-5.3 as the most cyber-capable open-weight model

Anthropic calls Zhipu open-weight GLM-5.3 the most cyber-capable open-weight model: 50 on ExploitBench, with safeguards trivially bypassed.

A metal door locked with a rusty chain and padlock

Anthropic published a threat report on September 30 naming Zhipu’s open-weight GLM-5.3 (released in August) the most cyber-capable open-weight model yet, with safeguards that are easily bypassed. QbitAI’s take: it reads like an advertisement for GLM — the report puts attack and defense numbers side by side.

Facts

  • ExploitBench: GLM-5.3 scored 50 (of 410 attempts) versus 56 for Anthropic’s own Claude Mythos Preview.
  • Third-party view: NIST/CAISI had already called GLM-5.3 the most cyber-capable open-weight model on September 17, about four months behind the US frontier.
  • Guardrail bypass: abliterated weights, produced for roughly $4,400, cut refusal rates from over 90% to low single digits.
  • Reproduced case: researchers built an ARM64 attack chain (Chrome CVE-2026-11645) in about 8 hours for $20.40 of API spend with GLM-5.3.
  • Context: follows Anthropic’s September 10 report accusing seven Chinese labs, including Zhipu, of distillation; the report does not dispute the legality of the open release.

Editorial take

One report, both sides of open weights: the capability is real and so is the fragility — that is the normal condition of open-weight models, not a Zhipu quirk. For model-admission security reviews, the two reusable numbers are the ExploitBench line and the $4,400 ablation cost. Remember the source is a competitor lab’s threat report; cross-reference how open-weight models score on agent benchmarks and read it both as marketing and as data.