Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. Anthropic Frontier Red Team: open-weights GLM-5.3 nearly…

Anthropic Frontier Red Team: open-weights GLM-5.3 nearly matches Claude Mythos Preview at exploit development, with weak safeguards

★★★★after cutoffpolicy-safetyAnthropicZhipu AIconfidence: high

On Sept 29, 2026 Anthropic's Frontier Red Team published an evaluation of Zhipu's open-weights GLM-5.3 and found exploit-development skills close to Claude Mythos Preview (ExploitBench V8 12% vs 14%; binary exploitation 4% vs 6%, all other tested models 0%). Its safeguards were bypassed 64–100% of the time with simple techniques, and the community had removed them by "abliteration" for $1,200–$4,400 of compute. Anthropic called it "a meaningful step change in the cyber capabilities available to attackers".

Key facts

What happened

Anthropic ran GLM-5.3, released with open weights in August, through the same offensive-cyber evaluations it uses for its own models. On the hardest tasks, building working exploits for browser and binary targets, it came within a few points of Claude Mythos Preview, the model Anthropic had kept restricted to vetted defenders under Project Glasswing. Because the weights are public, its refusals can be bypassed or removed.

Spread of refusal-removed versions (added 2026-10-03)

Within weeks of the weights' release, refusal-removed ("abliterated") GLM-5.3 builds became some of the most downloaded derivatives on Hugging Face. At least one was tuned specifically for offensive security. Several services also host them behind ordinary APIs: abliteration.ai, Venice and NanoGPT. On Oct 3 a startup founder offered such a model free "for cybersec" on X. Hugging Face disabled one repo whose name advertised offensive cyber use, but the model came back under a new name and on a mirror site (entry 2026-09-01-huggingface-disables-offensive-cyber-glm-5-3). None of the sources reviewed here, as of 2026-10-03, reports such a model being used in an actual attack. The evidence is about availability, not documented harm. NIST CAISI's independent assessment (entry 2026-09-17-caisi-glm-5-3-cyber-assessment) rated GLM-5.3 the most cyber-capable open-weight model but about four months behind the US frontier.

Why it matters

It is the first time a frontier lab has published evidence that an open-weights model reached the level of exploit capability it had judged too risky to release widely. That undercuts restricted-release strategies and strengthens the case for pre-release government testing. The source is a competitor's evaluation of a Chinese model, so independent replication would help.

Changelog

  • 2026-09-30: created (sweep 2026-09-29, via Simon Willison's Sept 29 quote)
  • 2026-09-30: added Nathan Lambert critique
  • 2026-10-03: added spread of abliterated GLM-5.3 (HF download counts, OrcaRouter, dealignai, abliteration.ai/Venice/NanoGPT hosting, HF takedown, Naihin free API), Gizmodo analysis, CAISI and Mindgard links

People

Nathan Lambert Simon Willison

Related posts (5)

Related events

  1. Zhipu (Z.ai) releases GLM-5.3, top open-weights coding/agent model ★★★
  2. NIST CAISI: GLM-5.3 is the most cyber-capable open-weight model yet, but trails the US frontier by about four months ★★★★
  3. Hugging Face disables an abliterated GLM-5.3 repo branded "for offensive cyber"; it is re-uploaded under a new name and mirrored on Pirate Face ★★★
  4. Moonshot opens internal review after Mindgard jailbreaks Kimi K2.6 and K3 Swarm into weapons and assassination guidance ★★★
  5. Anthropic reveals Claude Mythos Preview, withholds it over cyber risk and launches Project Glasswing ★★★★★
  6. Z.ai says GLM-5.3 largely built the inference stack that serves GLM-5.3-Flash, calling it an early step toward recursive self-improvement ★★★
  7. Artificial Analysis launches the Cyber Index and an industry alliance (IBM, NVIDIA, Vercel, Collinear) for AI vulnerability-fixing evals ★★
  8. Anthropic Frontier Red Team: frontier models reach superhuman photo geolocation and can write working drone strike software ★★★

Sources (12)

id: 2026-09-29-anthropic-glm-5-3-spread-of-cyber-capabilities · updated 2026-10-03 · open in the interactive timeline