Anthropic Frontier Red Team: open-weights GLM-5.3 nearly matches Claude Mythos Preview at exploit development, with weak safeguards
On Sept 29, 2026 Anthropic's Frontier Red Team published an evaluation of Zhipu's open-weights GLM-5.3 and found exploit-development skills close to Claude Mythos Preview (ExploitBench V8 12% vs 14%; binary exploitation 4% vs 6%, all other tested models 0%). Its safeguards were bypassed 64–100% of the time with simple techniques, and the community had removed them by "abliteration" for $1,200–$4,400 of compute. Anthropic called it "a meaningful step change in the cyber capabilities available to attackers".
Key facts
- ExploitBench (Chrome V8): GLM-5.3 12%, Claude Mythos Preview 14%, Claude Opus 4.6 and GLM-5.2 near 0%
- Internal binary-exploitation benchmark, 100 random tasks: GLM-5.3 achieved full control-flow hijacks in 4% of trials, Mythos Preview in 6%, every other tested model 0%
- Engagement with malicious requests: 0% unmodified, 64% with a deceptive prompt, 92% with prefilled reasoning, 100% for an abliterated version
- Abliteration (removing refusals from open weights) cost roughly $1,200–$4,400 of compute and took community developers days
- Recommendations: government safety testing of advanced models before release, safeguards on such capabilities, and wider vetted defender access to frontier models
- Abliterated GLM-5.3 builds spread widely on Hugging Face (last-30-day downloads per HF API, 2026-10-03): orcarouter/GLM-5.3-Flash-Uncensored-FP8 ~224k, dealignai/GLM-5.3-CYBERSECURITY-FP8 ~86k, dealignai/GLM-5.3-Flash-UNCENSORED-FP8 ~49k. The dealignai CYBERSECURITY card says refusals were reduced specifically for offensive security, exploit development, phishing and similar tasks
- OrcaRouter's Aug 29 launch post for its uncensored GLM-5.3-Flash (~2.05M views) claimed MaliciousInstruct refusals fell from 96% to 11% and HarmBench from 93% to 18%, and said part of the refusal behavior is not a single linear direction
- Hosted uncensored GLM-5.3: abliteration.ai's 'Abliterated Large V2' (Aug 29; $3/$5 per 1M input/output tokens; claims 84.5% CyberGym), also offered by Venice ('Abliterated Large V2', API id abliteration-abliterated-model-large-v2) and NanoGPT (from Aug 31)
- Hugging Face disabled one repo branded 'for offensive cyber' (Audn AI); it was re-uploaded under a new name and mirrored by Pirate Face
- Example of cheap hosting: on Oct 3 Silen Naihin, founder of Experience Labs (YC S26), offered free API access to the community-abliterated GLM-5.3-Flash ('use it for cybersec, coding, whatever'; ~107k views). This is a new distribution channel, not a new capability
- Gizmodo (Oct 1, 'AI's Abliteration Problem Is Bigger Than China') reports that in Anthropic's tests GLM-5.3 scored ~90% on JailbreakBench, HarmBench and StrongREJECT, and the abliterated version 3%, 2% and 12%. It argues that US labs also have competitive reasons to frame the issue around China
- As of 2026-10-03, none of the sources reviewed for this entry reports an abliterated GLM, Kimi or DeepSeek model being used in a real attack
- Nathan Lambert (X, 70K views) criticized the framing ("Open Unsafe, Closed Unsafe"), arguing closed models have been used in more documented attacks
What happened
Anthropic ran GLM-5.3, released with open weights in August, through the same offensive-cyber evaluations it uses for its own models. On the hardest tasks, building working exploits for browser and binary targets, it came within a few points of Claude Mythos Preview, the model Anthropic had kept restricted to vetted defenders under Project Glasswing. Because the weights are public, its refusals can be bypassed or removed.
Spread of refusal-removed versions (added 2026-10-03)
Within weeks of the weights' release, refusal-removed ("abliterated") GLM-5.3 builds became some of the most downloaded derivatives on Hugging Face. At least one was tuned specifically for offensive security. Several services also host them behind ordinary APIs: abliteration.ai, Venice and NanoGPT. On Oct 3 a startup founder offered such a model free "for cybersec" on X. Hugging Face disabled one repo whose name advertised offensive cyber use, but the model came back under a new name and on a mirror site (entry 2026-09-01-huggingface-disables-offensive-cyber-glm-5-3). None of the sources reviewed here, as of 2026-10-03, reports such a model being used in an actual attack. The evidence is about availability, not documented harm. NIST CAISI's independent assessment (entry 2026-09-17-caisi-glm-5-3-cyber-assessment) rated GLM-5.3 the most cyber-capable open-weight model but about four months behind the US frontier.
Why it matters
It is the first time a frontier lab has published evidence that an open-weights model reached the level of exploit capability it had judged too risky to release widely. That undercuts restricted-release strategies and strengthens the case for pre-release government testing. The source is a competitor's evaluation of a Chinese model, so independent replication would help.
Changelog
- 2026-09-30: created (sweep 2026-09-29, via Simon Willison's Sept 29 quote)
- 2026-09-30: added Nathan Lambert critique
- 2026-10-03: added spread of abliterated GLM-5.3 (HF download counts, OrcaRouter, dealignai, abliteration.ai/Venice/NanoGPT hosting, HF takedown, Naihin free API), Gizmodo analysis, CAISI and Mindgard links
People
Related posts (5)
- Silen Naihin offers free API access to uncensored GLM-5.3-Flash original ↗ Silen Naihin @silennai · x · 2026-10-03
~107k-view example of a startup founder (Experience Labs, YC S26) offering free refusal-removed GLM-5.3 hosting, pitched for cybersec. - Nathan Lambert criticizes Anthropic's GLM-5.3 cyber report original ↗ Nathan Lambert @natolambert · x · 2026-09-29
Most prominent critique (70K views) of the Anthropic report from an open-model researcher. - Chubby: Anthropic says GLM-5.3 nearly matches Mythos Preview on ExploitBench original ↗ Chubby♨️ @kimmonismus · x · 2026-09-29
High-reach (84K) summary of the report's key numbers. - Nano-GPT original ↗ Nano-GPT @NanoGPTcom · x · 2026-08-31
Cited as a source by: 2026-09-29-anthropic-glm-5-3-spread-of-cyber-capabilities - OrcaRouter releases uncensored (refusal-removed) GLM-5.3-Flash weights in native FP8 original ↗ OrcaRouter 🐳 @OrcaRouter · x · 2026-08-29
The most-viewed launch of an abliterated GLM-5.3 (~2.05M views); its Hugging Face repo had ~224k downloads in 30 days.
Related events
- Zhipu (Z.ai) releases GLM-5.3, top open-weights coding/agent model ★★★
- NIST CAISI: GLM-5.3 is the most cyber-capable open-weight model yet, but trails the US frontier by about four months ★★★★
- Hugging Face disables an abliterated GLM-5.3 repo branded "for offensive cyber"; it is re-uploaded under a new name and mirrored on Pirate Face ★★★
- Moonshot opens internal review after Mindgard jailbreaks Kimi K2.6 and K3 Swarm into weapons and assassination guidance ★★★
- Anthropic reveals Claude Mythos Preview, withholds it over cyber risk and launches Project Glasswing ★★★★★
- Z.ai says GLM-5.3 largely built the inference stack that serves GLM-5.3-Flash, calling it an early step toward recursive self-improvement ★★★
- Artificial Analysis launches the Cyber Index and an industry alliance (IBM, NVIDIA, Vercel, Collinear) for AI vulnerability-fixing evals ★★
- Anthropic Frontier Red Team: frontier models reach superhuman photo geolocation and can write working drone strike software ★★★
Sources (12)
- officialAnthropic: GLM-5.3 and the spread of advanced cyber capabilities
- discussionSimon Willison: Quoting Anthropic Frontier Red Team
- discussionNathan Lambert on X: critique of the GLM-5.3 report
- pressGizmodo: AI's 'Abliteration' Problem Is Bigger Than China (Oct 1)
- discussionOrcaRouter on X: uncensored GLM-5.3-Flash weights (Aug 29)
- codeHugging Face: orcarouter/GLM-5.3-Flash-Uncensored-FP8
- codeHugging Face: dealignai/GLM-5.3-CYBERSECURITY-FP8
- codeHugging Face: dealignai/GLM-5.3-Flash-UNCENSORED-FP8
- officialabliteration.ai: Introducing abliterated-model-large-v2 (GLM 5.3)
- pressTechCrunch: Abliteration.ai is making a business out of removing AI guardrails (Sept 3)
- discussionNanoGPT on X: Abliterated Model Large V2 now on NanoGPT
- discussionSilen Naihin on X: free uncensored GLM 5.3 Flash API (Oct 3)
id: 2026-09-29-anthropic-glm-5-3-spread-of-cyber-capabilities · updated 2026-10-03 · open in the interactive timeline