Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

A developer inspects a wall of code on a large monitor, highlighting a vulnerability marker beside a performance…
Products & LaunchesBreakthroughScore: 90

Zhipu GLM-5.3 Beats Anthropic Mythos on CyberGym, Lags ExploitBench

Zhipu's GLM-5.3 edged Anthropic's Mythos on CyberGym detection (84.5% vs 83.8%) but trailed badly on ExploitBench (54.4% vs 78%), exposing an exploitation gap.

·7h ago·3 min read··7 views·AI-Generated·Report error
Share:
Source: scmp.comvia scmp_tech, bloomberg_techSingle Source
How did Zhipu's GLM-5.3 perform against Anthropic's Mythos 5 on cybersecurity benchmarks?

Zhipu's GLM-5.3 scored 84.5% on CyberGym, beating Anthropic's Mythos 5 (83.8%) and OpenAI's GPT-5.6 Sol (83.6%) in source-code vulnerability identification. However, GLM-5.3 trailed badly on ExploitBench at 54.4% versus Mythos's 78%, exposing a gap in exploitation capability.

TL;DR

GLM-5.3 scores 84.5% on CyberGym, topping Mythos 5 · Zhipu trails on ExploitBench: 54.4% vs Mythos 78% · Model found 2,436 real-world vulnerabilities across 269 projects

Zhipu's GLM-5.3 scored 84.5% on CyberGym, beating Anthropic's Mythos 5 and OpenAI's gpt-5-6-sol" class="entity-chip">GPT-5.6 Sol in source-code vulnerability detection. But the Chinese model's 54.4% ExploitBench score versus Mythos's 78% reveals a capability gap the headline benchmark hides.

Key facts

  • GLM-5.3: 84.5% on CyberGym vs Mythos 83.8%
  • ExploitBench: GLM-5.3 54.4% vs Mythos 78%
  • 2,436 vulnerabilities found across 269 real projects
  • 1,097 vulnerabilities rated medium-to-high severity
  • GPT-5.6 Sol scored 83.6% on CyberGym

Beijing-based Zhipu, also known as Z.ai, unveiled GLM-5.3 claiming it beat Anthropic's frontier Mythos 5 model in a key cybersecurity test, as China races to counter Western advances in AI defence. According to the SCMP, GLM-5.3 achieved a success rate of 84.5 per cent on CyberGym, a benchmark measuring whether models can identify and validate security flaws from source code. That edged out Anthropic's Mythos at 83.8 per cent and OpenAI's GPT-5.6 Sol at 83.6 per cent, per Zhipu's own figures.

The gap widens dramatically on the second benchmark. On ExploitBench, which gauges how far models climb the exploitation ladder, GLM-5.3 scored 54.4 per cent — trailing Mythos's 78 per cent and GPT-5.6 Sol's 76.5 per cent. Zhipu did not dispute these numbers, and the company has not published methodology details for either benchmark run.

Detection is not exploitation. GLM-5.3 can find flaws but appears significantly weaker at chaining them into working exploits — the difference between a vulnerability scanner and a penetration tester. Zhipu's real-world validation is more concrete: it tested the model with security teams in China against real codebases, identifying 2,436 vulnerabilities across 269 projects after expert review. Of those, 1,097 were rated medium to high severity, according to the company.

The CyberGym headline is real but partial. Cyber defence procurement increasingly demands end-to-end capability — find, validate, exploit, patch. On that full spectrum, Anthropic's Mythos still leads by a wide margin. Zhipu's edge on detection is notable, but the ExploitBench gap suggests the Chinese model is not yet at Mythos-level for offensive operations, despite the framing in its announcement.

What to watch

Watch for third-party replication of Zhipu's CyberGym and ExploitBench results, and for whether Anthropic responds with a Mythos update specifically targeting vulnerability detection. Also track whether Zhipu ships GLM-5.3 to Chinese government security teams, which would signal real deployment beyond benchmark claims.

Zhipu says it has tested GLM-5.3 with security teams in China against real-world codebases, identifying 2,436 vulnerabilities across 269 projects afte


Source: scmp.com


Sources cited in this article

  1. Zhipu's
Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from 1 verified source, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

The CyberGym result is a genuine milestone — the first time a Chinese model has publicly topped a US frontier model on a security benchmark. But Zhipu's own numbers undermine the 'Mythos-level edge' framing in its announcement. A 23.6-point ExploitBench deficit is not a rounding error; it suggests fundamentally weaker capability in chaining vulnerabilities into exploits, which is what offensive security operations actually require. The structural read: detection benchmarks like CyberGym reward broad pattern recognition in source code, which plays to China's strength in large-scale data and training compute. Exploitation requires deep reasoning about runtime behavior, memory layouts, and system call sequences — areas where US labs like Anthropic have invested heavily in agentic reasoning. The gap mirrors the broader pattern in AI: Chinese models are closing the gap on curated benchmarks while still lagging on open-ended reasoning tasks. The real signal here is deployment intent. Zhipu's claim of testing with 'security teams in China' against real codebases — 2,436 vulnerabilities across 269 projects — is the kind of operational validation that matters more than any benchmark. If GLM-5.3 gets deployed in Chinese government or military cyber units, the CyberGym score becomes a procurement data point, not just a marketing one. Western labs should watch that deployment trajectory more closely than the benchmark leaderboard.
Compare side-by-side
Anthropic vs OpenAI
Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in Products & Launches

View all