Zhipu's GLM-5.3 scored 84.5% on CyberGym, beating Anthropic's Mythos 5 and OpenAI's gpt-5-6-sol" class="entity-chip">GPT-5.6 Sol in source-code vulnerability detection. But the Chinese model's 54.4% ExploitBench score versus Mythos's 78% reveals a capability gap the headline benchmark hides.
Key facts
- GLM-5.3: 84.5% on CyberGym vs Mythos 83.8%
- ExploitBench: GLM-5.3 54.4% vs Mythos 78%
- 2,436 vulnerabilities found across 269 real projects
- 1,097 vulnerabilities rated medium-to-high severity
- GPT-5.6 Sol scored 83.6% on CyberGym
Beijing-based Zhipu, also known as Z.ai, unveiled GLM-5.3 claiming it beat Anthropic's frontier Mythos 5 model in a key cybersecurity test, as China races to counter Western advances in AI defence. According to the SCMP, GLM-5.3 achieved a success rate of 84.5 per cent on CyberGym, a benchmark measuring whether models can identify and validate security flaws from source code. That edged out Anthropic's Mythos at 83.8 per cent and OpenAI's GPT-5.6 Sol at 83.6 per cent, per Zhipu's own figures.
The gap widens dramatically on the second benchmark. On ExploitBench, which gauges how far models climb the exploitation ladder, GLM-5.3 scored 54.4 per cent — trailing Mythos's 78 per cent and GPT-5.6 Sol's 76.5 per cent. Zhipu did not dispute these numbers, and the company has not published methodology details for either benchmark run.
Detection is not exploitation. GLM-5.3 can find flaws but appears significantly weaker at chaining them into working exploits — the difference between a vulnerability scanner and a penetration tester. Zhipu's real-world validation is more concrete: it tested the model with security teams in China against real codebases, identifying 2,436 vulnerabilities across 269 projects after expert review. Of those, 1,097 were rated medium to high severity, according to the company.
The CyberGym headline is real but partial. Cyber defence procurement increasingly demands end-to-end capability — find, validate, exploit, patch. On that full spectrum, Anthropic's Mythos still leads by a wide margin. Zhipu's edge on detection is notable, but the ExploitBench gap suggests the Chinese model is not yet at Mythos-level for offensive operations, despite the framing in its announcement.
What to watch
Watch for third-party replication of Zhipu's CyberGym and ExploitBench results, and for whether Anthropic responds with a Mythos update specifically targeting vulnerability detection. Also track whether Zhipu ships GLM-5.3 to Chinese government security teams, which would signal real deployment beyond benchmark claims.

Source: scmp.com








