The Chinese lab says GLM-5.3 now leads every model it tested at finding software vulnerabilities, and it will not publish the downloadable version until safety testing and hardening are done, about two weeks after the August 14 launch.
Z.ai released its new model, GLM-5.3, on August 14 through its paid coding service and held back the model’s core files, the version companies and individuals could download, change, and run on their own hardware. The company said it will publish the GLM-5.3 open weights, as those files are called, about two weeks after launch, once safety evaluation and hardening are complete.
Z.ai said it delayed the release because of the model’s cybersecurity abilities.
Z.ai says the model leads at finding software flaws
The company reported that GLM-5.3 scored higher than any other model it tested on CyberGym, an AI vulnerability discovery test that gives a model a program’s source code and checks whether it can locate a vulnerability and prove it’s real by making the program fail. GLM-5.3 scored 84.5% across 1,507 tasks, up from 77.2% for the previous version, GLM-5.2. Z.ai reported Anthropic’s Mythos 5 at 83.8% and OpenAI’s GPT-5.6 Sol at 83.6% on the same test.
The company says the ability grew faster than it expected
Z.ai added vulnerability discovery data and practice environments to the training mix and expected the model to get better at spotting individual flaws. It said the model went further than that. According to the company, GLM-5.3 “began to reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains,” meaning it strung individual weaknesses together into a working sequence rather than reporting them one at a time.
Z.ai also said the gains were largest on the tests that sit furthest along that chain. “Capability is growing fastest exactly where we are furthest behind,” the company wrote.
GLM-5.3 still trails Anthropic and OpenAI at turning flaws into attacks
On ExploitBench, which asks a model to reason about real vulnerabilities and how they would be exploited, GLM-5.3 scored 54.4%, more than double GLM-5.2’s 24.4%. Z.ai reported Mythos 5 at 78.0% and GPT-5.6 Sol at 76.5% on the same test.
On ExploitGym, which counts how many exploitation tasks a model finishes within a fixed time budget, GLM-5.3 completed 105 tasks in two hours and 130 in six, compared with 29 and 39 for GLM-5.2. Z.ai reported Mythos 5 at 181 and 247.
Z.ai reports 2,436 real vulnerabilities across 269 projects
The company said it has been running its models against real codebases with several security teams in China. After expert review, screening, and removing duplicates, it counted 2,436 vulnerabilities across 269 open-source projects, including 1,097 rated medium to high severity, of which 107 were critical.
Many had gone unnoticed for years. Z.ai said the oldest was introduced in 1981 and that, on average, a vulnerability lasted 26.6 years before the model found it.
Z.ai publishes a running record of this work, the Z.ai Security Disclosure Ledger. As of launch, it listed 53 vulnerabilities as publicly disclosed and 2,383 still under embargo while the affected projects work on fixes.
The cyber ability came out of training for coding
GLM-5.3 uses the same underlying model as GLM-5.2, and Z.ai said every gain came from the training it did afterward, mostly on long software engineering jobs.
The coding tests the company used all work the same way. Each one hands the model a real programming task and a working computer, then checks whether it finishes the job. On Terminal Bench 3.0, a public test of that kind, Z.ai said GLM-5.3 completed 28.3% of the tasks, up from 4.6% for GLM-5.2. On the company’s own version, which it does not publish, GLM-5.3 completed 34.5% against 23.4%, the 50% improvement Z.ai claims. It said Anthropic’s Claude Fable 5 completed 39.5% on that same unpublished test.
Z.ai said the downloadable files will follow about two weeks after the August 14 launch. Until then, GLM-5.3 runs only through the company’s paid coding plan and the coding tools connected to it, including ZCode, Claude Code, and OpenCode. Every number above is Z.ai’s own, and no outside party can check the results on its own hardware until the files are published.
Z.ai is not the first company to restrict how a model reaches the public over cybersecurity questions. Anthropic released Mythos 5, built to find software vulnerabilities and work out how they could be exploited, only to partners in its Project Glasswing program, with wider access planned for vetted organizations through an arrangement with the US government. In its August 2026 risk report, Anthropic said it has no current plans to release Model 2 outside the company, and raised its rating of the chance a model acts against the company’s interests from very low to low after models took unauthorized actions during cybersecurity tests.

