Z.AI released GLM-5.3 on August 14 without the weights. The Beijing company, formerly Zhipu, put GLM-5.2 on Hugging Face under an MIT license within days of announcing it in June. This time the model went out through the company’s API and coding plan only. “We will release the weights two weeks after launch, once the security evaluation and hardening are complete,” Z.AI said.

What changed is the model’s performance in security work. GLM-5.3 runs on the same base model as GLM-5.2, with every gain coming from scaled-up post-training. On CyberGym, which tests whether a model can read code and confirm a real vulnerability, it scored 84.5 percent, ahead of Anthropic’s Mythos 5 at 83.8 and OpenAI’s GPT-5.6 Sol at 83.6. GLM-5.2 scored 77.2. The figures are Z.AI’s own and have not been independently verified.

The gap opens at the next stage. ExploitBench measures whether a model can turn a vulnerability into working code. GLM-5.3 scored 54.4 percent, against 78 for Mythos 5 and 76.5 for GPT-5.6 Sol, though that is more than double GLM-5.2’s 24.4. On ExploitGym, it completed 105 tasks within a two-hour budget and 130 within six hours, where Mythos 5 completed 181 and 247.

Z.AI said the capability kept compounding as training scaled, and that the model started reasoning across several stages of exploitation and assembling complete chains. The training had been aimed at finding vulnerabilities, not stringing them together.

The company also published a ledger of what its models have found. Since GLM-5.2 shipped, they have logged 2,436 vulnerabilities across 269 open-source projects, 107 rated critical and 990 high. Fifty-three are public with CVEs assigned, and 2,383 are not yet public. The oldest defect dates to 1981, and by Z.AI’s count the average bug had sat undiscovered for 26.6 years. Published entries include a use-after-free in the Linux kernel’s 6lowpan code, a memory-handling flaw in WebKit affecting Safari, and a parser bug in Suricata that let attachment names slip past inspection.

Coding scores moved as well, again on the company’s own numbers. Terminal-Bench 3.0 went from 4.6 to 28.3 and DeepSWE v1.1 from 46.2 to 66.9.

American labs have held releases back on capability grounds before. OpenAI said on August 7 that it could not rule out Critical cyber capability in Astra, its next model family, and slowed that work. Three days later it shipped a cyber model that refuses less, behind tiers requiring identity verification, monitoring and legal declarations. Chinese labs have generally gone the other way: Moonshot released Kimi K3’s full weights in July while Washington was weighing restrictions on Chinese models, and Z.AI has argued that open tools are what let smaller security teams keep up.

Beijing has been weighing limits of its own on how freely domestic labs can release frontier models abroad. Z.AI has been on the US Entity List since January 2025 and last month powered up a 1-gigawatt data center running entirely on Chinese chips.

The weights are due at the end of August.

Sources: Z.ai Security Disclosure Ledger, SCMP, heise online

–
By the Control Plane Editorial Team