Z.ai has released GLM-5.3, a model that runs on the same 743B base as GLM-5.2 but gains performance through scaled post-training. The company reports improvements in complex coding and long-horizon tasks without retraining the underlying weights. Terminal-Bench 3.0 scores moved from 4.6 to 28.3. Cybersecurity benchmarks also saw significant gains, with CyberGym reaching 84.5%. Weights are not public yet. Z.ai says they will publish them roughly two weeks after launch, once safety evaluation and hardening finish.
In this article
Is It Deployable?
GLM-5.3 is live through the Z.ai API, the GLM Coding Plan, and ZCode. Weights are not out. Startups and mid-market engineering orgs can adopt it today via the Coding Plan or API. Enterprises with data-residency or vendor-review rules should wait for weights. Security vendors and MSSPs get the most signal, and the most policy exposure. Developer tooling, cloud infrastructure, application security, fintech and e-commerce engineering, and vendors shipping kernels, browser engines, or network stacks are the relevant industries. Applications include repository-scale refactors, long-horizon CLI agents, CI failure triage, white-box vulnerability discovery, crash triage, and secure code review.
Coding Results
Terminal-Bench 3.0 moves from 4.6 to 28.3 against GLM-5.2. DeepSWE v1.1 moves from 46.2 to 66.9. Agents’ Last Exam (CLI) moves from 23.8 to 28.5. On GDPval-AA v2, which spans 44 occupations, GLM-5.3 scores 1,769.
On Z.ai Code Bench, an internal evaluation, the company reports a 50% improvement over GLM-5.2. It reports 31.4% at roughly 50,000 output tokens per task. Claude Opus 4.8 scores 29.5% at 120,000 tokens. Claude Fable 5 still leads at 39.5% at maximum effort. Z.ai argues a private benchmark reduces contamination risk.
On public suites, GLM-5.3 trails GPT-5.6 Sol and Fable 5 on several harder coding evaluations. All figures are vendor-reported, with harness, context length, and sampling settings documented in the announcement.
The Cybersecurity Result
Z.ai flags this one as unplanned. It added vulnerability-discovery data expecting better single-bug reasoning. Instead, capability kept compounding as training scaled. The model began forming coherent plans across complete exploitation chains.
CyberGym, which tests discovery and validation from white-box source, moves from 77.2% to 84.5%. That edges past Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%. ExploitBench, which requires root-cause reasoning and a working exploit, moves from 24.4% to 54.4%. Mythos 5 sits at 78.0%. On ExploitGym, GLM-5.3 completes 105 tasks in two hours and 130 in six. GLM-5.2 completes 29 and 39. Mythos 5 completes 181 and 247.
The pattern is consistent. The deeper into the exploitation chain a benchmark sits, the larger the gain over GLM-5.2. The gap to closed frontier models also widens.
What it means
Developers using ZCode or the API can immediately test new capabilities for long-running scripts and complex code refactors. Security teams gain a tool capable of generating full exploitation chains rather than just identifying single vulnerabilities. The lack of open weights means large organisations must wait before integrating the model into internal pipelines.




