Just in! Z.AI releases GLM-5.3 model, "built for coding," with performance rivaling Fable 5

Wallstreetcn
2026.08.14 06:06

Z.AI has released its 700-billion-parameter flagship model, GLM-5.3, achieving significant breakthroughs in programming and cybersecurity through post-training scaling. Its programming capabilities on the Z.AI Code Bench improved by 50%, surpassing Claude Opus 4.8 on some high-difficulty tasks. The model demonstrated "emergent" capabilities in cybersecurity vulnerability mining, having identified a cumulative total of 2,436 vulnerabilities in real-world codebases. Model weights will be open-sourced within two weeks, further strengthening its market competitiveness as DeepSeek raises prices

Z.AI's latest flagship model, GLM-5.3, has officially debuted, marking major breakthroughs in programming capabilities and cybersecurity vulnerability discovery, launching a new round of challenges against AI leaders such as Anthropic and OpenAI.

According to Z.AI's official statement, GLM-5.3 is built on the same 700-billion-parameter base model as its predecessor GLM-5.2, with all performance improvements stemming from continuous scaling during the post-training phase. On its proprietary evaluation benchmark, the Z.AI Code Bench, GLM-5.3's programming capabilities improved by 50% compared to GLM-5.2; on the open-source programming benchmark Terminal-Bench 3.0, its score jumped from 4.6 to 28.3.

When introducing the model, Z.AI stated: GLM-5.3 is built for coding, ready to meet cyber defense challenges at any time. GLM-5.3 also exhibited an unexpected "emergent" leap in cybersecurity vulnerability mining capabilities, cumulatively identifying 2,436 security vulnerabilities in real-world codebases.

Z.AI announced that the model weights will be made public two weeks after release, by which time security assessments and hardening will be completed. This launch coincides with DeepSeek's price hike for its V4 models yesterday, further strengthening Z.AI's competitive position in the domestic AI market.

Significant Leap in Programming Capabilities, Surpassing Claude Opus 4.8

The core breakthroughs of GLM-5.3 focus on complex programming and long-chain tasks. According to official data from Z.AI, under the highest difficulty settings of the Z.AI Code Bench, GLM-5.3 completed 34.5% of tasks using approximately 75,000 output tokens, whereas GLM-5.2 achieved only a 23.4% completion rate while consuming more tokens (96,000). Under high-difficulty settings, GLM-5.3 achieved a 31.4% completion rate with approximately 50,000 tokens, surpassing Anthropic's Claude Opus 4.8 (29.5% completion rate, consuming 120,000 tokens). However, GLM-5.3 currently still lags behind Claude Fable 5, which achieved a 39.5% completion rate under the highest difficulty settings.

On public benchmarks, GLM-5.3 also performed impressively: its DeepSWE v1.1 score rose from 46.2 to 66.9, and its Agents' Last Exam score increased from 23.8 to 28.5.

To achieve these advancements, Z.AI significantly expanded the scale and diversity of training environments during the post-training phase, introducing tasks closer to real-world engineering scenarios—some of which equate to several days of work for an experienced engineer. Z.AI also constructed an automated pipeline for end-to-end synthesis of training environments and reinforcement learning reward signals.

"Emergent" Cybersecurity Capabilities, Over Two Thousand Vulnerabilities Disclosed

In the field of cybersecurity, GLM-5.3 demonstrated capability leaps exceeding expectations. After Z.AI introduced vulnerability discovery data and training environments, the model not only began to identify single defects but also started performing comprehensive reasoning across multiple exploitation stages, forming complete exploitation chain plans.

In the CyberGym benchmark, GLM-5.3 scored 84.5%, surpassing Anthropic's Mythos 5 (83.8%) and OpenAI's GPT-5.6 Sol (83.6%) to take the top spot. On ExploitBench, which requires deeper vulnerability reasoning, GLM-5.3's score jumped from 24.4% for GLM-5.2 to 54.4%, more than doubling; however, there remains a significant gap compared to Mythos 5 (78.0%) and GPT-5.6 Sol (76.5%). In the ExploitGym test, GLM-5.3 completed 105 exploitation tasks within two hours and 130 tasks within six hours, whereas the corresponding figures for GLM-5.2 were only 29 and 39 tasks; Mythos 5 maintained its lead with 181 and 247 tasks, respectively.

Z.AI stated that these capabilities have extended to real-world applications. Since GLM-5.2, Z.AI has collaborated with multiple security teams in China to run model tests on real-world codebases. After expert review, screening, and deduplication, 2,436 vulnerabilities were identified across 269 projects, including 1,097 medium-to-high severity vulnerabilities. These cover areas such as system kernels, operating systems, browser engines, open-source infrastructure, web applications, and network protocols. Some vulnerabilities had existed in codebases for decades, with the earliest dating back approximately 40 years. Z.AI has established the "Z.AI Security Disclosure Ledger" to continuously publicly disclose progress.

Open-Source Strategy Amid Shifting Competitive Landscape Opens Up Market Space

From a commercial perspective, the timing of the GLM-5.3 release is strategically significant. DeepSeek announced on Thursday that it would raise prices by up to four times for its V4 Flash and Pro models. Although the comprehensive pricing of V4 Pro remains lower than that of GLM-5.2, this move objectively provides a window for Z.AI to attract users and developers. Artificial Analysis gave DeepSeek's top-tier model and GLM-5.2 the same intelligence score of 53, with both lagging behind Kimi K3, Claude Fable 5, GPT-5.6, and other models.

Z.AI plans to release the GLM-5.3 model weights under a permissive license to attract a broader developer community. This strategy follows the same path as GLM-5.2—after its release, Z.AI's valuation once surged to $137 billion, surpassing internet giants such as Pinduoduo and NetEase. Although it has since retreated to around $80 billion, this represents more than a tenfold increase compared to its Hong Kong listing in January this year.

As previously reported by Bloomberg, Z.AI has completed the construction of a data center, deploying at least 10,000 domestically produced chips for the research, development, and inference of the GLM series models. Its annual recurring revenue (ARR) reached $1 billion in July this year.