News · High impactBack to News

Zhipu AI Releases GLM-5.3 with Scaled Post-Training RL and Emergent Cyber Capabilities, Open-Sourcing Weights in Two Weeks

Zhipu AI has officially released GLM-5.3, a flagship model enhanced through scaled post-training on a 743B base. Powered by IndexShare, SAO reinforcement learning, and the Slime async training framework, the model ranks first across CyberGym (84.5%), AutomationBench (48.2%), and GDPval-AA v2 (1769) while reaching 28.3 on Terminal Bench 3.0, accompanied by a two-week open-source roadmap.

Architecture Base and Post-Training Stack

Zhipu AI has officially released its next-generation flagship model, GLM-5.3. Retaining the 743B-parameter base from its predecessor, all performance gains stem from scaled post-training reinforcement learning.

Regarding its underlying infrastructure, GLM-5.3 leverages the IndexShare long-context engine, the SAO reinforcement learning algorithm designed for long-horizon tasks, and the open-source Slime asynchronous distributed training framework to collect signals across thousands of diverse interactive environments, driving a leap in long-sequence reasoning.

Coding and Cybersecurity Benchmark Performance

Official evaluation data demonstrates substantial progress on long-horizon terminal execution, software engineering repair, and vulnerability discovery. Notably in cybersecurity evaluations, the model exhibited emergent multi-step reconnaissance and exploit-chain reasoning capabilities.

Core Benchmark Performance

  • 28.3

    Terminal Bench 3.0 score, marking a major leap from 4.6 in the previous version.

  • 84.5%

    CyberGym real-world vulnerability benchmark accuracy, leading current evaluation tables.

  • 66.9

    DeepSWE v1.1 software engineering benchmark score, up from 46.2 previously.

Across benchmark comparisons with frontier models, GLM-5.3 exhibits comprehensive gains across 16 benchmarks spanning Coding, Cyber, and Agentic domains, establishing leading comprehensive engineering capability among open models:

Benchmark Category and MetricGLM-5.3GLM-5.2Kimi K3DeepSeek-V4 Pro-0813Qwen3.8-MaxOpus 4.8Fable 5 (w/ fallback)GPT-5.6 Sol
Terminal Bench 2.188.281.088.387.986.685.088.088.8
Terminal Bench 3.028.34.617.4--21.133.734.6
DeepSWE v1.166.946.267.562.756.658.069.772.7
NL2Repo58.048.958.061.155.969.7--
ProgramBench Almost Solved19.09.517.5-10.515.533.023.0
FrontierSWE78.167.5---66.588.2-
SWE-Marathon v1.142.519.448.1--48.833.142.5
PostTrainBench39.831.732.0--32.941.836.2
CyberGym84.577.280.083.378.578.183.883.6
ExploitGym 2h / 6h105 / 13029 / 3936 / 70-14 / 2680 / 120181 / 247216 / 293
ExploitBench54.424.432.2-28.840.078.076.5
Toolathlon Verified73.059.976.574.172.576.274.774.9
AutomationBench v1.0.648.226.246.743.239.841.046.245.8
Agents' Last Exam (ALE-CLI)28.523.827.625.727.025.723.828.6
HLE w/ Tools62.554.759.860.056.257.963.964.5
GDPval-AA v217691508168215901739158817431730

Across specific category evaluations, GLM-5.3 ranks first across major benchmarks including CyberGym vulnerability discovery (84.5%), AutomationBench workflow execution (48.2%), and the long-horizon benchmark GDPval-AA v2 (1769), surpassing contemporary proprietary and open-source models.

Mandatory Reasoning and Effort Levels

Regarding API interactions, GLM-5.3 alters its reasoning paradigm by deprecating direct output and enforcing mandatory thinking reasoning. Developers can configure computational budgets per task using explicit effort parameters.

The system provides three reasoning effort tiers: low for lightweight formatting and simple queries, high for balanced speed and token efficiency, and max as the default setting for intricate algorithmic design and multi-step debugging.

Availability and Open-Source Roadmap

GLM-5.3 is currently available through the official API and dedicated coding subscription plans. Addressing open-source community expectations, the team confirmed plans to open-source model weights within two weeks following safety reviews and alignment hardening.

Next step

Keep tracking GLM-5.3

Continue along the same topic.

Open entity record