GLM-5.3
Frontier open coding model with long-horizon reliability and three effort levels
About model
GLM-5.3 is Z.ai's frontier coding model, built on the same base as GLM-5.2 with every gain coming from scaled post-training. Training environments moved beyond coding exercises toward real units of expert work, some representing several days of an experienced engineer's effort, pushing the model to take ownership of substantial tasks end to end rather than relying on users to decompose and supervise each step. Thinking is always enabled, with reasoning effort adjustable across low, high, and max levels, and Z.ai reports higher accuracy at lower output-token budgets than GLM-5.2 at every effort level. Post-training also developed strong vulnerability-discovery capability, which Z.ai has directed into an ongoing coordinated disclosure program with a public ledger. Weights are released following the post-launch safety evaluation. Available on Together AI.
3
Low, high, and max, with max recommended for coding tasks
1M
Carried from the GLM-5.2 base for long-horizon engineering work
2,436
Identified across 269 open-source projects and tracked through Z.ai's public disclosure ledger
- Frontier Agentic Coding: Trained on environments built from real production workflows, taking ownership of multi-day engineering tasks end to end
- Long-Horizon Reliability: Post-training gains hold across long trajectories rather than only short tasks, carried by SAO with compaction
- Security Research: Identifies and validates real vulnerabilities in source code, with findings disclosed through Z.ai's public security ledger
- Production-Ready Infrastructure: 99.9% SLA, available on serverless and dedicated infrastructure
API usage
Endpoint:
Model card
Architecture Overview:
• Same base model as GLM-5.2, with all GLM-5.3 gains coming from post-training
• IndexShare architecture for efficient long-context processing across a 1M-token window
• Thinking always enabled, with reasoning effort configurable across low, high, and max levels (max default)
Training Methodology:
• Environment scaling toward real expert work: synthesized long-horizon environments with multi-step dependencies and hidden state, built by research agents from real task patterns and verified solvable by judge agents
• Verifiers synthesized without access to reference solutions, passing oracle, no-op, and unsolved-state checks before their binary rewards are used for training
• Carries over GLM-5.2's RL strategies, including SAO with compaction for long-horizon stability, with Z.ai reporting a 2.3x improvement in end-to-end RL training throughput from system optimizations
• Vulnerability-discovery data and environments included in the post-training mix
Performance Characteristics:
• Terminal coding (Terminal Bench 2.1 / 3.0, completing real tasks in a command-line environment): 88.2 / 28.3, with the 3.0 result up from 4.6 for GLM-5.2
• Repository-scale engineering (DeepSWE v1.1): 66.9, up from 46.2; repository generation (NL2Repo): 58.0; open-ended multi-hour projects (FrontierSWE): 78.1
• Ultra-long-horizon work (SWE-Marathon v1.1): 42.5; improving small models through post-training (PostTrainBench): 39.8
• General agent tasks (Toolathlon Verified multi-tool workflows: 73.0; AutomationBench: 48.2; Agents' Last Exam: 28.5, up from 23.8; Humanity's Last Exam with tools: 62.5)
• Knowledge work (GDPval-AA v2 Elo, evaluated by Artificial Analysis): 1769
• Security research (CyberGym, identifying and validating vulnerabilities from source code: 84.5, up from 77.2; ExploitBench: 54.4, up from 24.4)
• Z.ai reports a 50% improvement over GLM-5.2 on its in-house Z.ai Code Bench, reaching higher accuracy with fewer output tokens at every effort level
Prompting
Together AI API Access:
• Access GLM-5.3 via Together AI APIs using the endpoint zai-org/GLM-5.3
• Authenticate using your Together AI API key in request headers
• Thinking is always enabled; set reasoning effort to low, high, or max per request, with max as the default and recommended for coding tasks
• Compatible with Claude Code, OpenCode, and other major coding agent platforms
• Available on Together AI serverless and dedicated infrastructure
Applications & use cases
Long-Horizon Software Engineering:
• Hand off multi-day engineering tasks: diagnosing bottlenecks, implementing optimizations, and verifying end-to-end results
• Generate and modify entire repositories from natural language specifications
• Hold full codebases in the 1M-token window across extended sessions
Security & Vulnerability Research:
• Run authorized vulnerability discovery across codebases, identifying and validating real flaws from source
• Support security teams with triage and analysis grounded in the model's coordinated-disclosure track record
• Audit dependencies and legacy code where long-lived flaws hide
Knowledge Work & General Agents:
• Run multi-tool agent workflows with effort levels tuned to task difficulty
• Produce analysis and deliverables for real occupational tasks end to end
• Balance response speed against reasoning depth on a single deployment
- TypeChatReasoning
- IntelligenceHigh
- DeploymentServerlessDedicated
- Endpoint
- Context length1M
- Input price
$1.40 / 1M tokens
$0.26 (cached)/1M
- Output price
$4.40 / 1M tokens
- Input modalitiesText
- Output modalitiesText
- ReleasedAugust 13, 2026
- CategoryChat
