GLM-5.3
No public per-token API list price yet. Provider prices below can differ.
Agent scores
Terminal-Bench 2.x
shell agent loops
Can the model finish real terminal / shell-agent tasks end to end — install tools, run commands, recover from errors.
Editorial scores, not auto-scraped. Versions (2.0 vs 2.1) are not always comparable.
SWE-bench Pro
software engineering
Can the model solve real software-engineering problems — patches, reviews, and repo-level fixes — in a single pass.
This is SWE-bench Pro, not SWE-bench Verified. Vendor harness notes are marked (v).
MCP-Atlas
multi-tool use
Can the model use many Model Context Protocol tools together — pick the right tool, pass arguments, and chain calls.
SWE-Marathon & Long-Horizon
sustained multi-hour work
Whether the model stays useful over long agent sessions instead of collapsing after a few dozen steps.
Entries mix DeepSWE percentages, Marathon raw scores, and tool-call counts. Bars are editorial, not a shared scale.
Vendor launch scores
Z.ai launch · 2026-08-14. Vendor harness — not independent evals.
| Benchmark | Score | Note |
|---|---|---|
| Coding | ||
| Terminal-Bench 3.0 | 28.3 | (v) Claude Code |
| DeepSWE v1.1 | 66.9 | (v) |
| NL2Repo | 58.0 | (v) |
| ProgramBench Almost Solved | 19.0 | (v) |
| FrontierSWE | 78.1 | (v) Proximal |
| SWE-Marathon v1.1 | 42.5 | (v) |
| PostTrainBench | 39.8 | (v) |
| Cyber | ||
| CyberGym | 84.5 | (v) SOTA |
| ExploitGym 2h / 6h | 105 / 130 | (v) |
| ExploitBench | 54.4 | (v) |
| Agentic | ||
| Toolathlon Verified | 73.0 | (v) |
| AutomationBench v1.0.6 | 48.2 | (v) |
| Agents' Last Exam ALE-CLI | 28.5 | (v) |
| HLE w/ Tools | 62.5 | (v) |
| GDPval-AA v2 | 1769 | (v) Elo |
Source: z.ai/blog/glm-5.3. (v) = vendor harness.
Extended scores
BFCL
function calling
Berkeley Function Calling Leaderboard (V4): can the model call APIs with the right names, types, and arguments.
Coverage is thin on newest models. Sourced manually from the Berkeley leaderboard.
tau-bench
multi-turn agent tasks
Can the model hold a multi-turn tool-using conversation and finish a policy-constrained task.
AIME
math reasoning
American Invitational Mathematics Examination problems — contest math, not coding.
MMLU-Pro
knowledge breadth
Harder multiple-choice knowledge questions across academic subjects.
LiveCodeBench
live coding
Can the model solve fresh programming problems (contamination-resistant coding).
Where to buy
| Provider | Their price | Note | Link |
|---|---|---|---|
| Z.AI DevPackSub | sub | New flagship · Coding Plan | — |
Editorial routing
| Task | Role | Note |
|---|---|---|
| Long-horizon multi-tool agent | Alternative | Toolathlon 73.0 (v) |
| Terminal-agent loops | Alternative | 88.2 (v) |
| Security review | Primary | CyberGym 84.5 SOTA |
| Planning / spec mode | Alternative | GDPval-AA v2 1769 (v) |
| Long-horizon autonomous | Alternative | DeepSWE 66.9% |