GLM-5.2
Canonical API list price $0.76 in / $2.42 out per million tokens. Provider prices below can differ.
Agent scores
Terminal-Bench 2.x
shell agent loops
Can the model finish real terminal / shell-agent tasks end to end — install tools, run commands, recover from errors.
Editorial scores, not auto-scraped. Versions (2.0 vs 2.1) are not always comparable.
SWE-bench Pro
software engineering
Can the model solve real software-engineering problems — patches, reviews, and repo-level fixes — in a single pass.
This is SWE-bench Pro, not SWE-bench Verified. Vendor harness notes are marked (v).
MCP-Atlas
multi-tool use
Can the model use many Model Context Protocol tools together — pick the right tool, pass arguments, and chain calls.
SWE-Marathon & Long-Horizon
sustained multi-hour work
Whether the model stays useful over long agent sessions instead of collapsing after a few dozen steps.
Entries mix DeepSWE percentages, Marathon raw scores, and tool-call counts. Bars are editorial, not a shared scale.
Extended scores
BFCL
function calling
Berkeley Function Calling Leaderboard (V4): can the model call APIs with the right names, types, and arguments.
Coverage is thin on newest models. Sourced manually from the Berkeley leaderboard.
tau-bench
multi-turn agent tasks
Can the model hold a multi-turn tool-using conversation and finish a policy-constrained task.
AIME
math reasoning
American Invitational Mathematics Examination problems — contest math, not coding.
MMLU-Pro
knowledge breadth
Harder multiple-choice knowledge questions across academic subjects.
LiveCodeBench
live coding
Can the model solve fresh programming problems (contamination-resistant coding).
Where to buy
| Provider | Their price | Note | Link |
|---|---|---|---|
| CrofAI | $0.15/$0.52 | 1M ctx · 50% off list | Visit site |
| ZenMux | $1.35 | Best value on ZenMux | Sign upSupports this site |
| Together AI | $1.40 | Frontier MoE | — |
| Hyperbolic | $1.40 | Frontier MoE | — |
| OpenCode Go | $4.40 | Current main | Sign upSupports this site |
| Ollama CloudSub | sub | Frontier MoE · $20/mo Pro | — |
| OpenRouter | varies | Via proxy | — |
| Command CodeSub | sub | Frontier MoE | — |
| Kilo CodeSub | varies | Frontier MoE | — |
| FactorySub | sub | Droid Core | — |
| Z.AI DevPackSub | sub | Prior gen · MIT weights | — |
Editorial routing
| Task | Role | Note |
|---|---|---|
| Long-horizon multi-tool agent | Avoid | Tool-Decathlon ✗ |
| Single-pass code review | Alternative | 62.1 |
- Mid — $0.76/$2.42, TB 81, SWE-Pro 62.1
- GPT-5.6 Sol + GLM-5.2 — Peak TB + MIT executor
- GPT-5.6 Luna + GLM-5.2 — Cost-efficient frontier
- GLM-5.2 + Qwen3.8-27B — Self-hosted agents