Long-horizon multi-tool agent

Use is the highest Toolathlon score on a vendor card (80.6). GLM-5.3 73.0 and Qwen3.8 Max 72.5 sit Also. Opus Marathon 26.0 is a different bench. Skip GLM-5.2 (Tool-Decathlon fail).

Use

Claude Opus 5Toolathlon 80.6 (v)

Also

GLM-5.3Toolathlon 73.0 (v)
Qwen3.8 MaxToolathlon 72.5 (v)
Qwen3.7 Max35h proven
Opus 4.8Marathon 26.0 — different bench

Skip

GLM-5.2Tool-Decathlon ✗
DS V4 ProDeepSWE 8%

Terminal-agent loops

Use is the highest Terminal-Bench score not held out (88.8 SOTA). Luna is the cost companion at $6/M, TB 84.7. Grok 4.6 is AA TB 88.4 at $6/M. GLM-5.3 vendor TB 88.2 (Claude Code), Coding Plan only. Qwen3.8 Max vendor TB 86.6 at $2/$6

Use

GPT-5.6 Sol★ 88.8 SOTA
GPT-5.6 Luna84.7 · $6/M

Also

Grok 4.688.4 · $6/M
GLM-5.388.2 (v)
Qwen3.8 Max86.6 (v) · $6/M
Held out of Terminal-agent loops (3)
  • Kimi K388.3 — $15/M out — top open on the TB chart, not a default agent route
  • Claude Mythos 588 — Glasswing-only / restricted
  • Terra87.4 — Sol sibling; Luna is the cost primary on this row

Single-pass code review

Use is the highest SWE-Pro score not held out (76.2). Fable / Composer / Opus 5 / Mythos / SWE-1.7 / Mistral stay held out. Opus is the tight-loop reasoning Also. Step 3.7 is the budget Also; Step 3.5 is cheapest coding-per-$

Use

DS V4 ProSWE-Pro 76.2

Also

Opus 4.8SWE-Pro 69.2 · tight-loop reasoning
Grok 4.564.7 · $6/M
Step 3.756.3 ($1.15/M)
Step 3.574.4 V ($0.30/M)
Held out of Single-pass code review (6)
  • Claude Fable 580 — Highest SWE-Pro but classifier gap; on planning if available
  • Composer 2.579.8 — Cursor pair, not a review default
  • Claude Opus 579.2 — SWE-Pro near Fable; TB 2.1 deliberately unranked
  • Claude Mythos 577.8 — Glasswing-only / restricted
  • SWE-1.777.8 — Devin preview; multilingual SWE score on the model row
  • Mistral L376.2 — SWE-Pro is fine; Terminal-Bench 12 makes it a skip for agent work

Monorepo analysis

Context window is the gate

Also

Any 1M-context model

Skip

<1M ctx

Frontend / web dev

Qwen3.8 Max vendor QwenReactBench 1724. Hy3 won blind eval on frontend tasks

Use

Also

Qwen3.8 MaxReactBench 1724 (v)
Hy3blind eval win

High-volume batch / cheap

Agnes = new price floor, Step = free tier

Use

DS V4 FlashTB 2.1 82.7 · $0.28/M

Also

GPT-5.6 Luna$1/$6, TB 84.7
Agnes 2.0$0.20/M
Step 3.5$0.30/M
Step 3.7$1.15/M

Skip

Opus/GPTcost

Security review

Use is the highest CyberGym score on a vendor card (84.5 SOTA). Edges Mythos 83.8 and Sol 83.6. ExploitBench 54.4 still trails the closed frontier (~76–78)

Use

GLM-5.3CyberGym 84.5 SOTA

Also

Mythos 5CyberGym 83.8 · restricted
GPT-5.6 SolCyberGym 83.6

Planning / spec mode

Sonnet 5 = new default plan model. Fable 5 ceiling w/ classifier risk. GLM-5.3 vendor GDPval-AA v2 1769 leads its launch table

Use

Sonnet 5GDPval 1618
Fable 5if available

Also

GLM-5.3GDPval-AA v2 1769 (v)
Opus 4.8fallback

Long-horizon autonomous

Sol DeepSWE 72.7% = new #1. GLM-5.3 vendor DeepSWE 66.9%. Qwen3.8 Max vendor DeepSWE 56.6% / PaperBench 93.0. Qwen3.7 has public 1000+ call evidence

Use

GPT-5.6 SolDeepSWE 72.7% ★

Also

GLM-5.3DeepSWE 66.9%
Opus 4.8Marathon 26.0
Qwen3.8 MaxDeepSWE 56.6% (v)
Qwen3.7 Max35h / 1,158 calls

Skip

DS V4 Pro8% DeepSWE