Long-horizon multi-tool agent
Use is the highest Toolathlon score on a vendor card (80.6). GLM-5.3 73.0 and Qwen3.8 Max 72.5 sit Also. Opus Marathon 26.0 is a different bench. Skip GLM-5.2 (Tool-Decathlon fail).
Editorial picks per agent task: what to use, what else works, and what to skip. Scores and prices refresh from the model table. Use the wizard if you want a scored recommendation instead.
Roster 1 day ago. Routing reviewed Aug 16, 2026.
Use is the highest Toolathlon score on a vendor card (80.6). GLM-5.3 73.0 and Qwen3.8 Max 72.5 sit Also. Opus Marathon 26.0 is a different bench. Skip GLM-5.2 (Tool-Decathlon fail).
Use is the highest Terminal-Bench score not held out (88.8 SOTA). Luna is the cost companion at $6/M, TB 84.7. Grok 4.6 is AA TB 88.4 at $6/M. GLM-5.3 vendor TB 88.2 (Claude Code), Coding Plan only. Qwen3.8 Max vendor TB 86.6 at $2/$6
Use is the highest SWE-Pro score not held out (76.2). Fable / Composer / Opus 5 / Mythos / SWE-1.7 / Mistral stay held out. Opus is the tight-loop reasoning Also. Step 3.7 is the budget Also; Step 3.5 is cheapest coding-per-$
Context window is the gate
Qwen3.8 Max vendor QwenReactBench 1724. Hy3 won blind eval on frontend tasks
Agnes = new price floor, Step = free tier
Use is the highest CyberGym score on a vendor card (84.5 SOTA). Edges Mythos 83.8 and Sol 83.6. ExploitBench 54.4 still trails the closed frontier (~76–78)
Sonnet 5 = new default plan model. Fable 5 ceiling w/ classifier risk. GLM-5.3 vendor GDPval-AA v2 1769 leads its launch table
Sol DeepSWE 72.7% = new #1. GLM-5.3 vendor DeepSWE 66.9%. Qwen3.8 Max vendor DeepSWE 56.6% / PaperBench 93.0. Qwen3.7 has public 1000+ call evidence