52 models · 19 providers · 43 with Terminal-Bench

Benchmark Leaders

Terminal-Bench 2.x

shell agent loops

Can the model finish real terminal / shell-agent tasks end to end — install tools, run commands, recover from errors.

0–100, higher is better. 70+ means a usable coding agent in a terminal; 80+ is frontier agent quality.

Editorial scores, not auto-scraped. Versions (2.0 vs 2.1) are not always comparable.

#1 GPT-5.6 Sol ★ SOTA (91.9 ultra)88.8
#2 Grok 4.6 AA88.4
#3 Kimi K3 ★ top open88.3
#4 GLM-5.3 (v) Claude Code · TB3 28.388.2
#5 Claude Mythos 5 (v) restricted88

SWE-bench Pro

software engineering

Can the model solve real software-engineering problems — patches, reviews, and repo-level fixes — in a single pass.

0–100, higher is better. A high score means stronger code-review and patch quality, not just snippet completion.

This is SWE-bench Pro, not SWE-bench Verified. Vendor harness notes are marked (v).

#1 Claude Fable 5 (v) classifier reroutes80
#2 Claude Opus 5 (v)79.2
#3 Claude Mythos 5 (v) restricted77.8
#4 SWE-1.7 (m)77.8
#5 DS V4 Pro (a)76.2

MCP-Atlas

multi-tool use

Can the model use many Model Context Protocol tools together — pick the right tool, pass arguments, and chain calls.

0–100, higher is better. High scores matter if your agent talks to GitHub, browsers, databases, or other MCP servers.

#1 Muse Spark 1.1 ★★ new leader88.1
#2 Claude Opus 5 (v)85.8
#3 Kimi K3 84.2
#4 Gemini 3.5 Flash 83.6
#5 Claude Fable 5 83.3
What each bench measures

Data Freshness

Updated 1 day ago.
anthropic.comqwen.aihuggingface.coopenrouter.aiartificialanalysis.aiz.aillm-stats.comopenai.comx.aimeta.comswfte.comai.nahcrof.com