Terminal-Bench 2.xshell agent loops
SWE-bench Prosoftware engineering
MCP-Atlasmulti-tool use
SWE-Marathon & Long-Horizonsustained multi-hour work
BFCLfunction calling
tau-benchmulti-turn agent tasks
AIMEmath reasoning
MMLU-Proknowledge breadth
LiveCodeBenchlive coding

Agent-tier rankings

Terminal-Bench 2.x

shell agent loops

Can the model finish real terminal / shell-agent tasks end to end — install tools, run commands, recover from errors.

0–100, higher is better. 70+ means a usable coding agent in a terminal; 80+ is frontier agent quality.

Editorial scores, not auto-scraped. Versions (2.0 vs 2.1) are not always comparable.

#1 GPT-5.6 Sol ★ SOTA (91.9 ultra)88.8
#2 Grok 4.6 AA88.4
#3 Kimi K3 ★ top open88.3
#4 GLM-5.3 (v) Claude Code · TB3 28.388.2
#5 Claude Mythos 5 (v) restricted88
#6 Terra ★ half Sol cost87.4
#7 Qwen3.8 Max ★ beats Opus (v)86.6
#8 Luna ★ best $/TB84.7
#9 Claude Fable 5 ★ restored84.3
#10 Grok 4.5 ★ cost-efficient83.3
#11 DS V4 Flash official 073182.7
#12 SWE-1.7 (v)81.5
#13 GLM-5.2 81
#14 Sonnet 5 80.4
#15 Muse Spark 1.1 Meta API80
#16 Claude Opus 4.8 ~ (v: 85.0)79
#17 Gemini 3.6 Flash 78
#18 Gemini 3.5 Flash 76.2
#19 Qwen3.8-27B (v) Terminus73
#20 Hy3 2.171.7
#21 Seed 2.1 Pro (v)71
#22 DS V4 Pro 67.9
#23 Seed 2.1 Turbo (v)67.6
#24 MiniMax M3 (v)66
#25 Qwen3.7-Plus vision Plus61
#26 Step 3.7F 59.5
#27 Nemotron 3U 56.4
#28 Gemini 3.1P 54.2
#29 Gemini 3.5 Flash-Lite 54
#30 Step 3.5F (v)51
#31 Mistral L3 ✗ worst agent12
Not scored5

Claude Opus 5 (no vendor TB 2.1; AA ~89% at max effort, not ranked)

SWE-bench Pro

software engineering

Can the model solve real software-engineering problems — patches, reviews, and repo-level fixes — in a single pass.

0–100, higher is better. A high score means stronger code-review and patch quality, not just snippet completion.

This is SWE-bench Pro, not SWE-bench Verified. Vendor harness notes are marked (v).

#1 Claude Fable 5 (v) classifier reroutes80
#2 Claude Opus 5 (v)79.2
#3 Claude Mythos 5 (v) restricted77.8
#4 SWE-1.7 (m)77.8
#5 DS V4 Pro (a)76.2
#6 Fugu Ultra (v)73.7
#7 Claude Opus 4.8 (v)69.2
#8 Qwen3.8 Max (v)67.7
#9 Grok 4.5 (v) cost-efficient64.7
#10 GPT-5.6 Sol (v)64.6
#11 Terra (v) half Sol cost63.4
#12 Sonnet 5 63.2
#13 Luna (v) best $/SWE62.7
#14 GLM-5.2 (v)62.1
#15 DS V4.5 ★ open value62.1
#16 Qwen3.8-27B (v)61.7
#17 Muse Spark 1.1 Meta API61.5
#18 Qwen3.7 Max 60.6
#19 MiniMax M3 ~ (v)59
#20 Gemini 3.6 Flash (v)58.7
#21 Hy3 57.9
#22 Step 3.7F 56.3
#23 Gemini 3.5 Flash (v)55.1
#24 Gemini 3.5 Flash-Lite 54.2
#25 DS V4 Flash 52.6
#26 Gemini 3.1P ✗ (s)46.1
#27 Claude Haiku 4.5 (s)39.5
Not scored

Kimi K2.7 Code (has SWE-V 62.0 not Pro)

Kimi K3 (has SWE-V 76.8 & FrontierSWE 81.2, no Pro yet)

GLM-5.3 (no SWE-Pro at launch; DeepSWE v1.1 66.9, SWE-Marathon v1.1 42.5)

Llama 4 Maverick (has SWE-V 70.4 not Pro)

Mistral Large 3 (has SWE-V 76.2 not Pro)

Nemotron 3 Ultra (has SWE-V 71.9 not Pro)

Step 3.5 Flash (has SWE-V 74.4 not Pro)

MiMo v2.5 Pro, Ring 2.6 1T, Qwen3.7-Plus, Command A+, Cohere North Mini Code — no SWE-Pro or SWE-V data

Agnes 2.0 Flash, ERNIE 5.1, Gemma 4 12B, DiffusionGemma 26B-A4B, Seed 2.1 Pro, Seed 2.1 Turbo — no SWE data available

DeepSeek V4 Flash (SWE-Pro 52.6, SWE-V 79.0 — no newer official SWE-Pro)

Gemini 3.5 Flash Cyber — restricted pilot, no public scores

MCP-Atlas

multi-tool use

Can the model use many Model Context Protocol tools together — pick the right tool, pass arguments, and chain calls.

0–100, higher is better. High scores matter if your agent talks to GitHub, browsers, databases, or other MCP servers.

#1 Muse Spark 1.1 ★★ new leader88.1
#2 Claude Opus 5 (v)85.8
#3 Kimi K3 84.2
#4 Gemini 3.5 Flash 83.6
#5 Claude Fable 5 83.3
#6 Hy3 79.1
#7 Claude Opus 4.8 77.8
#8 GLM-5.2 76.8
#9 Qwen3.7 Max 76.4
#10 Kimi K2.7 76
#11 DS V4 Pro 73.6
#12 Qwen3.7-Plus vision Plus73.2
#13 Mistral L3 70.4
#14 DS V4 Flash 69

SWE-Marathon & Long-Horizon

sustained multi-hour work

Whether the model stays useful over long agent sessions instead of collapsing after a few dozen steps.

Mixed — not one number. Use this as supporting evidence for long-running agents, not as a ranking against TB or SWE-Pro.

Entries mix DeepSWE percentages, Marathon raw scores, and tool-call counts. Bars are editorial, not a shared scale.

#1 GPT-5.6 Sol DeepSWE #1 ★72.7%
#2 Claude Fable 5 DeepSWE ★70%
#3 GPT-5.6 Terra DeepSWE69.6%
#4 Claude Opus 5 DeepSWE (v)68.8%
#5 GPT-5.6 Luna DeepSWE67.2%
#6 GLM-5.3 DeepSWE (v)66.9%
#7 Grok 4.6 DeepSWE (v)65.9%
#8 Qwen3.8 Max DeepSWE (v)56.6%
#9 Claude Opus 4.8 Marathon ★26.0
#10 Qwen3.7 Max 35h proven ★1,158 calls
#11 GLM-5.2 Marathon13.0
#12 DeepSeek V4 Pro DeepSWE ✗8%

Entries mix DeepSWE percentages, Marathon raw scores, and tool-call counts. Bars are editorial, not a shared scale.

Not on this chart

Grok 4.5, Muse Spark 1.1, DeepSeek V4.5, Gemini 3.5 Flash, Step 3.7, Hy3, Sonnet 5, MiniMax M3, SWE-1.7 — no Marathon scores yet. Kimi K3 DeepSWE v1.1 ~67.5% not charted (vendor board).

Extended benches

BFCL

function calling

Berkeley Function Calling Leaderboard (V4): can the model call APIs with the right names, types, and arguments.

0–100, higher is better. Useful when the job is structured tool calls rather than a full agent loop.

Coverage is thin on newest models. Sourced manually from the Berkeley leaderboard.

#1 Claude Opus 4.8 75.5
#2 GLM-5.2 74.1

tau-bench

multi-turn agent tasks

Can the model hold a multi-turn tool-using conversation and finish a policy-constrained task.

0–100, higher is better. High scores help for customer-support-style agents that must follow rules across turns.

#1 Gemini 3.1P 99.3
#2 Llama 4M 93.7
#3 GLM-5.2 89.7
#4 Claude Opus 4.8 85.2
#5 Gemma 4 12B (v)69
#6 Nemotron 3S 61.15
#7 DiffusionGemma 26B 56.2

AIME

math reasoning

American Invitational Mathematics Examination problems — contest math, not coding.

0–100, higher is better. A proxy for careful multi-step reasoning. Weak signal for terminal or repo agents.

#1 GLM-5.2 99.2
#2 Nemotron 3S 90.21
#3 Claude Fable 5 (v)87.3
#4 Claude Opus 4.8 82.1
#5 Gemini 3.5 Flash 78.4
#6 Gemma 4 12B 77.5
#7 DiffusionGemma 26B 69.1

MMLU-Pro

knowledge breadth

Harder multiple-choice knowledge questions across academic subjects.

0–100, higher is better. Shows general knowledge. Does not measure tool use or coding agents.

#1 Qwen3.7 Max 89.6
#2 DS V4.5 89.4
#3 Qwen3.7-Plus 88.5
#4 DS V4 Pro Max 87.5
#5 DS V4 Pro 87.5
#6 Nemotron 3U 86.8
#7 Claude Opus 4.8 86.5
#8 DS V4 Flash 86.2
#9 Claude Fable 5 (v)85.8
#10 Nemotron 3S 83.7
#11 Gemini 3.5 Flash 82.3
#12 Llama 4M 80.5
#13 GLM-5.2 80.1
#14 DiffusionGemma 26B 77.6
#15 Gemma 4 12B 77.2
#16 MiMo v2.5 Pro 68.5

LiveCodeBench

live coding

Can the model solve fresh programming problems (contamination-resistant coding).

0–100, higher is better. Complements SWE-Pro: LiveCode is contest-style code, SWE-Pro is repo engineering.

#1 DS V4 Pro Max 93.5
#2 DS V4 Pro 93.5
#3 DS V4.5 93.1
#4 DS V4 Flash 91.6
#5 Qwen3.8-27B (v) LCB v690.3
#6 Nemotron 3U 89
#7 Nemotron 3S 81.2
#8 Claude Fable 5 (v)79.5
#9 Grok 4.5 79
#10 Claude Opus 4.8 74.2
#11 Gemma 4 12B (v)72
#12 DiffusionGemma 26B (v)69.1
#13 Gemini 3.5 Flash 68.1
#14 GLM-5.2 65.3
#15 Llama 4M 43.4
#16 Mistral L3 34.4

Cost-Performance Map

Up and left is better · Terminal-Bench vs $/M output (log scale)

  • Hy3TB 71.7 · $0.53/M
  • DS Flash 0731TB 82.7 · $0.28/M
  • Step 3.5FTB 51 · $0.3/M
  • Agnes 2.0TB 0 · $0.2/M
  • DS V4 ProTB 67.9 · $0.87/M
  • NemotronTB 56.4 · $2.2/M
  • Step 3.7FTB 59.5 · $1.15/M
  • MiniMax M3TB 66 · $1.2/M
  • Mistral L3TB 12 · $1.5/M
  • Muse SparkTB 80 · $4.25/M
  • GLM-5.2TB 81 · $2.42/M
  • Grok 4.5TB 83.3 · $6/M
  • Grok 4.6TB 88.4 · $6/M
  • Qwen3.8 MaxTB 86.6 · $6/M
  • LunaTB 84.7 · $6/M
  • Haiku 4.5TB 44.2 · $5/M
  • Gemini 3.5FTB 76.2 · $9/M
  • Sonnet 5 ★TB 80.4 · $10/M
  • Gemini 3.1PTB 54.2 · $12/M
  • Opus 4.8TB 79 · $25/M
  • TerraTB 87.4 · $15/M
  • GPT-5.6 Sol ★TB 88.8 · $30/M
  • Fable 5 ★TB 84.3 · $50/M