Benchmarks
Each score answers a different question. Agent-tier benches measure tool-using work; extended benches add function calling, math, knowledge, and contest coding. Every chart card states what the bench measures.
Agent-tier rankings
Bar color is the score band. The dashed mark at 80 is frontier-quality. Click a bar to open the model.
Terminal-Bench 2.x
shell agent loops
Can the model finish real terminal / shell-agent tasks end to end — install tools, run commands, recover from errors.
Editorial scores, not auto-scraped. Versions (2.0 vs 2.1) are not always comparable.
Claude Opus 5 (no vendor TB 2.1; AA ~89% at max effort, not ranked)
SWE-bench Pro
software engineering
Can the model solve real software-engineering problems — patches, reviews, and repo-level fixes — in a single pass.
This is SWE-bench Pro, not SWE-bench Verified. Vendor harness notes are marked (v).
Kimi K2.7 Code (has SWE-V 62.0 not Pro)
Kimi K3 (has SWE-V 76.8 & FrontierSWE 81.2, no Pro yet)
GLM-5.3 (no SWE-Pro at launch; DeepSWE v1.1 66.9, SWE-Marathon v1.1 42.5)
Llama 4 Maverick (has SWE-V 70.4 not Pro)
Mistral Large 3 (has SWE-V 76.2 not Pro)
Nemotron 3 Ultra (has SWE-V 71.9 not Pro)
Step 3.5 Flash (has SWE-V 74.4 not Pro)
MiMo v2.5 Pro, Ring 2.6 1T, Qwen3.7-Plus, Command A+, Cohere North Mini Code — no SWE-Pro or SWE-V data
Agnes 2.0 Flash, ERNIE 5.1, Gemma 4 12B, DiffusionGemma 26B-A4B, Seed 2.1 Pro, Seed 2.1 Turbo — no SWE data available
DeepSeek V4 Flash (SWE-Pro 52.6, SWE-V 79.0 — no newer official SWE-Pro)
Gemini 3.5 Flash Cyber — restricted pilot, no public scores
MCP-Atlas
multi-tool use
Can the model use many Model Context Protocol tools together — pick the right tool, pass arguments, and chain calls.
Not scored30
- Claude Sonnet 5
- GPT-5.6 Sol
- GPT-5.6 Terra
- GPT-5.6 Luna
- Claude Mythos 5
- Grok 4.5
- Grok 4.6
- Grok Build 0.1
- GLM-5.3
- DeepSeek V4 Pro Max
- DeepSeek V4.5
- Nemotron 3 Ultra
- Nemotron 3 Super
- Step 3.7 Flash
- Step 3.5 Flash
- Llama 4 Maverick
- Qwen3-Coder-Next
- Ring 2.6 1T
- MiMo v2.5 Pro
- ERNIE 5.1
- Agnes 2.0 Flash
- Seed 2.1 Turbo
- Seed 2.1 Pro
- DiffusionGemma 26B-A4B
- Gemma 4 12B
- Composer 2.5
- SWE-1.7
- Gemini 3.6 Flash
- Gemini 3.5 Flash-Lite
- Gemini 3.5 Flash Cyber
Scale MCP-Atlas leaderboard last updated April 2026 — does not include models released after that date
SWE-Marathon & Long-Horizon
sustained multi-hour work
Whether the model stays useful over long agent sessions instead of collapsing after a few dozen steps.
Entries mix DeepSWE percentages, Marathon raw scores, and tool-call counts. Bars are editorial, not a shared scale.
Entries mix DeepSWE percentages, Marathon raw scores, and tool-call counts. Bars are editorial, not a shared scale.
Not on this chart
Grok 4.5, Muse Spark 1.1, DeepSeek V4.5, Gemini 3.5 Flash, Step 3.7, Hy3, Sonnet 5, MiniMax M3, SWE-1.7 — no Marathon scores yet. Kimi K3 DeepSWE v1.1 ~67.5% not charted (vendor board).
Extended benches
BFCL
function calling
Berkeley Function Calling Leaderboard (V4): can the model call APIs with the right names, types, and arguments.
Coverage is thin on newest models. Sourced manually from the Berkeley leaderboard.
Not scored34
- Claude Sonnet 5
- Claude Fable 5
- GPT-5.6 Sol
- Claude Mythos 5
- Gemini 3.5 Flash
- Grok 4.5
- Grok 4.6
- Grok Build 0.1
- GLM-5.3
- Kimi K2.7 Code
- Qwen3.7 Max
- DeepSeek V4 Pro Max
- DeepSeek V4 Pro
- DeepSeek V4 Flash
- DeepSeek V4.5
- MiniMax M3
- Nemotron 3 Ultra
- Step 3.7 Flash
- Step 3.5 Flash
- Qwen3-Coder-Next
- Mistral Large 3
- Ring 2.6 1T
- MiMo v2.5 Pro
- ERNIE 5.1
- Agnes 2.0 Flash
- Fugu Ultra
- Cohere North Mini Code
- SWE-1.7
- Muse Spark 1.1
- Claude Opus 5
- Gemini 3.6 Flash
- Kimi K3
- Gemini 3.5 Flash-Lite
- Gemini 3.5 Flash Cyber
BFCL V4 leaderboard last updated April 2026 — does not include frontier models released after that date
tau-bench
multi-turn agent tasks
Can the model hold a multi-turn tool-using conversation and finish a policy-constrained task.
Not scored35
- Claude Sonnet 5
- Claude Fable 5
- Claude Haiku 4.5
- GPT-5.6 Sol
- Claude Mythos 5
- Gemini 3.5 Flash
- Grok 4.5
- Grok 4.6
- Grok Build 0.1
- GLM-5.3
- Kimi K2.7 Code
- Qwen3.7 Max
- DeepSeek V4 Pro Max
- DeepSeek V4 Pro
- DeepSeek V4 Flash
- DeepSeek V4.5
- MiniMax M3
- Nemotron 3 Ultra
- Step 3.7 Flash
- Step 3.5 Flash
- Qwen3-Coder-Next
- Mistral Large 3
- Ring 2.6 1T
- MiMo v2.5 Pro
- ERNIE 5.1
- Agnes 2.0 Flash
- Fugu Ultra
- Cohere North Mini Code
- SWE-1.7
- Muse Spark 1.1
- Claude Opus 5
- Gemini 3.6 Flash
- Kimi K3
- Gemini 3.5 Flash-Lite
- Gemini 3.5 Flash Cyber
AIME
math reasoning
American Invitational Mathematics Examination problems — contest math, not coding.
Not scored35
- Claude Sonnet 5
- Claude Haiku 4.5
- GPT-5.6 Sol
- Claude Mythos 5
- Gemini 3.1 Pro
- Grok 4.5
- Grok 4.6
- Grok Build 0.1
- GLM-5.3
- Kimi K2.7 Code
- Qwen3.7 Max
- DeepSeek V4 Pro Max
- DeepSeek V4 Pro
- DeepSeek V4 Flash
- DeepSeek V4.5
- MiniMax M3
- Nemotron 3 Ultra
- Step 3.7 Flash
- Step 3.5 Flash
- Llama 4 Maverick
- Qwen3-Coder-Next
- Mistral Large 3
- Ring 2.6 1T
- MiMo v2.5 Pro
- ERNIE 5.1
- Agnes 2.0 Flash
- Fugu Ultra
- Cohere North Mini Code
- SWE-1.7
- Muse Spark 1.1
- Claude Opus 5
- Gemini 3.6 Flash
- Kimi K3
- Gemini 3.5 Flash-Lite
- Gemini 3.5 Flash Cyber
MMLU-Pro
knowledge breadth
Harder multiple-choice knowledge questions across academic subjects.
Not scored27
- Claude Sonnet 5
- Claude Haiku 4.5
- GPT-5.6 Sol
- Claude Mythos 5
- Gemini 3.1 Pro
- Grok 4.5
- Grok 4.6
- Grok Build 0.1
- GLM-5.3
- Kimi K2.7 Code
- MiniMax M3
- Step 3.7 Flash
- Step 3.5 Flash
- Qwen3-Coder-Next
- Mistral Large 3
- Ring 2.6 1T
- ERNIE 5.1
- Agnes 2.0 Flash
- Fugu Ultra
- Cohere North Mini Code
- SWE-1.7
- Muse Spark 1.1
- Claude Opus 5
- Gemini 3.6 Flash
- Kimi K3
- Gemini 3.5 Flash-Lite
- Gemini 3.5 Flash Cyber
LiveCodeBench
live coding
Can the model solve fresh programming problems (contamination-resistant coding).
Not scored27
- Claude Sonnet 5
- Claude Haiku 4.5
- GPT-5.6 Sol
- Claude Mythos 5
- Gemini 3.1 Pro
- Grok 4.6
- Grok Build 0.1
- GLM-5.3
- Kimi K2.7 Code
- Qwen3.7 Max
- MiniMax M3
- Step 3.7 Flash
- Step 3.5 Flash
- Qwen3-Coder-Next
- Ring 2.6 1T
- MiMo v2.5 Pro
- ERNIE 5.1
- Agnes 2.0 Flash
- Fugu Ultra
- Cohere North Mini Code
- SWE-1.7
- Muse Spark 1.1
- Claude Opus 5
- Gemini 3.6 Flash
- Kimi K3
- Gemini 3.5 Flash-Lite
- Gemini 3.5 Flash Cyber
Cost-Performance Map
Up and left is better · Terminal-Bench vs $/M output (log scale)
- Hy3TB 71.7 · $0.53/M
- DS Flash 0731TB 82.7 · $0.28/M
- Step 3.5FTB 51 · $0.3/M
- Agnes 2.0TB 0 · $0.2/M
- DS V4 ProTB 67.9 · $0.87/M
- NemotronTB 56.4 · $2.2/M
- Step 3.7FTB 59.5 · $1.15/M
- MiniMax M3TB 66 · $1.2/M
- Mistral L3TB 12 · $1.5/M
- Muse SparkTB 80 · $4.25/M
- GLM-5.2TB 81 · $2.42/M
- Grok 4.5TB 83.3 · $6/M
- Grok 4.6TB 88.4 · $6/M
- Qwen3.8 MaxTB 86.6 · $6/M
- LunaTB 84.7 · $6/M
- Haiku 4.5TB 44.2 · $5/M
- Gemini 3.5FTB 76.2 · $9/M
- Sonnet 5 ★TB 80.4 · $10/M
- Gemini 3.1PTB 54.2 · $12/M
- Opus 4.8TB 79 · $25/M
- TerraTB 87.4 · $15/M
- GPT-5.6 Sol ★TB 88.8 · $30/M
- Fable 5 ★TB 84.3 · $50/M
Most attractive is up and left: stronger Terminal-Bench at a lower price. Dashed line is the Pareto frontier — nothing below-left of it is both cheaper and stronger. Larger rings are frontier or leader points. Color is the Terminal-Bench band. Click a point to open the model.
no score