$0.76
Input $/M
$2.42
Output $/M
1M
Context
109
Speed tok/s

Agent scores

Terminal-Bench 2.x

shell agent loops

Can the model finish real terminal / shell-agent tasks end to end — install tools, run commands, recover from errors.

0–100, higher is better. 70+ means a usable coding agent in a terminal; 80+ is frontier agent quality.

Editorial scores, not auto-scraped. Versions (2.0 vs 2.1) are not always comparable.

81

SWE-bench Pro

software engineering

Can the model solve real software-engineering problems — patches, reviews, and repo-level fixes — in a single pass.

0–100, higher is better. A high score means stronger code-review and patch quality, not just snippet completion.

This is SWE-bench Pro, not SWE-bench Verified. Vendor harness notes are marked (v).

62.1 (v)

MCP-Atlas

multi-tool use

Can the model use many Model Context Protocol tools together — pick the right tool, pass arguments, and chain calls.

0–100, higher is better. High scores matter if your agent talks to GitHub, browsers, databases, or other MCP servers.

76.8

SWE-Marathon & Long-Horizon

sustained multi-hour work

Whether the model stays useful over long agent sessions instead of collapsing after a few dozen steps.

Mixed — not one number. Use this as supporting evidence for long-running agents, not as a ranking against TB or SWE-Pro.

Entries mix DeepSWE percentages, Marathon raw scores, and tool-call counts. Bars are editorial, not a shared scale.

13.0 Marathon

Extended scores

BFCL

function calling

Berkeley Function Calling Leaderboard (V4): can the model call APIs with the right names, types, and arguments.

0–100, higher is better. Useful when the job is structured tool calls rather than a full agent loop.

Coverage is thin on newest models. Sourced manually from the Berkeley leaderboard.

74.1

tau-bench

multi-turn agent tasks

Can the model hold a multi-turn tool-using conversation and finish a policy-constrained task.

0–100, higher is better. High scores help for customer-support-style agents that must follow rules across turns.

89.7

AIME

math reasoning

American Invitational Mathematics Examination problems — contest math, not coding.

0–100, higher is better. A proxy for careful multi-step reasoning. Weak signal for terminal or repo agents.

99.2

MMLU-Pro

knowledge breadth

Harder multiple-choice knowledge questions across academic subjects.

0–100, higher is better. Shows general knowledge. Does not measure tool use or coding agents.

80.1

LiveCodeBench

live coding

Can the model solve fresh programming problems (contamination-resistant coding).

0–100, higher is better. Complements SWE-Pro: LiveCode is contest-style code, SWE-Pro is repo engineering.

65.3

Where to buy

ProviderTheir priceNoteLink
CrofAI$0.15/$0.521M ctx · 50% off listVisit site
ZenMux$1.35Best value on ZenMux
Together AI$1.40Frontier MoE
Hyperbolic$1.40Frontier MoE
OpenCode Go$4.40Current main
Ollama CloudSubsubFrontier MoE · $20/mo Pro
OpenRoutervariesVia proxy
Command CodeSubsubFrontier MoE
Kilo CodeSubvariesFrontier MoE
FactorySubsubDroid Core
Z.AI DevPackSubsubPrior gen · MIT weights

Editorial routing

Task matrix
TaskRoleNote
Long-horizon multi-tool agentAvoidTool-Decathlon ✗
Single-pass code reviewAlternative62.1
Budget and pairs