$1.25
Input $/M
$3.75
Output $/M
1M
Context
197
Speed tok/s

Agent scores

Terminal-Bench 2.x

shell agent loops

Can the model finish real terminal / shell-agent tasks end to end — install tools, run commands, recover from errors.

0–100, higher is better. 70+ means a usable coding agent in a terminal; 80+ is frontier agent quality.

Editorial scores, not auto-scraped. Versions (2.0 vs 2.1) are not always comparable.

69.7

SWE-bench Pro

software engineering

Can the model solve real software-engineering problems — patches, reviews, and repo-level fixes — in a single pass.

0–100, higher is better. A high score means stronger code-review and patch quality, not just snippet completion.

This is SWE-bench Pro, not SWE-bench Verified. Vendor harness notes are marked (v).

60.6

MCP-Atlas

multi-tool use

Can the model use many Model Context Protocol tools together — pick the right tool, pass arguments, and chain calls.

0–100, higher is better. High scores matter if your agent talks to GitHub, browsers, databases, or other MCP servers.

76.4

SWE-Marathon & Long-Horizon

sustained multi-hour work

Whether the model stays useful over long agent sessions instead of collapsing after a few dozen steps.

Mixed — not one number. Use this as supporting evidence for long-running agents, not as a ranking against TB or SWE-Pro.

Entries mix DeepSWE percentages, Marathon raw scores, and tool-call counts. Bars are editorial, not a shared scale.

1,158 calls 35h proven ★

Extended scores

BFCL

function calling

Berkeley Function Calling Leaderboard (V4): can the model call APIs with the right names, types, and arguments.

0–100, higher is better. Useful when the job is structured tool calls rather than a full agent loop.

Coverage is thin on newest models. Sourced manually from the Berkeley leaderboard.

tau-bench

multi-turn agent tasks

Can the model hold a multi-turn tool-using conversation and finish a policy-constrained task.

0–100, higher is better. High scores help for customer-support-style agents that must follow rules across turns.

AIME

math reasoning

American Invitational Mathematics Examination problems — contest math, not coding.

0–100, higher is better. A proxy for careful multi-step reasoning. Weak signal for terminal or repo agents.

MMLU-Pro

knowledge breadth

Harder multiple-choice knowledge questions across academic subjects.

0–100, higher is better. Shows general knowledge. Does not measure tool use or coding agents.

89.6

LiveCodeBench

live coding

Can the model solve fresh programming problems (contamination-resistant coding).

0–100, higher is better. Complements SWE-Pro: LiveCode is contest-style code, SWE-Pro is repo engineering.

Where to buy

ProviderTheir priceNoteLink
Together AI$1.25Long-horizon
Fireworks AI$1.25Long-horizon
Hyperbolic$1.25Long-horizon
ZenMux$1.2935h proven
OpenCode Go$3.75Long-horizon
OpenRoutervariesVia proxy
Command CodeSubsubLong-horizon
Kilo CodeSubvariesLong-horizon

Editorial routing

Task matrix
TaskRoleNote
Long-horizon multi-tool agentAlternative35h proven
Long-horizon autonomousAlternative35h / 1,158 calls