Input $/M
Output $/M
256K
Context
Speed tok/s

Agent scores

Terminal-Bench 2.x

shell agent loops

Can the model finish real terminal / shell-agent tasks end to end — install tools, run commands, recover from errors.

0–100, higher is better. 70+ means a usable coding agent in a terminal; 80+ is frontier agent quality.

Editorial scores, not auto-scraped. Versions (2.0 vs 2.1) are not always comparable.

SWE-bench Pro

software engineering

Can the model solve real software-engineering problems — patches, reviews, and repo-level fixes — in a single pass.

0–100, higher is better. A high score means stronger code-review and patch quality, not just snippet completion.

This is SWE-bench Pro, not SWE-bench Verified. Vendor harness notes are marked (v).

MCP-Atlas

multi-tool use

Can the model use many Model Context Protocol tools together — pick the right tool, pass arguments, and chain calls.

0–100, higher is better. High scores matter if your agent talks to GitHub, browsers, databases, or other MCP servers.

Extended scores

BFCL

function calling

Berkeley Function Calling Leaderboard (V4): can the model call APIs with the right names, types, and arguments.

0–100, higher is better. Useful when the job is structured tool calls rather than a full agent loop.

Coverage is thin on newest models. Sourced manually from the Berkeley leaderboard.

tau-bench

multi-turn agent tasks

Can the model hold a multi-turn tool-using conversation and finish a policy-constrained task.

0–100, higher is better. High scores help for customer-support-style agents that must follow rules across turns.

56.2

AIME

math reasoning

American Invitational Mathematics Examination problems — contest math, not coding.

0–100, higher is better. A proxy for careful multi-step reasoning. Weak signal for terminal or repo agents.

69.1

MMLU-Pro

knowledge breadth

Harder multiple-choice knowledge questions across academic subjects.

0–100, higher is better. Shows general knowledge. Does not measure tool use or coding agents.

77.6

LiveCodeBench

live coding

Can the model solve fresh programming problems (contamination-resistant coding).

0–100, higher is better. Complements SWE-Pro: LiveCode is contest-style code, SWE-Pro is repo engineering.

69.1

Where to buy

No tracked provider lists this model yet.