Overview
Compare models and providers for agent work. Start by finding a model, or skip ahead if you already know which ones you need.
52 models · 19 providers · 43 with Terminal-Bench
Benchmark Leaders
Top five on each agent bench. Color is the score band. 80 is frontier-quality.
Terminal-Bench 2.x
shell agent loops
Can the model finish real terminal / shell-agent tasks end to end — install tools, run commands, recover from errors.
Editorial scores, not auto-scraped. Versions (2.0 vs 2.1) are not always comparable.
SWE-bench Pro
software engineering
Can the model solve real software-engineering problems — patches, reviews, and repo-level fixes — in a single pass.
This is SWE-bench Pro, not SWE-bench Verified. Vendor harness notes are marked (v).
MCP-Atlas
multi-tool use
Can the model use many Model Context Protocol tools together — pick the right tool, pass arguments, and chain calls.