AI Coding Benchmarks
Short map of the benchmarks most cited for coding agents, and how much to trust them. Current leaderboard numbers were not retrievable on 2026-09-30 or 2026-10-02 (pages are dynamic or cached), so no scores are quoted as current; check the leaderboards directly.
SWE-bench family
- The SWE-bench site (swebench.com) hosts separate leaderboards: Verified, Lite, Full, Multimodal, Multilingual and Bash Only.
- SWE-bench Verified is the human-filtered subset (500 tasks per OpenAI’s 2024 announcement; Epoch AI’s hosted version uses 484 validated samples from 12 Python repositories, 93 annotators, estimated 5-10% error rate; Epoch upgraded scaffolding, environments and token limits in Feb 2026 (v2.0.0+, 2M tokens), which raised scores, so numbers before and after are not comparable). Its saturation and possible training-data contamination are discussed in secondary sources (e.g. codesota.com; not independently assessed here). A report that OpenAI stopped evaluating on Verified could not be checked (OpenAI page returned 403): unverified.
- SWE-Bench Pro (Scale AI) uses 1,865 problems from 41 repositories (731 public instances, 276 private proprietary, 858 held out; 11 public, 12 held-out and 18 commercial repositories per the arXiv paper 2509.16941, submitted 2025-09-21), mixing GPL-licensed open source and private proprietary code to reduce contamination. Scale’s page, as cached, quotes only launch-era results (about 23% for the best models, GPT-5 and Claude Opus 4.1, on the public set), so treat current scores as unknown here.
Terminal-Bench
Benchmark of terminal-based agent tasks, hosted by the Harbor framework and the Laude Institute (with Stanford), led by Ryan Marten, Alex Shaw, Andy Konwinski and Ludwig Schmidt. The leaderboard page currently shows a 4.0 dataset/leaderboard version with per-agent resolution rates, 95% confidence intervals, cost and token usage. The task count and top entries were not visible to the fetch. The GitHub repo (harbor-framework/terminal-bench) lists compute sponsors Modal, Anthropic, OpenAI and Google and data partners Scale AI and Snorkel, and says the latest dataset is on Harbor Hub.
METR developer-productivity research
- July 2025 RCT: 16 experienced open-source developers, 246 issues; with AI tools (mostly Cursor Pro with Claude models) they took 19% longer, while believing AI had sped them up about 20%.
- 24 Feb 2026 update: METR said selection effects from wider AI adoption undermined its follow-up: developers increasingly refuse to work without AI and skip tasks. Reported speedup estimates: original cohort -18% (CI -38% to +9%), newly recruited developers -4% (CI -15% to +9%); METR is exploring other designs. Pay in the new study was lower (150/h).
- May 2026: a survey of 349 technical workers found median 1.4 to 2x self-reported value changes (self-report, not measured output). METR’s research page lists “Time Horizon 1.1” (2026-01-29: more tasks, new eval infrastructure; original doubling time about 7 months) and the MirrorCode benchmark (2026-04-10; agents reimplementing a 16,000-line codebase).
Why benchmarks are hard for agents
Multi-file changes, long runs, environment setup, test quality and training-data contamination: public tasks can leak into training data, which may inflate scores over time (the author’s view; opinion). No current scores are given in this note except where a primary source was reachable; vendor and press claims below are marked.
More detail on the main benchmarks (merged from the 2026-10-02 agentic-coding-benchmarks note)
- SWE-bench (Princeton/Stanford; ICLR 2024 oral): a patch that resolves a real GitHub issue in a Python repo, graded by tests. Variants: full, Lite, Verified, Multimodal, Multilingual; cloud evaluation via Modal or
sb-cli; companion tools SWE-agent and SWE-smith. The full set is 2,294 tasks from 12 Python repos (original paper figure, not re-checked). Official repo active (github.com/SWE-bench/SWE-bench). - SWE-bench Verified saturation: secondary sources say OpenAI published an audit in February 2026 (“Why SWE-bench Verified no longer measures frontier coding progress”), found frontier models could reproduce gold patches and that about 59% of the hardest unsolved tasks had flawed tests, and stopped reporting it (OpenAI’s page returned 403; unverified at primary level).
- SWE-bench Pro: original paper says frontier models scored below 45% Pass@1. Secondary press (alphasignal.ai) says OpenAI retracted its recommendation in July 2026 after judging about 30% of the 731 public tasks broken; unverified.
- Terminal-Bench 2.0 (secondary leaderboards, September 2026): Claude Fable 5 at about 84% (model-card self-report) and GPT-5.5 at about 82-83%; scaffold plus model combinations around 82% (ForgeCode). Not verified against the official board (the board itself was seen at version 4.0 on 2026-09-30, see above).
Other benchmarks
| Benchmark | What it measures | Status (2026) |
|---|---|---|
| LiveCodeBench | Contamination-free code generation from new LeetCode, AtCoder and Codeforces problems; also self-repair, code execution, test-output prediction. Filter by release date. | Primary: GitHub; release_v6 covers May 2023 - Apr 2025 with 1,055 problems. |
| LiveSWEBench (LiveBench team) | Coding assistants (Copilot, Cursor, Aider, SWE-agent, OpenHands) on agentic, targeted-edit and autocomplete tasks from real issue and PR pairs (C++, Java, TypeScript, Python) | Exists (liveswebench.ai). Earlier scores came from search snippets, not current. |
| AgentBench (THUDM, ICLR 2024) | General LLM-as-agent benchmark over 8 environments (OS, database, knowledge graph, card game, lateral thinking, ALFWorld, WebShop, Mind2Web); not coding-specific | Primary: GitHub |
| DevQualityEval (Symflower) | Code-generation quality in Go, Java, Ruby etc. (unit-test writing, compile success, coverage, token economy); monthly releases since April 2024 | Exists (docs.symflower.com) |
| HumanEval / HumanEval+ | 164 hand-written Python function problems; + adds more tests | Described as saturated (opinion; not sourced here) |
| METR time horizon | Length of human task (in time) an agent completes with 50% success; doubling roughly every 7 months since 2019, closer to 4 months over 2024-2025 (secondary summaries; the 7-month figure appears on METR’s research page) | Secondary summaries: a Feb-Mar 2026 pilot estimated the most capable model at roughly 16-20 hours at 50%, but METR says measurements above about 16 hours are unreliable with the current suite (Time Horizon 1.1, January 2026) |
Practical guidance (author’s suggestions, opinion)
- Use several benchmarks and prefer fresh or private ones (LiveCodeBench windows, SWE-bench Pro commercial/held-out) over saturated ones.
- Compare model plus scaffold: the same model differs by many points across harnesses; treat vendor self-reported scores as unverified until reproduced.
- Test on your own repositories and track cost and tokens, not only resolve rate; add human review.
Reading benchmark claims
Vendor-reported scores depend on scaffold, harness and model settings; compare like with like. Self-reported productivity and measured productivity have diverged in METR’s data.
Related
agent-evals-and-observability, llm-as-judge-and-evals, task-decomposition-patterns, claude-agent-sdk.
Sources
- https://www.swebench.com (leaderboard list; accessed 2026-09-30)
- https://labs.scale.com/leaderboard/swe_bench_pro_public (accessed 2026-09-30; cached launch-era content)
- https://www.tbench.ai/leaderboard and https://github.com/harbor-framework/terminal-bench (accessed 2026-09-30)
- https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ (accessed 2026-09-30)
- https://metr.org/blog/2026-02-24-uplift-update/ (accessed 2026-09-30)
- https://metr.org/research (accessed 2026-09-30)
- https://epoch.ai/benchmarks/swe-bench-verified (accessed 2026-10-02)
- https://arxiv.org/abs/2509.16941 (SWE-Bench Pro paper; accessed 2026-10-02)
- https://github.com/harbor-framework/terminal-bench (accessed 2026-10-02)
Open items
- No current leaderboard scores for SWE-bench Verified, SWE-Bench Pro or Terminal-Bench 4.0 could be fetched (dynamic pages; Scale’s page shows launch-era numbers only), so the note carries none.
- Whether OpenAI stopped reporting SWE-bench Verified is unverified (OpenAI page returned 403).
- Terminal-Bench 4.0 task count not found.
- The May 2026 METR survey (349 workers) was not re-fetched.
- https://github.com/SWE-bench/SWE-bench, https://github.com/LiveCodeBench/LiveCodeBench, https://github.com/THUDM/AgentBench, https://docs.symflower.com/docs/devqualityeval/, https://liveswebench.ai/, https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ (accessed 2026-10-02 by the merged note)
- Secondary: https://www.codesota.com/news/swe-bench-contamination-debate, https://alphasignal.ai/news/openai-retracts-swe-bench-pro-after-finding-30-of-tasks-broken, https://benchlm.ai/benchmarks/terminal-bench-2
- Open items carried over: the Verified retirement and SWE-bench Pro retraction rest on secondary sources; Terminal-Bench 2.0 leaderboard figures are secondary; the Verified release date (August 2024) was not re-verified.