AI Coding Benchmarks

Short map of the benchmarks most cited for coding agents, and how much to trust them. Current leaderboard numbers were not retrievable on 2026-09-30 or 2026-10-02 (pages are dynamic or cached), so no scores are quoted as current; check the leaderboards directly.

SWE-bench family

  • The SWE-bench site (swebench.com) hosts separate leaderboards: Verified, Lite, Full, Multimodal, Multilingual and Bash Only.
  • SWE-bench Verified is the human-filtered subset (500 tasks per OpenAI’s 2024 announcement; Epoch AI’s hosted version uses 484 validated samples from 12 Python repositories, 93 annotators, estimated 5-10% error rate; Epoch upgraded scaffolding, environments and token limits in Feb 2026 (v2.0.0+, 2M tokens), which raised scores, so numbers before and after are not comparable). Its saturation and possible training-data contamination are discussed in secondary sources (e.g. codesota.com; not independently assessed here). A report that OpenAI stopped evaluating on Verified could not be checked (OpenAI page returned 403): unverified.
  • SWE-Bench Pro (Scale AI) uses 1,865 problems from 41 repositories (731 public instances, 276 private proprietary, 858 held out; 11 public, 12 held-out and 18 commercial repositories per the arXiv paper 2509.16941, submitted 2025-09-21), mixing GPL-licensed open source and private proprietary code to reduce contamination. Scale’s page, as cached, quotes only launch-era results (about 23% for the best models, GPT-5 and Claude Opus 4.1, on the public set), so treat current scores as unknown here.

Terminal-Bench

Benchmark of terminal-based agent tasks, hosted by the Harbor framework and the Laude Institute (with Stanford), led by Ryan Marten, Alex Shaw, Andy Konwinski and Ludwig Schmidt. The leaderboard page currently shows a 4.0 dataset/leaderboard version with per-agent resolution rates, 95% confidence intervals, cost and token usage. The task count and top entries were not visible to the fetch. The GitHub repo (harbor-framework/terminal-bench) lists compute sponsors Modal, Anthropic, OpenAI and Google and data partners Scale AI and Snorkel, and says the latest dataset is on Harbor Hub.

METR developer-productivity research

  • July 2025 RCT: 16 experienced open-source developers, 246 issues; with AI tools (mostly Cursor Pro with Claude models) they took 19% longer, while believing AI had sped them up about 20%.
  • 24 Feb 2026 update: METR said selection effects from wider AI adoption undermined its follow-up: developers increasingly refuse to work without AI and skip tasks. Reported speedup estimates: original cohort -18% (CI -38% to +9%), newly recruited developers -4% (CI -15% to +9%); METR is exploring other designs. Pay in the new study was lower (150/h).
  • May 2026: a survey of 349 technical workers found median 1.4 to 2x self-reported value changes (self-report, not measured output). METR’s research page lists “Time Horizon 1.1” (2026-01-29: more tasks, new eval infrastructure; original doubling time about 7 months) and the MirrorCode benchmark (2026-04-10; agents reimplementing a 16,000-line codebase).

Why benchmarks are hard for agents

Multi-file changes, long runs, environment setup, test quality and training-data contamination: public tasks can leak into training data, which may inflate scores over time (the author’s view; opinion). No current scores are given in this note except where a primary source was reachable; vendor and press claims below are marked.

More detail on the main benchmarks (merged from the 2026-10-02 agentic-coding-benchmarks note)

  • SWE-bench (Princeton/Stanford; ICLR 2024 oral): a patch that resolves a real GitHub issue in a Python repo, graded by tests. Variants: full, Lite, Verified, Multimodal, Multilingual; cloud evaluation via Modal or sb-cli; companion tools SWE-agent and SWE-smith. The full set is 2,294 tasks from 12 Python repos (original paper figure, not re-checked). Official repo active (github.com/SWE-bench/SWE-bench).
  • SWE-bench Verified saturation: secondary sources say OpenAI published an audit in February 2026 (“Why SWE-bench Verified no longer measures frontier coding progress”), found frontier models could reproduce gold patches and that about 59% of the hardest unsolved tasks had flawed tests, and stopped reporting it (OpenAI’s page returned 403; unverified at primary level).
  • SWE-bench Pro: original paper says frontier models scored below 45% Pass@1. Secondary press (alphasignal.ai) says OpenAI retracted its recommendation in July 2026 after judging about 30% of the 731 public tasks broken; unverified.
  • Terminal-Bench 2.0 (secondary leaderboards, September 2026): Claude Fable 5 at about 84% (model-card self-report) and GPT-5.5 at about 82-83%; scaffold plus model combinations around 82% (ForgeCode). Not verified against the official board (the board itself was seen at version 4.0 on 2026-09-30, see above).

Other benchmarks

BenchmarkWhat it measuresStatus (2026)
LiveCodeBenchContamination-free code generation from new LeetCode, AtCoder and Codeforces problems; also self-repair, code execution, test-output prediction. Filter by release date.Primary: GitHub; release_v6 covers May 2023 - Apr 2025 with 1,055 problems.
LiveSWEBench (LiveBench team)Coding assistants (Copilot, Cursor, Aider, SWE-agent, OpenHands) on agentic, targeted-edit and autocomplete tasks from real issue and PR pairs (C++, Java, TypeScript, Python)Exists (liveswebench.ai). Earlier scores came from search snippets, not current.
AgentBench (THUDM, ICLR 2024)General LLM-as-agent benchmark over 8 environments (OS, database, knowledge graph, card game, lateral thinking, ALFWorld, WebShop, Mind2Web); not coding-specificPrimary: GitHub
DevQualityEval (Symflower)Code-generation quality in Go, Java, Ruby etc. (unit-test writing, compile success, coverage, token economy); monthly releases since April 2024Exists (docs.symflower.com)
HumanEval / HumanEval+164 hand-written Python function problems; + adds more testsDescribed as saturated (opinion; not sourced here)
METR time horizonLength of human task (in time) an agent completes with 50% success; doubling roughly every 7 months since 2019, closer to 4 months over 2024-2025 (secondary summaries; the 7-month figure appears on METR’s research page)Secondary summaries: a Feb-Mar 2026 pilot estimated the most capable model at roughly 16-20 hours at 50%, but METR says measurements above about 16 hours are unreliable with the current suite (Time Horizon 1.1, January 2026)

Practical guidance (author’s suggestions, opinion)

  • Use several benchmarks and prefer fresh or private ones (LiveCodeBench windows, SWE-bench Pro commercial/held-out) over saturated ones.
  • Compare model plus scaffold: the same model differs by many points across harnesses; treat vendor self-reported scores as unverified until reproduced.
  • Test on your own repositories and track cost and tokens, not only resolve rate; add human review.

Reading benchmark claims

Vendor-reported scores depend on scaffold, harness and model settings; compare like with like. Self-reported productivity and measured productivity have diverged in METR’s data.

agent-evals-and-observability, llm-as-judge-and-evals, task-decomposition-patterns, claude-agent-sdk.

Sources

Open items