AI Benchmarks and Evals

A benchmark is a fixed task set with a scoring rule used to compare models; an eval is the broader practice of measuring a system on tasks that matter to you. Benchmarks can saturate and leak, and new ones are regularly introduced.

Named benchmarks

  • MMLU (2020): multiple-choice knowledge; reportedly saturated by frontier models (not rechecked) (massive-multitask-language-understanding).
  • SWE-bench (2023): resolving real GitHub issues in repositories; variants exist, see ai-coding-benchmarks.
  • Humanity’s Last Exam (2025): expert-written, hard questions (arXiv 2501.14249).
  • ARC / ARC-AGI (Chollet, 2019; ARC-AGI-2, 2025): visual puzzles that, per the ARC paper, are meant to test skill acquisition rather than memorised knowledge.
  • Chatbot Arena / MT-Bench (2023): human-preference rankings and the LLM-as-judge method (arXiv 2306.05685) (llm-as-judge-and-evals).
  • Tool use: the Berkeley function-calling leaderboard (berkeley-function-calling-leaderboard-v3—aka-berkeley-tool-calling-leaderboard-v3-).

Pitfalls

Contamination (test items in training data), overfitting to public leaderboards, vendor-reported versus independently reproduced scores, and scaffold differences (scores can differ between harnesses; opinion, not measured here). In the author’s view, task-specific evals are preferable: agent-evals-and-observability, rag-evaluation.

Related: reasoning-models, artificial-general-intelligence, artificial-analysis-llm-api-provider-leaderboard.

Sources

Open items

  • MMLU (2020) and the Berkeley leaderboard not re-checked this pass.
  • No current leaderboard scores are quoted on purpose; add dated figures only from the official leaderboards.