AI Benchmarks and Evals
A benchmark is a fixed task set with a scoring rule used to compare models; an eval is the broader practice of measuring a system on tasks that matter to you. Benchmarks can saturate and leak, and new ones are regularly introduced.
Named benchmarks
- MMLU (2020): multiple-choice knowledge; reportedly saturated by frontier models (not rechecked) (massive-multitask-language-understanding).
- SWE-bench (2023): resolving real GitHub issues in repositories; variants exist, see ai-coding-benchmarks.
- Humanity’s Last Exam (2025): expert-written, hard questions (arXiv 2501.14249).
- ARC / ARC-AGI (Chollet, 2019; ARC-AGI-2, 2025): visual puzzles that, per the ARC paper, are meant to test skill acquisition rather than memorised knowledge.
- Chatbot Arena / MT-Bench (2023): human-preference rankings and the LLM-as-judge method (arXiv 2306.05685) (llm-as-judge-and-evals).
- Tool use: the Berkeley function-calling leaderboard (berkeley-function-calling-leaderboard-v3—aka-berkeley-tool-calling-leaderboard-v3-).
Pitfalls
Contamination (test items in training data), overfitting to public leaderboards, vendor-reported versus independently reproduced scores, and scaffold differences (scores can differ between harnesses; opinion, not measured here). In the author’s view, task-specific evals are preferable: agent-evals-and-observability, rag-evaluation.
Related: reasoning-models, artificial-general-intelligence, artificial-analysis-llm-api-provider-leaderboard.
Sources
- https://arxiv.org/abs/2310.06770 (SWE-bench, 2023)
- https://arxiv.org/abs/2501.14249 (Humanity’s Last Exam, 2025)
- https://arxiv.org/abs/1911.01547 (On the Measure of Intelligence / ARC, 2019)
- https://arxiv.org/abs/2505.11831 (ARC-AGI-2, 2025)
- https://arxiv.org/abs/2306.05685 (LLM-as-a-Judge, MT-Bench, Chatbot Arena, 2023)
- Titles and submission dates confirmed via arXiv API 2026-10-07 (primary): SWE-bench 2023-10, HLE 2025-01, ARC 2019-11, ARC-AGI-2 2025-05, MT-Bench/Chatbot Arena paper 2023-06.
Open items
- MMLU (2020) and the Berkeley leaderboard not re-checked this pass.
- No current leaderboard scores are quoted on purpose; add dated figures only from the official leaderboards.