Massive Multitask Language Understanding (MMLU)
MMLU is a benchmark, not a training approach. The earlier version of this note described it as a multitask training method; that was wrong. MMLU is a multiple-choice test introduced in “Measuring Massive Multitask Language Understanding” (Hendrycks, Burns, Basart, Zou, Mazeika, Song, Steinhardt; arXiv Sept 2020, ICLR 2021). It covers 57 tasks, including elementary mathematics, US history, computer science and law. At the time most models scored near random; the largest GPT-3 was about 20 points above chance, with weak results on morality and law.
Why it is no longer a frontier yardstick
- Saturation and errors: “Are We Done with MMLU?” (arXiv June 2024) estimated 6.49% of MMLU questions contain errors (57% in the virology subset) and released MMLU-Redux, 5,700 re-annotated questions. A score ceiling below 100% therefore partly reflects bad questions. (That frontier models now cluster at the top of MMLU is widely reported; I did not find a primary leaderboard to cite, so no score is quoted here.)
- MMLU-Pro (arXiv June 2024, NeurIPS 2024): 10 answer options instead of 4, more reasoning-heavy questions, trivial and noisy ones removed; accuracy drops 16 to 33% versus MMLU and results are less sensitive to prompt wording.
What frontier evaluation looks like in 2026
Vendors now headline harder, agentic or expert-level tests. Anthropic’s Fable 5.1 announcement, for example, cites Terminal-Bench 4.0 (55.8%), Humanity’s Last Exam (60.9% without tools, 65.0% with tools), OSWorld 2.0, GDPval-AA v2, CursorBench and AutomationBench (vendor-reported; see claude-fable-5). Main benchmarks:
| Benchmark | What it tests | Verified facts |
|---|---|---|
| GPQA (Diamond is its hardest subset) | Graduate-level “Google-proof” science questions | 448 multiple-choice biology, physics, chemistry questions (arXiv Nov 2023); PhD experts 65%, skilled non-experts with web access 34%, GPT-4 39% |
| Humanity’s Last Exam (HLE) | Expert-written closed-ended academic questions | 2,500 questions, 100+ subjects, from the Center for AI Safety and Scale AI; published in Nature 28 Jan 2026 (vol. 649) |
| SWE-bench Verified | Fixing real GitHub issues | Part of the SWE-bench family (Verified, Multimodal, Multilingual, Lite); scores above 70% reported for top models |
| SWE-bench Pro (Scale AI) | Harder, contamination-resistant software engineering | 1,865 tasks, 41 repositories (731 public, 276 private, 858 held-out); leaderboard page lists Muse Spark 1.1 at 61.5% and gpt-5.4 (xHigh) at 59.1% on the public set; the same page says top models scored about 23% on the original release versus 70%+ on SWE-bench Verified |
| ARC-AGI (1, 2, 3) | Fluid intelligence, skill acquisition | ARC-AGI-3 is an interactive benchmark where agents explore novel environments and learn goals on the fly; ARC Prize 2026 is running on Kaggle |
| Terminal-Bench, OSWorld, GDPval | Agentic terminal, computer use and economically valuable tasks | Cited in vendor launch posts; definitions not independently verified here |
Caveats: leaderboard pages change constantly (the HLE page’s own table was visibly older than current models), vendor numbers use different tool and effort settings, and contamination is an ongoing concern, which is why private held-out sets (SWE-bench Pro) and interactive tests (ARC-AGI-3) emerged. Compare with independent aggregators, e.g. llm-comparisons-artificialanalysis.ai.
Related: large-language-model, generative-pretrained-transformer, artificial-general-intelligence.
Sources
- https://arxiv.org/abs/2009.03300 (accessed 2026-09-30)
- https://arxiv.org/abs/2406.01574 - MMLU-Pro (accessed 2026-09-30)
- https://arxiv.org/abs/2406.04127 - Are We Done with MMLU? (accessed 2026-09-30)
- https://arxiv.org/abs/2311.12022 - GPQA (accessed 2026-09-30)
- https://lastexam.ai (accessed 2026-09-30)
- https://labs.scale.com/leaderboard/swe_bench_pro_public (accessed 2026-09-30)
- https://www.swebench.com (accessed 2026-09-30)
- https://arcprize.org/arc-agi and https://arcprize.org/arc-agi/3 (accessed 2026-09-30)
- https://www.anthropic.com/claude-fable-and-mythos-5-1 (accessed 2026-09-30)
- OpenAI’s SWE-bench Verified page returned 403; its claims about Verified being retired are unverified.