A systematic review of 445 benchmarks reports pervasive validity failures
A team of about forty researchers led from the Oxford Internet Institute reviewed 445 large-language-model benchmarks drawn from the major machine-learning and computational-linguistics conferences. About half claimed to measure abstract properties — reasoning, harmlessness — without defining them or specifying how they would be measured; 27% relied on convenience sampling; and only 16% used uncertainty estimates or statistical tests when comparing systems. The work was presented at NeurIPS in San Diego in December 2025. It does not say benchmark scores are worthless; it says most published comparisons lack the apparatus that would let a reader tell a real difference from a sampling artefact.
Observed — measured or witnessed, with the observation cited.
Subjects
- Construct validity in AI evaluation · concept · not located
- The difficulty of evaluating machine learning systems · phenomenon · not located
Evidence
Verified — Every field is supported by a cited authoritative source.
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
A systematic review of 445 large-language-model benchmarks, conducted by about forty researchers led from the Oxford Internet Institute, found that roughly half claim to measure abstract phenomena such as reasoning or harmlessness without defining them or specifying how they are measured, that 27% rely on convenience sampling, and that only 16% use uncertainty estimates or statistical tests when comparing model performance; the paper offers eight recommendations for benchmark construction. V55 verification basis: the paper was not fetched, but search results carried the 445-benchmark count, the three statistics, the arXiv identifier 2511.04703, the NeurIPS 2025 venue and the OII lead and attributed them to this work and to the Institute's own announcement.
NeurIPS 2025 proceedings; Oxford Internet Institute, University of Oxford · 2025-11 · Verified
Bean, A. M. et al., arXiv:2511.04703; NeurIPS 2025, San Diego, 2–7 December 2025; Oxford Internet Institute announcement, "Study identifies weaknesses in how AI systems are evaluated"; Oxford University Research Archive record
Recorded on Construct validity in AI evaluation
- A human-parity claim in translation fails under document-level evaluation
- An audit finds the most-cited preference leaderboard is optimised rather than merely measured
- A systematic review of 445 benchmarks reports pervasive validity failures
- A randomised trial finds experienced developers slower with AI assistance
- The Commission proposes deferring the high-risk obligations
- The AI Index records both fast benchmark movement and doubts about benchmarks
Related
Related because they share a subject in the Atlas — never because the text looks similar.
- The AI Index records both fast benchmark movement and doubts about benchmarks · 2026
- The Commission proposes deferring the high-risk obligations · 2025-11-19
- A randomised trial finds experienced developers slower with AI assistance · 2025-07-10
- An audit finds the most-cited preference leaderboard is optimised rather than merely measured · 2025-04
- A human-parity claim in translation fails under document-level evaluation · 2018