← Signals
Recentmonthscience

A systematic review of 445 benchmarks reports pervasive validity failures

A team of about forty researchers led from the Oxford Internet Institute reviewed 445 large-language-model benchmarks drawn from the major machine-learning and computational-linguistics conferences. About half claimed to measure abstract properties — reasoning, harmlessness — without defining them or specifying how they would be measured; 27% relied on convenience sampling; and only 16% used uncertainty estimates or statistical tests when comparing systems. The work was presented at NeurIPS in San Diego in December 2025. It does not say benchmark scores are worthless; it says most published comparisons lack the apparatus that would let a reader tell a real difference from a sampling artefact.

Observed — measured or witnessed, with the observation cited.

Subjects

Evidence

Verified — Every field is supported by a cited authoritative source.

Recorded on Construct validity in AI evaluation

  1. A human-parity claim in translation fails under document-level evaluation
  2. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
  3. A systematic review of 445 benchmarks reports pervasive validity failures
  4. A randomised trial finds experienced developers slower with AI assistance
  5. The Commission proposes deferring the high-risk obligations
  6. The AI Index records both fast benchmark movement and doubts about benchmarks

Related

Related because they share a subject in the Atlas — never because the text looks similar.