Skip to content
RichseenAtlasAtlasSign in

A systematic review of 445 benchmarks reports pervasive validity failures

node

A team of about forty researchers led from the Oxford Internet Institute reviewed 445 large-language-model benchmarks drawn from the major machine-learning and computational-linguistics conferences. About half claimed to measure abstract properties — reasoning, harmlessness — without defining them or specifying how they would be measured; 27% relied on convenience sampling; and only 16% used uncertainty estimates or statistical tests when comparing systems. The work was presented at NeurIPS in San Diego in December 2025. It does not say benchmark scores are worthless; it says most published comparisons lack the apparatus that would let a reader tell a real difference from a sampling artefact.

A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Read in · 1

Evidence · 1
Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 2
Assembled narrative · 1

Assembled from 26 blocks · 1 evidence · 19 related

  1. Story
  2. A team of about forty researchers led from the Oxford Internet Institute reviewed 445 large-language-model benchmarks drawn from the major machine-learning and computational-linguistics conferences. About half claimed to measure abstract properties — reasoning, harmlessness — without defining them or specifying how they would be measured; 27% relied on convenience sampling; and only 16% used uncertainty estimates or statistical tests when comparing systems. The work was presented at NeurIPS in San Diego in December 2025. It does not say benchmark scores are worthless; it says most published comparisons lack the apparatus that would let a reader tell a real difference from a sampling artefact.
  3. Knowledge
  4. Construct validity in AI evaluation
  5. The difficulty of evaluating machine learning systems
  6. A systematic review of 445 benchmarks reports pervasive validity failures
  7. Connections
  8. Construct validity in AI evaluation
  9. The difficulty of evaluating machine learning systems
  10. The difficulty of evaluating machine learning systems
  11. A human-parity claim in translation fails under document-level evaluation
  12. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
  13. A systematic review of 445 benchmarks reports pervasive validity failures
  14. CASP — Critical Assessment of Structure Prediction
  15. Construct validity in AI evaluation
  16. Benchmark saturation
  17. METR
  18. European Union Artificial Intelligence Act
  19. A human-parity claim in translation fails under document-level evaluation
  20. A randomised trial finds experienced developers slower with AI assistance
  21. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
  22. A systematic review of 445 benchmarks reports pervasive validity failures
  23. The Commission proposes deferring the high-risk obligations
  24. The AI Index records both fast benchmark movement and doubts about benchmarks
  25. Evidence
  26. A systematic review of 445 large-language-model benchmarks, conducted by about forty researchers led from the Oxford Internet Institute, found that roughly half claim to measure abstract phenomena such as reasoning or harmlessness without defining them or specifying how they are measured, that 27% rely on convenience sampling, and that only 16% use uncertainty estimates or statistical tests when comparing model performance; the paper offers eight recommendations for benchmark construction. V55 verification basis: the paper was not fetched, but search results carried the 445-benchmark count, the three statistics, the arXiv identifier 2511.04703, the NeurIPS 2025 venue and the OII lead and attributed them to this work and to the Institute's own announcement.
Close the narrative
Observed changes · 0

No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

Actions

Read the assembled narrativeContinue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

/atlas?object=EVT_CONSTRUCT_VALIDITY_REVIEW_2025&experience=EVT_CONSTRUCT_VALIDITY_REVIEW_2025