A systematic review of 445 benchmarks reports pervasive validity failures
node
A team of about forty researchers led from the Oxford Internet Institute reviewed 445 large-language-model benchmarks drawn from the major machine-learning and computational-linguistics conferences. About half claimed to measure abstract properties — reasoning, harmlessness — without defining them or specifying how they would be measured; 27% relied on convenience sampling; and only 16% used uncertainty estimates or statistical tests when comparing systems. The work was presented at NeurIPS in San Diego in December 2025. It does not say benchmark scores are worthless; it says most published comparisons lack the apparatus that would let a reader tell a real difference from a sampling artefact.
A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.
Read in · 1
Timeline
No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.
Connections · 2
Assembled narrative · 1
An assembled narrative is available for this object.
Observed changes · 0
No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.
Actions
/atlas?object=EVT_CONSTRUCT_VALIDITY_REVIEW_2025