Skip to content
RichseenAtlasAtlasSign in

A systematic review of 445 benchmarks reports pervasive validity failures

node

A team of about forty researchers led from the Oxford Internet Institute reviewed 445 large-language-model benchmarks drawn from the major machine-learning and computational-linguistics conferences. About half claimed to measure abstract properties — reasoning, harmlessness — without defining them or specifying how they would be measured; 27% relied on convenience sampling; and only 16% used uncertainty estimates or statistical tests when comparing systems. The work was presented at NeurIPS in San Diego in December 2025. It does not say benchmark scores are worthless; it says most published comparisons lack the apparatus that would let a reader tell a real difference from a sampling artefact.

A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Read in · 1

Evidence · 1
Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 2
Assembled narrative · 1

An assembled narrative is available for this object.

Observed changes · 0

No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

Actions

Read the assembled narrativeContinue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

/atlas?object=EVT_CONSTRUCT_VALIDITY_REVIEW_2025