Measuring what Matters: Construct Validity in Large Language Model Benchmarks
evidence
A systematic review of 445 large-language-model benchmarks, conducted by about forty researchers led from the Oxford Internet Institute, found that roughly half claim to measure abstract phenomena such as reasoning or harmlessness without defining them or specifying how they are measured, that 27% rely on convenience sampling, and that only 16% use uncertainty estimates or statistical tests when comparing model performance; the paper offers eight recommendations for benchmark construction. V55 verification basis: the paper was not fetched, but search results carried the 445-benchmark count, the three statistics, the arXiv identifier 2511.04703, the NeurIPS 2025 venue and the OII lead and attributed them to this work and to the Institute's own announcement.
A evidence is not a place. Drawing it on a map would assert something about the world that no stored fact supports.
Evidence · 0
This object cites no evidence.
Timeline
No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.
Connections · 0
Assembled narrative · 0
This object does not clear the publishing floor: an assembled narrative needs a description and at least one cited piece of evidence.
Observed changes · 0
No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.
Actions
/atlas?object=EVIDENCE_EV_OXFORD_CONSTRUCT_VALIDITY_2025