Skip to content
RichseenAtlasAtlasSign in

Construct validity in AI evaluation

node

The question of whether a test measures the thing it claims to measure — borrowed from psychometrics, and the sharpest available tool for reading machine-learning results. A benchmark has construct validity when the construct it names is defined, when the items sample that construct rather than whatever data was convenient, and when the scoring supports the inference being drawn. A systematic review of 445 large-language-model benchmarks presented at NeurIPS in December 2025 found that about half claimed to measure abstract properties such as "reasoning" or "harmlessness" without defining them, that 27% relied on convenience sampling, and that only 16% used uncertainty estimates or statistical tests when comparing models. The practical consequence is not that benchmark scores are meaningless but that a score is evidence about performance on that item set, and any wider claim — about capability, about a profession, about an economy — is a separate inference requiring separate evidence.

A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Read in · 1

Evidence · 2
Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 1
Assembled narrative · 1

Assembled from 24 blocks · 2 evidence · 19 related

  1. Story
  2. The question of whether a test measures the thing it claims to measure — borrowed from psychometrics, and the sharpest available tool for reading machine-learning results. A benchmark has construct validity when the construct it names is defined, when the items sample that construct rather than whatever data was convenient, and when the scoring supports the inference being drawn. A systematic review of 445 large-language-model benchmarks presented at NeurIPS in December 2025 found that about half claimed to measure abstract properties such as "reasoning" or "harmlessness" without defining them, that 27% relied on convenience sampling, and that only 16% used uncertainty estimates or statistical tests when comparing models. The practical consequence is not that benchmark scores are meaningless but that a score is evidence about performance on that item set, and any wider claim — about capability, about a profession, about an economy — is a separate inference requiring separate evidence.
  3. Knowledge
  4. Construct validity in AI evaluation
  5. The difficulty of evaluating machine learning systems
  6. Connections
  7. The difficulty of evaluating machine learning systems
  8. A human-parity claim in translation fails under document-level evaluation
  9. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
  10. A systematic review of 445 benchmarks reports pervasive validity failures
  11. CASP — Critical Assessment of Structure Prediction
  12. Construct validity in AI evaluation
  13. Benchmark saturation
  14. METR
  15. European Union Artificial Intelligence Act
  16. A human-parity claim in translation fails under document-level evaluation
  17. A randomised trial finds experienced developers slower with AI assistance
  18. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
  19. A systematic review of 445 benchmarks reports pervasive validity failures
  20. The Commission proposes deferring the high-risk obligations
  21. The AI Index records both fast benchmark movement and doubts about benchmarks
  22. Evidence
  23. A systematic review of 445 large-language-model benchmarks, conducted by about forty researchers led from the Oxford Internet Institute, found that roughly half claim to measure abstract phenomena such as reasoning or harmlessness without defining them or specifying how they are measured, that 27% rely on convenience sampling, and that only 16% use uncertainty estimates or statistical tests when comparing model performance; the paper offers eight recommendations for benchmark construction. V55 verification basis: the paper was not fetched, but search results carried the 445-benchmark count, the three statistics, the arXiv identifier 2511.04703, the NeurIPS 2025 venue and the OII lead and attributed them to this work and to the Institute's own announcement.
  24. A systematic study of roughly two million pairwise battles across 243 models from 42 providers over sixteen months (January 2024 – April 2025) of a leading public human-preference leaderboard found undisclosed private testing allowing providers to evaluate multiple variants and publish only the best — one provider tested 27 private variants before a public release — that up to 26.5% of prompts were duplicates or near-duplicates, that proprietary models received a disproportionate share of evaluation data, and that fine-tuning on arena-style prompts raised an arena win rate by 112% with no improvement, or a slight decline, on general benchmarks such as MMLU. V55 verification basis: the paper was not fetched, but search results carried the author list, the arXiv identifier 2504.20879, and each of these figures with attribution to this paper.
Close the narrative
Observed changes · 0

No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

Actions

Read the assembled narrativeContinue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

/atlas?object=CONCEPT_CONSTRUCT_VALIDITY&experience=CONCEPT_CONSTRUCT_VALIDITY