Skip to content
RichseenAtlasAtlasSign in

Construct validity in AI evaluation

node

The question of whether a test measures the thing it claims to measure — borrowed from psychometrics, and the sharpest available tool for reading machine-learning results. A benchmark has construct validity when the construct it names is defined, when the items sample that construct rather than whatever data was convenient, and when the scoring supports the inference being drawn. A systematic review of 445 large-language-model benchmarks presented at NeurIPS in December 2025 found that about half claimed to measure abstract properties such as "reasoning" or "harmlessness" without defining them, that 27% relied on convenience sampling, and that only 16% used uncertainty estimates or statistical tests when comparing models. The practical consequence is not that benchmark scores are meaningless but that a score is evidence about performance on that item set, and any wider claim — about capability, about a profession, about an economy — is a separate inference requiring separate evidence.

A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Read in · 1

Evidence · 2
Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 1
Assembled narrative · 1

An assembled narrative is available for this object.

Observed changes · 0

No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

Actions

Read the assembled narrativeContinue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

/atlas?object=CONCEPT_CONSTRUCT_VALIDITY