The difficulty of evaluating machine learning systems
node
The structural reason most confident statements about AI capability are weaker than they sound. A machine-learning system is characterised by its behaviour on samples, so every claim about it is a claim about a sample and an inference beyond it, and four things attack that inference at once. Training corpora are large enough that test material is frequently inside them, so a score may be measuring recall rather than generalisation. Public leaderboards are optimisation targets, and a 2025 audit of the most-cited human-preference leaderboard found undisclosed private testing at a scale that turns publication into selection of the best of many attempts. The constructs being scored are often undefined, and most benchmark papers do not report the statistics needed to say whether two systems differ. And the systems now act over long horizons with tools, where scoring an end state says nothing about the path taken. The consequence is not nihilism about measurement; it is that the strongest evidence in the field comes from designs — blind held-back tests, randomised trials, operational deployment with accountability — that are rare precisely because they are expensive.
A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.
Read in · 1
Evidence · 4
Timeline
No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.
Connections · 0
Assembled narrative · 1
An assembled narrative is available for this object.
Observed changes · 0
No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.
Actions
/atlas?object=PHENOMENON_EVALUATION_DIFFICULTY