Skip to content
RichseenAtlasAtlasSign in

The difficulty of evaluating machine learning systems

node

The structural reason most confident statements about AI capability are weaker than they sound. A machine-learning system is characterised by its behaviour on samples, so every claim about it is a claim about a sample and an inference beyond it, and four things attack that inference at once. Training corpora are large enough that test material is frequently inside them, so a score may be measuring recall rather than generalisation. Public leaderboards are optimisation targets, and a 2025 audit of the most-cited human-preference leaderboard found undisclosed private testing at a scale that turns publication into selection of the best of many attempts. The constructs being scored are often undefined, and most benchmark papers do not report the statistics needed to say whether two systems differ. And the systems now act over long horizons with tools, where scoring an end state says nothing about the path taken. The consequence is not nihilism about measurement; it is that the strongest evidence in the field comes from designs — blind held-back tests, randomised trials, operational deployment with accountability — that are rare precisely because they are expensive.

A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Read in · 1

Evidence · 4
Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 0
    Assembled narrative · 1

    Assembled from 21 blocks · 4 evidence · 19 related

    1. Story
    2. The structural reason most confident statements about AI capability are weaker than they sound. A machine-learning system is characterised by its behaviour on samples, so every claim about it is a claim about a sample and an inference beyond it, and four things attack that inference at once. Training corpora are large enough that test material is frequently inside them, so a score may be measuring recall rather than generalisation. Public leaderboards are optimisation targets, and a 2025 audit of the most-cited human-preference leaderboard found undisclosed private testing at a scale that turns publication into selection of the best of many attempts. The constructs being scored are often undefined, and most benchmark papers do not report the statistics needed to say whether two systems differ. And the systems now act over long horizons with tools, where scoring an end state says nothing about the path taken. The consequence is not nihilism about measurement; it is that the strongest evidence in the field comes from designs — blind held-back tests, randomised trials, operational deployment with accountability — that are rare precisely because they are expensive.
    3. Knowledge
    4. The difficulty of evaluating machine learning systems
    5. Connections
    6. CASP — Critical Assessment of Structure Prediction
    7. Construct validity in AI evaluation
    8. Benchmark saturation
    9. METR
    10. European Union Artificial Intelligence Act
    11. A human-parity claim in translation fails under document-level evaluation
    12. A randomised trial finds experienced developers slower with AI assistance
    13. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
    14. A systematic review of 445 benchmarks reports pervasive validity failures
    15. The Commission proposes deferring the high-risk obligations
    16. The AI Index records both fast benchmark movement and doubts about benchmarks
    17. Evidence
    18. A systematic review of 445 large-language-model benchmarks, conducted by about forty researchers led from the Oxford Internet Institute, found that roughly half claim to measure abstract phenomena such as reasoning or harmlessness without defining them or specifying how they are measured, that 27% rely on convenience sampling, and that only 16% use uncertainty estimates or statistical tests when comparing model performance; the paper offers eight recommendations for benchmark construction. V55 verification basis: the paper was not fetched, but search results carried the 445-benchmark count, the three statistics, the arXiv identifier 2511.04703, the NeurIPS 2025 venue and the OII lead and attributed them to this work and to the Institute's own announcement.
    19. A systematic study of roughly two million pairwise battles across 243 models from 42 providers over sixteen months (January 2024 – April 2025) of a leading public human-preference leaderboard found undisclosed private testing allowing providers to evaluate multiple variants and publish only the best — one provider tested 27 private variants before a public release — that up to 26.5% of prompts were duplicates or near-duplicates, that proprietary models received a disproportionate share of evaluation data, and that fine-tuning on arena-style prompts raised an arena win rate by 112% with no improvement, or a slight decline, on general benchmarks such as MMLU. V55 verification basis: the paper was not fetched, but search results carried the author list, the arXiv identifier 2504.20879, and each of these figures with attribution to this paper.
    20. In a randomised trial, sixteen experienced open-source developers working on 246 real issues in repositories they knew well — about five years of prior experience with the projects on average — took 19% longer to complete tasks when allowed to use AI tools available between February and June 2025, while estimating afterwards that the tools had made them about 20% faster; METR has since labelled the result historical, as not necessarily reflecting current tools or workflows. Separately, METR reports a task time-horizon metric — the human-measured duration of task a model completes autonomously with 50% reliability — as having doubled approximately every seven months over 2019–2025, measured across 170 tasks against over 800 human baselines with thirteen frontier models. V55 verification basis: metr.org was not fetched. The trial figures, the arXiv identifier 2507.09089, the task and participant counts and METR's own historical caveat were carried in search results attributing them to METR; the seven-month doubling, the 170-task suite and the 800-plus human baselines were carried with attribution to METR's March 2025 post. Any specific model's time horizon, and reports of a faster doubling in the most recent period, come from secondary summaries and are NOT asserted in this pack.
    21. AI benchmark scores rose rapidly across language, reasoning, coding and mathematics in 2025 while benchmarks themselves faced growing questions about reliability: resolution on SWE-bench Verified rose from about 60% to near 100% within a year; accuracy on OSWorld rose from roughly 12% to 66.3%; frontier models gained about thirty percentage points in one year on Humanity's Last Exam; and robots succeed at only about 12% of real household tasks. V55 verification basis: the report was not fetched, but these figures were returned by a search restricted to hai.stanford.edu and are attributed there to the Index itself. Model-by-model scores, leaderboard ratings and adoption percentages that appeared alongside them in retrieved synthesis were NOT separately confirmed and are excluded from this pack.
    Close the narrative
    Observed changes · 0

    No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

    Actions

    Read the assembled narrativeContinue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

    /atlas?object=PHENOMENON_EVALUATION_DIFFICULTY&experience=PHENOMENON_EVALUATION_DIFFICULTY