Skip to content
RichseenAtlasAtlasSign in

The difficulty of evaluating machine learning systems

node

The structural reason most confident statements about AI capability are weaker than they sound. A machine-learning system is characterised by its behaviour on samples, so every claim about it is a claim about a sample and an inference beyond it, and four things attack that inference at once. Training corpora are large enough that test material is frequently inside them, so a score may be measuring recall rather than generalisation. Public leaderboards are optimisation targets, and a 2025 audit of the most-cited human-preference leaderboard found undisclosed private testing at a scale that turns publication into selection of the best of many attempts. The constructs being scored are often undefined, and most benchmark papers do not report the statistics needed to say whether two systems differ. And the systems now act over long horizons with tools, where scoring an end state says nothing about the path taken. The consequence is not nihilism about measurement; it is that the strongest evidence in the field comes from designs — blind held-back tests, randomised trials, operational deployment with accountability — that are rare precisely because they are expensive.

A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Read in · 1

Evidence · 4
Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 0
    Assembled narrative · 1

    An assembled narrative is available for this object.

    Observed changes · 0

    No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

    Actions

    Read the assembled narrativeContinue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

    /atlas?object=PHENOMENON_EVALUATION_DIFFICULTY