Skip to content
RichseenAtlasJourneySign in

What AI Changes

How these connect

The difficulty of evaluating machine learning systems and what it relates to. Every line is an evidence-backed relation, and the arrangement is fixed — this address draws this picture, today and next year.

constrained_by — The high-risk conformity regime requires assessments that the standards to perform them do not yet specify, and the deferral of its deadline followed that gap.contradicts — A blind experiment with withheld targets and independent assessors removes the specific failure modes — contamination, leaderboard optimisation, self-selected test sets — that make most machine-learning evaluation weak.evidence_for — The audit documents undisclosed private testing and leaderboard overfitting, which are mechanisms by which a public score can rise without the underlying capability changing.assesses — The gap between measured and self-reported speed in METR's trial is a direct measurement of how unreliable perceived-productivity evidence is.part_of — A saturating benchmark is one specific way an evaluation stops carrying information; it is a component of the wider measurement problem rather than the whole of it.governs — Whether a benchmark result licenses any wider claim is decided by whether the test measures the construct it names, which is the question construct validity asks.Construct validity in AI evaluationConstruct validity in AI …An audit finds the most-cited preference leaderboard is optimised rather than merely measuredAn audit finds the most-c…CASP — Critical Assessment of Structure PredictionCASP — Critical Assessmen…METRMETRBenchmark saturationBenchmark saturationThe difficulty of evaluating machine learning systemsThe difficulty of evaluat…European Union Artificial Intelligence ActEuropean Union Artificial…
7 drawn · 7 reachable and not drawn · positions derived from object ids, never from a simulation
Open more4
Every relation drawn, in words6
  1. European Union Artificial Intelligence Act · constrained_by · The difficulty of evaluating machine learning systems

    The high-risk conformity regime requires assessments that the standards to perform them do not yet specify, and the deferral of its deadline followed that gap.

  2. CASP — Critical Assessment of Structure Prediction · contradicts · The difficulty of evaluating machine learning systems

    A blind experiment with withheld targets and independent assessors removes the specific failure modes — contamination, leaderboard optimisation, self-selected test sets — that make most machine-learning evaluation weak.

  3. An audit finds the most-cited preference leaderboard is optimised rather than merely measured · evidence_for · The difficulty of evaluating machine learning systems

    The audit documents undisclosed private testing and leaderboard overfitting, which are mechanisms by which a public score can rise without the underlying capability changing.

  4. METR · assesses · The difficulty of evaluating machine learning systems

    The gap between measured and self-reported speed in METR's trial is a direct measurement of how unreliable perceived-productivity evidence is.

  5. Benchmark saturation · part_of · The difficulty of evaluating machine learning systems

    A saturating benchmark is one specific way an evaluation stops carrying information; it is a component of the wider measurement problem rather than the whole of it.

  6. Construct validity in AI evaluation · governs · The difficulty of evaluating machine learning systems

    Whether a benchmark result licenses any wider claim is decided by whether the test measures the construct it names, which is the question construct validity asks.