What AI Changes
How these connect
The difficulty of evaluating machine learning systems and what it relates to. Every line is an evidence-backed relation, and the arrangement is fixed — this address draws this picture, today and next year.
Open more4
- Expand CASP — Critical Assessment of Structure Prediction — 1 more
- Expand METR — 1 more
- Expand Benchmark saturation — 2 more
- Expand European Union Artificial Intelligence Act — 3 more
Every relation drawn, in words6
- European Union Artificial Intelligence Act · constrained_by · The difficulty of evaluating machine learning systems
The high-risk conformity regime requires assessments that the standards to perform them do not yet specify, and the deferral of its deadline followed that gap.
- CASP — Critical Assessment of Structure Prediction · contradicts · The difficulty of evaluating machine learning systems
A blind experiment with withheld targets and independent assessors removes the specific failure modes — contamination, leaderboard optimisation, self-selected test sets — that make most machine-learning evaluation weak.
- An audit finds the most-cited preference leaderboard is optimised rather than merely measured · evidence_for · The difficulty of evaluating machine learning systems
The audit documents undisclosed private testing and leaderboard overfitting, which are mechanisms by which a public score can rise without the underlying capability changing.
- METR · assesses · The difficulty of evaluating machine learning systems
The gap between measured and self-reported speed in METR's trial is a direct measurement of how unreliable perceived-productivity evidence is.
- Benchmark saturation · part_of · The difficulty of evaluating machine learning systems
A saturating benchmark is one specific way an evaluation stops carrying information; it is a component of the wider measurement problem rather than the whole of it.
- Construct validity in AI evaluation · governs · The difficulty of evaluating machine learning systems
Whether a benchmark result licenses any wider claim is decided by whether the test measures the construct it names, which is the question construct validity asks.