An audit finds the most-cited preference leaderboard is optimised rather than merely measured
node
Singh and colleagues analysed roughly two million pairwise battles across 243 models from 42 providers over sixteen months of a public human-preference leaderboard. They reported undisclosed private testing under which a provider could evaluate many variants and publish only the best — one provider tested 27 private variants before releasing one publicly — that up to 26.5% of prompts were duplicates or near-duplicates, and that fine-tuning on arena-style prompts raised an arena win rate by 112% with no gain, or a slight loss, on general benchmarks. The last result is the important one: it separates performing well on the instrument from performing well on the thing the instrument was meant to stand for.
A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.
Read in · 1
Evidence · 1
Timeline
No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.
Connections · 3
- The difficulty of evaluating machine learning systemsevidence_for
- Construct validity in AI evaluationconcerns
- Benchmark saturationconcerns
Assembled narrative · 1
Assembled from 34 blocks · 1 evidence · 23 related
- Story
- Singh and colleagues analysed roughly two million pairwise battles across 243 models from 42 providers over sixteen months of a public human-preference leaderboard. They reported undisclosed private testing under which a provider could evaluate many variants and publish only the best — one provider tested 27 private variants before releasing one publicly — that up to 26.5% of prompts were duplicates or near-duplicates, and that fine-tuning on arena-style prompts raised an arena win rate by 112% with no gain, or a slight loss, on general benchmarks. The last result is the important one: it separates performing well on the instrument from performing well on the thing the instrument was meant to stand for.
- Knowledge
- Construct validity in AI evaluation
- The difficulty of evaluating machine learning systems
- Benchmark saturation
- An audit finds the most-cited preference leaderboard is optimised rather than merely measured
- Connections
- The difficulty of evaluating machine learning systems
- Construct validity in AI evaluation
- Benchmark saturation
- CASP — Critical Assessment of Structure Prediction
- Construct validity in AI evaluation
- Benchmark saturation
- METR
- European Union Artificial Intelligence Act
- A human-parity claim in translation fails under document-level evaluation
- A randomised trial finds experienced developers slower with AI assistance
- An audit finds the most-cited preference leaderboard is optimised rather than merely measured
- A systematic review of 445 benchmarks reports pervasive validity failures
- The Commission proposes deferring the high-risk obligations
- The AI Index records both fast benchmark movement and doubts about benchmarks
- The difficulty of evaluating machine learning systems
- A human-parity claim in translation fails under document-level evaluation
- An audit finds the most-cited preference leaderboard is optimised rather than merely measured
- A systematic review of 445 benchmarks reports pervasive validity failures
- The difficulty of evaluating machine learning systems
- Measured adoption of AI by firms
- Fei-Fei Li
- GLUE is saturated and SuperGLUE is built to replace it
- An audit finds the most-cited preference leaderboard is optimised rather than merely measured
- The AI Index records both fast benchmark movement and doubts about benchmarks
- Evidence
- A systematic study of roughly two million pairwise battles across 243 models from 42 providers over sixteen months (January 2024 – April 2025) of a leading public human-preference leaderboard found undisclosed private testing allowing providers to evaluate multiple variants and publish only the best — one provider tested 27 private variants before a public release — that up to 26.5% of prompts were duplicates or near-duplicates, that proprietary models received a disproportionate share of evaluation data, and that fine-tuning on arena-style prompts raised an arena win rate by 112% with no improvement, or a slight decline, on general benchmarks such as MMLU. V55 verification basis: the paper was not fetched, but search results carried the author list, the arXiv identifier 2504.20879, and each of these figures with attribution to this paper.
Observed changes · 0
No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.
Actions
/atlas?object=EVT_LEADERBOARD_ILLUSION_2025&experience=EVT_LEADERBOARD_ILLUSION_2025