The Leaderboard Illusion
evidence
A systematic study of roughly two million pairwise battles across 243 models from 42 providers over sixteen months (January 2024 – April 2025) of a leading public human-preference leaderboard found undisclosed private testing allowing providers to evaluate multiple variants and publish only the best — one provider tested 27 private variants before a public release — that up to 26.5% of prompts were duplicates or near-duplicates, that proprietary models received a disproportionate share of evaluation data, and that fine-tuning on arena-style prompts raised an arena win rate by 112% with no improvement, or a slight decline, on general benchmarks such as MMLU. V55 verification basis: the paper was not fetched, but search results carried the author list, the arXiv identifier 2504.20879, and each of these figures with attribution to this paper.
A evidence is not a place. Drawing it on a map would assert something about the world that no stored fact supports.
Evidence · 0
This object cites no evidence.
Timeline
No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.
Connections · 0
Assembled narrative · 0
This object does not clear the publishing floor: an assembled narrative needs a description and at least one cited piece of evidence.
Observed changes · 0
No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.
Actions
/atlas?object=EVIDENCE_EV_LEADERBOARD_ILLUSION_2025