An audit finds the most-cited preference leaderboard is optimised rather than merely measured
node
Singh and colleagues analysed roughly two million pairwise battles across 243 models from 42 providers over sixteen months of a public human-preference leaderboard. They reported undisclosed private testing under which a provider could evaluate many variants and publish only the best — one provider tested 27 private variants before releasing one publicly — that up to 26.5% of prompts were duplicates or near-duplicates, and that fine-tuning on arena-style prompts raised an arena win rate by 112% with no gain, or a slight loss, on general benchmarks. The last result is the important one: it separates performing well on the instrument from performing well on the thing the instrument was meant to stand for.
A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.
Read in · 1
Evidence · 1
Timeline
No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.
Connections · 3
- The difficulty of evaluating machine learning systemsevidence_for
- Construct validity in AI evaluationconcerns
- Benchmark saturationconcerns
Assembled narrative · 1
An assembled narrative is available for this object.
Observed changes · 0
No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.
Actions
/atlas?object=EVT_LEADERBOARD_ILLUSION_2025