Skip to content
RichseenAtlasAtlasSign in

An audit finds the most-cited preference leaderboard is optimised rather than merely measured

node

Singh and colleagues analysed roughly two million pairwise battles across 243 models from 42 providers over sixteen months of a public human-preference leaderboard. They reported undisclosed private testing under which a provider could evaluate many variants and publish only the best — one provider tested 27 private variants before releasing one publicly — that up to 26.5% of prompts were duplicates or near-duplicates, and that fine-tuning on arena-style prompts raised an arena win rate by 112% with no gain, or a slight loss, on general benchmarks. The last result is the important one: it separates performing well on the instrument from performing well on the thing the instrument was meant to stand for.

A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Read in · 1

Evidence · 1
Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 3
Assembled narrative · 1

Assembled from 34 blocks · 1 evidence · 23 related

  1. Story
  2. Singh and colleagues analysed roughly two million pairwise battles across 243 models from 42 providers over sixteen months of a public human-preference leaderboard. They reported undisclosed private testing under which a provider could evaluate many variants and publish only the best — one provider tested 27 private variants before releasing one publicly — that up to 26.5% of prompts were duplicates or near-duplicates, and that fine-tuning on arena-style prompts raised an arena win rate by 112% with no gain, or a slight loss, on general benchmarks. The last result is the important one: it separates performing well on the instrument from performing well on the thing the instrument was meant to stand for.
  3. Knowledge
  4. Construct validity in AI evaluation
  5. The difficulty of evaluating machine learning systems
  6. Benchmark saturation
  7. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
  8. Connections
  9. The difficulty of evaluating machine learning systems
  10. Construct validity in AI evaluation
  11. Benchmark saturation
  12. CASP — Critical Assessment of Structure Prediction
  13. Construct validity in AI evaluation
  14. Benchmark saturation
  15. METR
  16. European Union Artificial Intelligence Act
  17. A human-parity claim in translation fails under document-level evaluation
  18. A randomised trial finds experienced developers slower with AI assistance
  19. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
  20. A systematic review of 445 benchmarks reports pervasive validity failures
  21. The Commission proposes deferring the high-risk obligations
  22. The AI Index records both fast benchmark movement and doubts about benchmarks
  23. The difficulty of evaluating machine learning systems
  24. A human-parity claim in translation fails under document-level evaluation
  25. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
  26. A systematic review of 445 benchmarks reports pervasive validity failures
  27. The difficulty of evaluating machine learning systems
  28. Measured adoption of AI by firms
  29. Fei-Fei Li
  30. GLUE is saturated and SuperGLUE is built to replace it
  31. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
  32. The AI Index records both fast benchmark movement and doubts about benchmarks
  33. Evidence
  34. A systematic study of roughly two million pairwise battles across 243 models from 42 providers over sixteen months (January 2024 – April 2025) of a leading public human-preference leaderboard found undisclosed private testing allowing providers to evaluate multiple variants and publish only the best — one provider tested 27 private variants before a public release — that up to 26.5% of prompts were duplicates or near-duplicates, that proprietary models received a disproportionate share of evaluation data, and that fine-tuning on arena-style prompts raised an arena win rate by 112% with no improvement, or a slight decline, on general benchmarks such as MMLU. V55 verification basis: the paper was not fetched, but search results carried the author list, the arXiv identifier 2504.20879, and each of these figures with attribution to this paper.
Close the narrative
Observed changes · 0

No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

Actions

Read the assembled narrativeContinue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

/atlas?object=EVT_LEADERBOARD_ILLUSION_2025&experience=EVT_LEADERBOARD_ILLUSION_2025