Atlas — the interactive world
Earth · Watch · All time · 60 of 198 subjects drawn · editorial corpus and accepted observations · official source connected
Watch
Close ✕An audit finds the most-cited preference leaderboard is optimised rather than merely measured
2025-04 · Recorded · science
The source states this happened, and its date is not in the future. The word the source used for it is observed.
Singh and colleagues analysed roughly two million pairwise battles across 243 models from 42 providers over sixteen months of a public human-preference leaderboard. They reported undisclosed private testing under which a provider could evaluate many variants and publish only the best — one provider tested 27 private variants before releasing one publicly — that up to 26.5% of prompts were duplicates or near-duplicates, and that fine-tuning on arena-style prompts raised an arena win rate by 112% with no gain, or a slight loss, on general benchmarks. The last result is the important one: it separates performing well on the instrument from performing well on the thing the instrument was meant to stand for.
- Source
- NeurIPS 2025; Cohere Labs and collaborating institutions · Singh, S., Nan, Y., Wang, A., D'souza, D., Kapoor, S., Üstün, A., Koyejo, S., Deng, Y., Longpre, S., Smith, N., Ermis, B., Fadaee, M. and Hooker, S., arXiv:2504.20879; NeurIPS 2025
- Standing
- This record: verified · publisher: verified
- How it arrived
- Editorial — A dated record from the Richseen Atlas editorial corpus, written and cited by a curator. Not captured from anywhere, and not live.
- Where
- At a canonical Atlas object, placed and checked by a curator.
- Subject
- PHENOMENON_EVALUATION_DIFFICULTY · CONCEPT_CONSTRUCT_VALIDITY · PHENOMENON_BENCHMARK_SATURATION
Sources polled — usgs: newest 2026-09-10 · nws: newest 2026-09-10 · jpl-cneos: states no publication time, last reached 2026-09-10 · Official source connected · Basemap: Natural Earth (public domain) — marks are drawn only from sourced coordinates