GLUE is saturated and SuperGLUE is built to replace it
The GLUE language-understanding suite was released in 2018 and within about a year pretrained transformer models had passed its non-expert human baseline. SuperGLUE was constructed in 2019 explicitly as a harder successor, opening with a gap of roughly eighteen points between the best available model and the human baseline — and closed similarly quickly. The two-year cycle established here has repeated at every subsequent level of difficulty and is the origin of the pattern the Journey later examines: an instrument built to discriminate, answered, and replaced faster than the interpretation of its scores could be established.
Historical — it happened, and the record is settled.
Subjects
- Benchmark saturation · phenomenon · not located
- The transformer architecture · technology · not located
Evidence
Verified — Every field is supported by a cited authoritative source.
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
A systematic review of 445 large-language-model benchmarks, conducted by about forty researchers led from the Oxford Internet Institute, found that roughly half claim to measure abstract phenomena such as reasoning or harmlessness without defining them or specifying how they are measured, that 27% rely on convenience sampling, and that only 16% use uncertainty estimates or statistical tests when comparing model performance; the paper offers eight recommendations for benchmark construction. V55 verification basis: the paper was not fetched, but search results carried the 445-benchmark count, the three statistics, the arXiv identifier 2511.04703, the NeurIPS 2025 venue and the OII lead and attributed them to this work and to the Institute's own announcement.
NeurIPS 2025 proceedings; Oxford Internet Institute, University of Oxford · 2025-11 · Verified
Bean, A. M. et al., arXiv:2511.04703; NeurIPS 2025, San Diego, 2–7 December 2025; Oxford Internet Institute announcement, "Study identifies weaknesses in how AI systems are evaluated"; Oxford University Research Archive record
Recorded on Benchmark saturation
- The transformer is presented at NeurIPS
- GLUE is saturated and SuperGLUE is built to replace it
- An audit finds the most-cited preference leaderboard is optimised rather than merely measured
- The AI Index records both fast benchmark movement and doubts about benchmarks
Related
Related because they share a subject in the Atlas — never because the text looks similar.