Skip to content
RichseenAtlasAtlasSign in

Benchmark saturation

node

The pattern in which a benchmark built to be hard is answered, then ceases to discriminate, then is replaced — and the interval between those steps has been shortening. GLUE, released in 2018 as a general language-understanding suite, was surpassed on its non-expert human baseline within about a year; SuperGLUE was constructed in 2019 explicitly because of that, opening with a roughly eighteen-point gap between the best model and the human baseline, and closed similarly. The same shape repeats at higher difficulty: the Stanford AI Index for 2026 reports resolution on SWE-bench Verified rising from around 60% to near 100% in a single year, accuracy on the OSWorld computer-use benchmark rising from roughly 12% to 66.3%, and a thirty-percentage-point gain in one year on Humanity's Last Exam, a set written by subject-matter experts specifically to resist this. Saturation measures two things at once and they must not be collapsed: systems improving, and the instrument losing its ability to tell systems apart. The same report notes that robots succeed at only about 12% of real household tasks, which is a useful check on how far a saturated text benchmark generalises.

A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Read in · 1

Evidence · 3
Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 1
Assembled narrative · 1

Assembled from 27 blocks · 3 evidence · 23 related

  1. Story
  2. The pattern in which a benchmark built to be hard is answered, then ceases to discriminate, then is replaced — and the interval between those steps has been shortening. GLUE, released in 2018 as a general language-understanding suite, was surpassed on its non-expert human baseline within about a year; SuperGLUE was constructed in 2019 explicitly because of that, opening with a roughly eighteen-point gap between the best model and the human baseline, and closed similarly. The same shape repeats at higher difficulty: the Stanford AI Index for 2026 reports resolution on SWE-bench Verified rising from around 60% to near 100% in a single year, accuracy on the OSWorld computer-use benchmark rising from roughly 12% to 66.3%, and a thirty-percentage-point gain in one year on Humanity's Last Exam, a set written by subject-matter experts specifically to resist this. Saturation measures two things at once and they must not be collapsed: systems improving, and the instrument losing its ability to tell systems apart. The same report notes that robots succeed at only about 12% of real household tasks, which is a useful check on how far a saturated text benchmark generalises.
  3. Knowledge
  4. The difficulty of evaluating machine learning systems
  5. Benchmark saturation
  6. Connections
  7. The difficulty of evaluating machine learning systems
  8. Measured adoption of AI by firms
  9. Fei-Fei Li
  10. GLUE is saturated and SuperGLUE is built to replace it
  11. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
  12. The AI Index records both fast benchmark movement and doubts about benchmarks
  13. CASP — Critical Assessment of Structure Prediction
  14. Construct validity in AI evaluation
  15. Benchmark saturation
  16. METR
  17. European Union Artificial Intelligence Act
  18. A human-parity claim in translation fails under document-level evaluation
  19. A randomised trial finds experienced developers slower with AI assistance
  20. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
  21. A systematic review of 445 benchmarks reports pervasive validity failures
  22. The Commission proposes deferring the high-risk obligations
  23. The AI Index records both fast benchmark movement and doubts about benchmarks
  24. Evidence
  25. AI benchmark scores rose rapidly across language, reasoning, coding and mathematics in 2025 while benchmarks themselves faced growing questions about reliability: resolution on SWE-bench Verified rose from about 60% to near 100% within a year; accuracy on OSWorld rose from roughly 12% to 66.3%; frontier models gained about thirty percentage points in one year on Humanity's Last Exam; and robots succeed at only about 12% of real household tasks. V55 verification basis: the report was not fetched, but these figures were returned by a search restricted to hai.stanford.edu and are attributed there to the Index itself. Model-by-model scores, leaderboard ratings and adoption percentages that appeared alongside them in retrieved synthesis were NOT separately confirmed and are excluded from this pack.
  26. SWE-bench Verified is a 500-instance human-validated subset of SWE-bench, released in August 2024, curated with 93 software developers to remove instances with incorrect grading of correct solutions, under-specified problem statements and overly specific unit tests; each task is a real GitHub issue from one of twelve open-source Python repositories, and a model must produce a patch that passes the repository's tests. V55 verification basis: neither page was fetched, but search results carried the 500-instance count, the August 2024 release, the 93 annotators, the twelve-repository composition and the stated purpose of the human validation, attributing them to OpenAI's release page and to swebench.com.
  27. A systematic review of 445 large-language-model benchmarks, conducted by about forty researchers led from the Oxford Internet Institute, found that roughly half claim to measure abstract phenomena such as reasoning or harmlessness without defining them or specifying how they are measured, that 27% rely on convenience sampling, and that only 16% use uncertainty estimates or statistical tests when comparing model performance; the paper offers eight recommendations for benchmark construction. V55 verification basis: the paper was not fetched, but search results carried the 445-benchmark count, the three statistics, the arXiv identifier 2511.04703, the NeurIPS 2025 venue and the OII lead and attributed them to this work and to the Institute's own announcement.
Close the narrative
Observed changes · 0

No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

Actions

Read the assembled narrativeContinue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

/atlas?object=PHENOMENON_BENCHMARK_SATURATION&experience=PHENOMENON_BENCHMARK_SATURATION