Benchmark saturation
node
The pattern in which a benchmark built to be hard is answered, then ceases to discriminate, then is replaced — and the interval between those steps has been shortening. GLUE, released in 2018 as a general language-understanding suite, was surpassed on its non-expert human baseline within about a year; SuperGLUE was constructed in 2019 explicitly because of that, opening with a roughly eighteen-point gap between the best model and the human baseline, and closed similarly. The same shape repeats at higher difficulty: the Stanford AI Index for 2026 reports resolution on SWE-bench Verified rising from around 60% to near 100% in a single year, accuracy on the OSWorld computer-use benchmark rising from roughly 12% to 66.3%, and a thirty-percentage-point gain in one year on Humanity's Last Exam, a set written by subject-matter experts specifically to resist this. Saturation measures two things at once and they must not be collapsed: systems improving, and the instrument losing its ability to tell systems apart. The same report notes that robots succeed at only about 12% of real household tasks, which is a useful check on how far a saturated text benchmark generalises.
A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.
Read in · 1
Evidence · 3
Timeline
No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.
Connections · 1
Assembled narrative · 1
An assembled narrative is available for this object.
Observed changes · 0
No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.
Actions
/atlas?object=PHENOMENON_BENCHMARK_SATURATION