The AI Index records both fast benchmark movement and doubts about benchmarks
Stanford HAI's 2026 AI Index reported resolution on SWE-bench Verified rising from around 60% to near 100% within a year, accuracy on the OSWorld computer-use benchmark rising from roughly 12% to 66.3%, and a thirty-percentage-point gain on Humanity's Last Exam — while stating that benchmarks face growing questions about their reliability, and recording that robots succeed at only about 12% of real household tasks. Publishing the rate of progress and the doubt about how it is measured in the same document is the correct posture, and it is the posture this Journey takes.
Historical — it happened, and the record is settled.
Subjects
- Benchmark saturation · phenomenon · not located
- The difficulty of evaluating machine learning systems · phenomenon · not located
- Fei-Fei Li · person · United States · not located
- Machine-generated code · technology · not located
Evidence
Partly verified — Core facts are sourced; some optional detail is deliberately absent.
The 2026 AI Index Report — Technical Performance chapter and takeaways
AI benchmark scores rose rapidly across language, reasoning, coding and mathematics in 2025 while benchmarks themselves faced growing questions about reliability: resolution on SWE-bench Verified rose from about 60% to near 100% within a year; accuracy on OSWorld rose from roughly 12% to 66.3%; frontier models gained about thirty percentage points in one year on Humanity's Last Exam; and robots succeed at only about 12% of real household tasks. V55 verification basis: the report was not fetched, but these figures were returned by a search restricted to hai.stanford.edu and are attributed there to the Index itself. Model-by-model scores, leaderboard ratings and adoption percentages that appeared alongside them in retrieved synthesis were NOT separately confirmed and are excluded from this pack.
Stanford Institute for Human-Centered Artificial Intelligence (Stanford HAI) · 2026 · Partly verified
https://hai.stanford.edu/ai-index/2026-ai-index-report and its Technical Performance chapter; "Inside the AI Index: 12 Takeaways from the 2026 Report"
Recorded on Benchmark saturation
- ImageNet is assembled and released
- GLUE is saturated and SuperGLUE is built to replace it
- A human-parity claim in translation fails under document-level evaluation
- Payroll data shows a relative decline in entry-level employment in AI-exposed occupations
- An audit finds the most-cited preference leaderboard is optimised rather than merely measured
- A systematic review of 445 benchmarks reports pervasive validity failures
- A randomised trial finds experienced developers slower with AI assistance
- The Commission proposes deferring the high-risk obligations
- The second International AI Safety Report reports mixed labour findings
- The AI Index records both fast benchmark movement and doubts about benchmarks
Related
Related because they share a subject in the Atlas — never because the text looks similar.
- The second International AI Safety Report reports mixed labour findings · 2026-02
- The Commission proposes deferring the high-risk obligations · 2025-11-19
- A randomised trial finds experienced developers slower with AI assistance · 2025-07-10
- Payroll data shows a relative decline in entry-level employment in AI-exposed occupations · 2025
- An audit finds the most-cited preference leaderboard is optimised rather than merely measured · 2025-04
- A systematic review of 445 benchmarks reports pervasive validity failures · 2025-11
- GLUE is saturated and SuperGLUE is built to replace it · 2018–2019
- A human-parity claim in translation fails under document-level evaluation · 2018