Skip to content
RichseenAtlasAtlasSign in

The AI Index records both fast benchmark movement and doubts about benchmarks

node

Stanford HAI's 2026 AI Index reported resolution on SWE-bench Verified rising from around 60% to near 100% within a year, accuracy on the OSWorld computer-use benchmark rising from roughly 12% to 66.3%, and a thirty-percentage-point gain on Humanity's Last Exam — while stating that benchmarks face growing questions about their reliability, and recording that robots succeed at only about 12% of real household tasks. Publishing the rate of progress and the doubt about how it is measured in the same document is the correct posture, and it is the posture this Journey takes.

A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Read in · 1

Evidence · 1
Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 4
Assembled narrative · 1

Assembled from 44 blocks · 1 evidence · 33 related

  1. Story
  2. Stanford HAI's 2026 AI Index reported resolution on SWE-bench Verified rising from around 60% to near 100% within a year, accuracy on the OSWorld computer-use benchmark rising from roughly 12% to 66.3%, and a thirty-percentage-point gain on Humanity's Last Exam — while stating that benchmarks face growing questions about their reliability, and recording that robots succeed at only about 12% of real household tasks. Publishing the rate of progress and the doubt about how it is measured in the same document is the correct posture, and it is the posture this Journey takes.
  3. Knowledge
  4. Machine-generated code
  5. The difficulty of evaluating machine learning systems
  6. Benchmark saturation
  7. Fei-Fei Li
  8. The AI Index records both fast benchmark movement and doubts about benchmarks
  9. Connections
  10. Benchmark saturation
  11. The difficulty of evaluating machine learning systems
  12. Fei-Fei Li
  13. Machine-generated code
  14. The difficulty of evaluating machine learning systems
  15. Measured adoption of AI by firms
  16. Fei-Fei Li
  17. GLUE is saturated and SuperGLUE is built to replace it
  18. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
  19. The AI Index records both fast benchmark movement and doubts about benchmarks
  20. CASP — Critical Assessment of Structure Prediction
  21. Construct validity in AI evaluation
  22. Benchmark saturation
  23. METR
  24. European Union Artificial Intelligence Act
  25. A human-parity claim in translation fails under document-level evaluation
  26. A randomised trial finds experienced developers slower with AI assistance
  27. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
  28. A systematic review of 445 benchmarks reports pervasive validity failures
  29. The Commission proposes deferring the high-risk obligations
  30. The AI Index records both fast benchmark movement and doubts about benchmarks
  31. ImageNet
  32. Benchmark saturation
  33. ImageNet is assembled and released
  34. The AI Index records both fast benchmark movement and doubts about benchmarks
  35. Labour-market effects of machine learning
  36. The transformer architecture
  37. METR
  38. A randomised trial finds experienced developers slower with AI assistance
  39. Payroll data shows a relative decline in entry-level employment in AI-exposed occupations
  40. The second International AI Safety Report reports mixed labour findings
  41. The AI Index records both fast benchmark movement and doubts about benchmarks
  42. Introducing SWE-bench Verified
  43. Evidence
  44. AI benchmark scores rose rapidly across language, reasoning, coding and mathematics in 2025 while benchmarks themselves faced growing questions about reliability: resolution on SWE-bench Verified rose from about 60% to near 100% within a year; accuracy on OSWorld rose from roughly 12% to 66.3%; frontier models gained about thirty percentage points in one year on Humanity's Last Exam; and robots succeed at only about 12% of real household tasks. V55 verification basis: the report was not fetched, but these figures were returned by a search restricted to hai.stanford.edu and are attributed there to the Index itself. Model-by-model scores, leaderboard ratings and adoption percentages that appeared alongside them in retrieved synthesis were NOT separately confirmed and are excluded from this pack.
Close the narrative
Observed changes · 0

No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

Actions

Read the assembled narrativeContinue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

/atlas?object=EVT_AI_INDEX_2026&experience=EVT_AI_INDEX_2026