Skip to content
RichseenAtlasAtlasSign in

The AI Index records both fast benchmark movement and doubts about benchmarks

node

Stanford HAI's 2026 AI Index reported resolution on SWE-bench Verified rising from around 60% to near 100% within a year, accuracy on the OSWorld computer-use benchmark rising from roughly 12% to 66.3%, and a thirty-percentage-point gain on Humanity's Last Exam — while stating that benchmarks face growing questions about their reliability, and recording that robots succeed at only about 12% of real household tasks. Publishing the rate of progress and the doubt about how it is measured in the same document is the correct posture, and it is the posture this Journey takes.

A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Read in · 1

Evidence · 1
Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 4
Assembled narrative · 1

An assembled narrative is available for this object.

Observed changes · 0

No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

Actions

Read the assembled narrativeContinue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

/atlas?object=EVT_AI_INDEX_2026