Skip to content
RichseenAtlasAtlasSign in

The 2026 AI Index Report — Technical Performance chapter and takeaways

evidence

AI benchmark scores rose rapidly across language, reasoning, coding and mathematics in 2025 while benchmarks themselves faced growing questions about reliability: resolution on SWE-bench Verified rose from about 60% to near 100% within a year; accuracy on OSWorld rose from roughly 12% to 66.3%; frontier models gained about thirty percentage points in one year on Humanity's Last Exam; and robots succeed at only about 12% of real household tasks. V55 verification basis: the report was not fetched, but these figures were returned by a search restricted to hai.stanford.edu and are attributed there to the Index itself. Model-by-model scores, leaderboard ratings and adoption percentages that appeared alongside them in retrieved synthesis were NOT separately confirmed and are excluded from this pack.

A evidence is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Evidence · 0

This object cites no evidence.

Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 0
    Assembled narrative · 0

    This object does not clear the publishing floor: an assembled narrative needs a description and at least one cited piece of evidence.

    Observed changes · 0

    No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

    Actions

    Read the assembled narrativeThis object has no public Experience: one needs a description and at least one cited piece of evidence.Continue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

    /atlas?object=EVIDENCE_EV_AI_INDEX_2026