Skip to content
RichseenAtlasAtlasSign in

Introducing SWE-bench Verified

evidence

SWE-bench Verified is a 500-instance human-validated subset of SWE-bench, released in August 2024, curated with 93 software developers to remove instances with incorrect grading of correct solutions, under-specified problem statements and overly specific unit tests; each task is a real GitHub issue from one of twelve open-source Python repositories, and a model must produce a patch that passes the repository's tests. V55 verification basis: neither page was fetched, but search results carried the 500-instance count, the August 2024 release, the 93 annotators, the twelve-repository composition and the stated purpose of the human validation, attributing them to OpenAI's release page and to swebench.com.

A evidence is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Evidence · 0

This object cites no evidence.

Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 1
Assembled narrative · 0

This object does not clear the publishing floor: an assembled narrative needs a description and at least one cited piece of evidence.

Observed changes · 0

No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

Actions

Read the assembled narrativeThis object has no public Experience: one needs a description and at least one cited piece of evidence.Continue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

/atlas?object=EVIDENCE_EV_SWE_BENCH_VERIFIED