Skip to content
RichseenAtlasAtlasSign in

Machine-generated code

technology

Writing and modifying software with model assistance — the most heavily instrumented deployment of language models, and the one where measured benchmark movement and measured human productivity most conspicuously fail to line up. On SWE-bench Verified, a 500-task set of real GitHub issues human-validated by 93 developers and released with OpenAI in August 2024, the 2026 AI Index reports resolution rates rising from around 60% to near 100% within a single year. Against that, a randomised controlled trial published by METR in July 2025 found that sixteen experienced open-source developers working on 246 issues in repositories they already knew took 19% longer with early-2025 AI tools than without, while estimating afterwards that the tools had made them about 20% faster. Both results are real measurements of different things: one of whether a patch passes a held-out test suite, one of how long a familiar expert takes on a familiar codebase. Neither licenses a claim about software employment, and the gap between the perceived and measured speed-up in the trial is itself a finding worth more than either number.

A technology is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Read in · 1

Evidence · 3
Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 2
Assembled narrative · 1

Assembled from 33 blocks · 3 evidence · 27 related

  1. Story
  2. Writing and modifying software with model assistance — the most heavily instrumented deployment of language models, and the one where measured benchmark movement and measured human productivity most conspicuously fail to line up. On SWE-bench Verified, a 500-task set of real GitHub issues human-validated by 93 developers and released with OpenAI in August 2024, the 2026 AI Index reports resolution rates rising from around 60% to near 100% within a single year. Against that, a randomised controlled trial published by METR in July 2025 found that sixteen experienced open-source developers working on 246 issues in repositories they already knew took 19% longer with early-2025 AI tools than without, while estimating afterwards that the tools had made them about 20% faster. Both results are real measurements of different things: one of whether a patch passes a held-out test suite, one of how long a familiar expert takes on a familiar codebase. Neither licenses a claim about software employment, and the gap between the perceived and measured speed-up in the trial is itself a finding worth more than either number.
  3. Knowledge
  4. The transformer architecture
  5. Machine-generated code
  6. Labour-market effects of machine learning
  7. Connections
  8. Labour-market effects of machine learning
  9. The transformer architecture
  10. METR
  11. A randomised trial finds experienced developers slower with AI assistance
  12. Payroll data shows a relative decline in entry-level employment in AI-exposed occupations
  13. The second International AI Safety Report reports mixed labour findings
  14. The AI Index records both fast benchmark movement and doubts about benchmarks
  15. Introducing SWE-bench Verified
  16. Machine translation
  17. Machine-generated code
  18. Measured adoption of AI by firms
  19. Payroll data shows a relative decline in entry-level employment in AI-exposed occupations
  20. A whole-labour-market analysis finds no economy-wide disruption yet
  21. The second International AI Safety Report reports mixed labour findings
  22. Radiologist workforce projections and employment outlook, against the 2016 prediction
  23. The AI accelerator
  24. Deep learning
  25. Machine translation
  26. Machine-generated code
  27. Growth in frontier training compute
  28. The transformer is presented at NeurIPS
  29. GLUE is saturated and SuperGLUE is built to replace it
  30. Evidence
  31. SWE-bench Verified is a 500-instance human-validated subset of SWE-bench, released in August 2024, curated with 93 software developers to remove instances with incorrect grading of correct solutions, under-specified problem statements and overly specific unit tests; each task is a real GitHub issue from one of twelve open-source Python repositories, and a model must produce a patch that passes the repository's tests. V55 verification basis: neither page was fetched, but search results carried the 500-instance count, the August 2024 release, the 93 annotators, the twelve-repository composition and the stated purpose of the human validation, attributing them to OpenAI's release page and to swebench.com.
  32. In a randomised trial, sixteen experienced open-source developers working on 246 real issues in repositories they knew well — about five years of prior experience with the projects on average — took 19% longer to complete tasks when allowed to use AI tools available between February and June 2025, while estimating afterwards that the tools had made them about 20% faster; METR has since labelled the result historical, as not necessarily reflecting current tools or workflows. Separately, METR reports a task time-horizon metric — the human-measured duration of task a model completes autonomously with 50% reliability — as having doubled approximately every seven months over 2019–2025, measured across 170 tasks against over 800 human baselines with thirteen frontier models. V55 verification basis: metr.org was not fetched. The trial figures, the arXiv identifier 2507.09089, the task and participant counts and METR's own historical caveat were carried in search results attributing them to METR; the seven-month doubling, the 170-task suite and the 800-plus human baselines were carried with attribution to METR's March 2025 post. Any specific model's time horizon, and reports of a faster doubling in the most recent period, come from secondary summaries and are NOT asserted in this pack.
  33. AI benchmark scores rose rapidly across language, reasoning, coding and mathematics in 2025 while benchmarks themselves faced growing questions about reliability: resolution on SWE-bench Verified rose from about 60% to near 100% within a year; accuracy on OSWorld rose from roughly 12% to 66.3%; frontier models gained about thirty percentage points in one year on Humanity's Last Exam; and robots succeed at only about 12% of real household tasks. V55 verification basis: the report was not fetched, but these figures were returned by a search restricted to hai.stanford.edu and are attributed there to the Index itself. Model-by-model scores, leaderboard ratings and adoption percentages that appeared alongside them in retrieved synthesis were NOT separately confirmed and are excluded from this pack.
Close the narrative
Observed changes · 0

No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

Actions

Read the assembled narrativeContinue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

/atlas?object=DOMAIN_CODE_GENERATION&experience=DOMAIN_CODE_GENERATION