Machine-generated code
technology
Writing and modifying software with model assistance — the most heavily instrumented deployment of language models, and the one where measured benchmark movement and measured human productivity most conspicuously fail to line up. On SWE-bench Verified, a 500-task set of real GitHub issues human-validated by 93 developers and released with OpenAI in August 2024, the 2026 AI Index reports resolution rates rising from around 60% to near 100% within a single year. Against that, a randomised controlled trial published by METR in July 2025 found that sixteen experienced open-source developers working on 246 issues in repositories they already knew took 19% longer with early-2025 AI tools than without, while estimating afterwards that the tools had made them about 20% faster. Both results are real measurements of different things: one of whether a patch passes a held-out test suite, one of how long a familiar expert takes on a familiar codebase. Neither licenses a claim about software employment, and the gap between the perceived and measured speed-up in the trial is itself a finding worth more than either number.
A technology is not a place. Drawing it on a map would assert something about the world that no stored fact supports.
Read in · 1
Evidence · 3
Timeline
No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.
Connections · 2
Assembled narrative · 1
Assembled from 33 blocks · 3 evidence · 27 related
- Story
- Writing and modifying software with model assistance — the most heavily instrumented deployment of language models, and the one where measured benchmark movement and measured human productivity most conspicuously fail to line up. On SWE-bench Verified, a 500-task set of real GitHub issues human-validated by 93 developers and released with OpenAI in August 2024, the 2026 AI Index reports resolution rates rising from around 60% to near 100% within a single year. Against that, a randomised controlled trial published by METR in July 2025 found that sixteen experienced open-source developers working on 246 issues in repositories they already knew took 19% longer with early-2025 AI tools than without, while estimating afterwards that the tools had made them about 20% faster. Both results are real measurements of different things: one of whether a patch passes a held-out test suite, one of how long a familiar expert takes on a familiar codebase. Neither licenses a claim about software employment, and the gap between the perceived and measured speed-up in the trial is itself a finding worth more than either number.
- Knowledge
- The transformer architecture
- Machine-generated code
- Labour-market effects of machine learning
- Connections
- Labour-market effects of machine learning
- The transformer architecture
- METR
- A randomised trial finds experienced developers slower with AI assistance
- Payroll data shows a relative decline in entry-level employment in AI-exposed occupations
- The second International AI Safety Report reports mixed labour findings
- The AI Index records both fast benchmark movement and doubts about benchmarks
- Introducing SWE-bench Verified
- Machine translation
- Machine-generated code
- Measured adoption of AI by firms
- Payroll data shows a relative decline in entry-level employment in AI-exposed occupations
- A whole-labour-market analysis finds no economy-wide disruption yet
- The second International AI Safety Report reports mixed labour findings
- Radiologist workforce projections and employment outlook, against the 2016 prediction
- The AI accelerator
- Deep learning
- Machine translation
- Machine-generated code
- Growth in frontier training compute
- The transformer is presented at NeurIPS
- GLUE is saturated and SuperGLUE is built to replace it
- Evidence
- SWE-bench Verified is a 500-instance human-validated subset of SWE-bench, released in August 2024, curated with 93 software developers to remove instances with incorrect grading of correct solutions, under-specified problem statements and overly specific unit tests; each task is a real GitHub issue from one of twelve open-source Python repositories, and a model must produce a patch that passes the repository's tests. V55 verification basis: neither page was fetched, but search results carried the 500-instance count, the August 2024 release, the 93 annotators, the twelve-repository composition and the stated purpose of the human validation, attributing them to OpenAI's release page and to swebench.com.
- In a randomised trial, sixteen experienced open-source developers working on 246 real issues in repositories they knew well — about five years of prior experience with the projects on average — took 19% longer to complete tasks when allowed to use AI tools available between February and June 2025, while estimating afterwards that the tools had made them about 20% faster; METR has since labelled the result historical, as not necessarily reflecting current tools or workflows. Separately, METR reports a task time-horizon metric — the human-measured duration of task a model completes autonomously with 50% reliability — as having doubled approximately every seven months over 2019–2025, measured across 170 tasks against over 800 human baselines with thirteen frontier models. V55 verification basis: metr.org was not fetched. The trial figures, the arXiv identifier 2507.09089, the task and participant counts and METR's own historical caveat were carried in search results attributing them to METR; the seven-month doubling, the 170-task suite and the 800-plus human baselines were carried with attribution to METR's March 2025 post. Any specific model's time horizon, and reports of a faster doubling in the most recent period, come from secondary summaries and are NOT asserted in this pack.
- AI benchmark scores rose rapidly across language, reasoning, coding and mathematics in 2025 while benchmarks themselves faced growing questions about reliability: resolution on SWE-bench Verified rose from about 60% to near 100% within a year; accuracy on OSWorld rose from roughly 12% to 66.3%; frontier models gained about thirty percentage points in one year on Humanity's Last Exam; and robots succeed at only about 12% of real household tasks. V55 verification basis: the report was not fetched, but these figures were returned by a search restricted to hai.stanford.edu and are attributed there to the Index itself. Model-by-model scores, leaderboard ratings and adoption percentages that appeared alongside them in retrieved synthesis were NOT separately confirmed and are excluded from this pack.
Observed changes · 0
No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.
Actions
/atlas?object=DOMAIN_CODE_GENERATION&experience=DOMAIN_CODE_GENERATION