A randomised trial finds experienced developers slower with AI assistance
node
METR published a randomised controlled trial in which sixteen experienced open-source developers worked on 246 real issues in repositories they knew well — about five years of prior experience with the projects on average — using AI tools available between February and June 2025. They took 19% longer with the tools than without, and estimated afterwards that the tools had made them about 20% faster. The gap between the measured and the perceived effect is the finding that generalises furthest, because self-reported productivity gains are the evidence base for a very large share of public claims about this technology.
A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.
Read in · 1
Timeline
No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.
Connections · 3
- Machine-generated codeconcerns
- METRconcerns
- The difficulty of evaluating machine learning systemsconcerns
Assembled narrative · 1
Assembled from 35 blocks · 1 evidence · 28 related
- Story
- METR published a randomised controlled trial in which sixteen experienced open-source developers worked on 246 real issues in repositories they knew well — about five years of prior experience with the projects on average — using AI tools available between February and June 2025. They took 19% longer with the tools than without, and estimated afterwards that the tools had made them about 20% faster. The gap between the measured and the perceived effect is the finding that generalises furthest, because self-reported productivity gains are the evidence base for a very large share of public claims about this technology.
- Knowledge
- Machine-generated code
- The difficulty of evaluating machine learning systems
- METR
- A randomised trial finds experienced developers slower with AI assistance
- Connections
- Machine-generated code
- METR
- The difficulty of evaluating machine learning systems
- Labour-market effects of machine learning
- The transformer architecture
- METR
- A randomised trial finds experienced developers slower with AI assistance
- Payroll data shows a relative decline in entry-level employment in AI-exposed occupations
- The second International AI Safety Report reports mixed labour findings
- The AI Index records both fast benchmark movement and doubts about benchmarks
- Introducing SWE-bench Verified
- Machine-generated code
- The difficulty of evaluating machine learning systems
- A randomised trial finds experienced developers slower with AI assistance
- CASP — Critical Assessment of Structure Prediction
- Construct validity in AI evaluation
- Benchmark saturation
- METR
- European Union Artificial Intelligence Act
- A human-parity claim in translation fails under document-level evaluation
- A randomised trial finds experienced developers slower with AI assistance
- An audit finds the most-cited preference leaderboard is optimised rather than merely measured
- A systematic review of 445 benchmarks reports pervasive validity failures
- The Commission proposes deferring the high-risk obligations
- The AI Index records both fast benchmark movement and doubts about benchmarks
- Evidence
- In a randomised trial, sixteen experienced open-source developers working on 246 real issues in repositories they knew well — about five years of prior experience with the projects on average — took 19% longer to complete tasks when allowed to use AI tools available between February and June 2025, while estimating afterwards that the tools had made them about 20% faster; METR has since labelled the result historical, as not necessarily reflecting current tools or workflows. Separately, METR reports a task time-horizon metric — the human-measured duration of task a model completes autonomously with 50% reliability — as having doubled approximately every seven months over 2019–2025, measured across 170 tasks against over 800 human baselines with thirteen frontier models. V55 verification basis: metr.org was not fetched. The trial figures, the arXiv identifier 2507.09089, the task and participant counts and METR's own historical caveat were carried in search results attributing them to METR; the seven-month doubling, the 170-task suite and the 800-plus human baselines were carried with attribution to METR's March 2025 post. Any specific model's time horizon, and reports of a faster doubling in the most recent period, come from secondary summaries and are NOT asserted in this pack.
Observed changes · 0
No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.
Actions
/atlas?object=EVT_METR_DEVELOPER_RCT_2025&experience=EVT_METR_DEVELOPER_RCT_2025