Skip to content
RichseenAtlasAtlasSign in

A randomised trial finds experienced developers slower with AI assistance

node

METR published a randomised controlled trial in which sixteen experienced open-source developers worked on 246 real issues in repositories they knew well — about five years of prior experience with the projects on average — using AI tools available between February and June 2025. They took 19% longer with the tools than without, and estimated afterwards that the tools had made them about 20% faster. The gap between the measured and the perceived effect is the finding that generalises furthest, because self-reported productivity gains are the evidence base for a very large share of public claims about this technology.

A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Read in · 1

Evidence · 1
Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 3
Assembled narrative · 1

Assembled from 35 blocks · 1 evidence · 28 related

  1. Story
  2. METR published a randomised controlled trial in which sixteen experienced open-source developers worked on 246 real issues in repositories they knew well — about five years of prior experience with the projects on average — using AI tools available between February and June 2025. They took 19% longer with the tools than without, and estimated afterwards that the tools had made them about 20% faster. The gap between the measured and the perceived effect is the finding that generalises furthest, because self-reported productivity gains are the evidence base for a very large share of public claims about this technology.
  3. Knowledge
  4. Machine-generated code
  5. The difficulty of evaluating machine learning systems
  6. METR
  7. A randomised trial finds experienced developers slower with AI assistance
  8. Connections
  9. Machine-generated code
  10. METR
  11. The difficulty of evaluating machine learning systems
  12. Labour-market effects of machine learning
  13. The transformer architecture
  14. METR
  15. A randomised trial finds experienced developers slower with AI assistance
  16. Payroll data shows a relative decline in entry-level employment in AI-exposed occupations
  17. The second International AI Safety Report reports mixed labour findings
  18. The AI Index records both fast benchmark movement and doubts about benchmarks
  19. Introducing SWE-bench Verified
  20. Machine-generated code
  21. The difficulty of evaluating machine learning systems
  22. A randomised trial finds experienced developers slower with AI assistance
  23. CASP — Critical Assessment of Structure Prediction
  24. Construct validity in AI evaluation
  25. Benchmark saturation
  26. METR
  27. European Union Artificial Intelligence Act
  28. A human-parity claim in translation fails under document-level evaluation
  29. A randomised trial finds experienced developers slower with AI assistance
  30. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
  31. A systematic review of 445 benchmarks reports pervasive validity failures
  32. The Commission proposes deferring the high-risk obligations
  33. The AI Index records both fast benchmark movement and doubts about benchmarks
  34. Evidence
  35. In a randomised trial, sixteen experienced open-source developers working on 246 real issues in repositories they knew well — about five years of prior experience with the projects on average — took 19% longer to complete tasks when allowed to use AI tools available between February and June 2025, while estimating afterwards that the tools had made them about 20% faster; METR has since labelled the result historical, as not necessarily reflecting current tools or workflows. Separately, METR reports a task time-horizon metric — the human-measured duration of task a model completes autonomously with 50% reliability — as having doubled approximately every seven months over 2019–2025, measured across 170 tasks against over 800 human baselines with thirteen frontier models. V55 verification basis: metr.org was not fetched. The trial figures, the arXiv identifier 2507.09089, the task and participant counts and METR's own historical caveat were carried in search results attributing them to METR; the seven-month doubling, the 170-task suite and the 800-plus human baselines were carried with attribution to METR's March 2025 post. Any specific model's time horizon, and reports of a faster doubling in the most recent period, come from secondary summaries and are NOT asserted in this pack.
Close the narrative
Observed changes · 0

No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

Actions

Read the assembled narrativeContinue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

/atlas?object=EVT_METR_DEVELOPER_RCT_2025&experience=EVT_METR_DEVELOPER_RCT_2025