Skip to content
RichseenAtlasAtlasSign in

METR

node · earth

The non-profit research organisation, based in the United States, whose distinctive contribution has been to measure what models do in conditions closer to work than to examinations. Two of its outputs are load-bearing here. In July 2025 it published a randomised controlled trial in which sixteen experienced open-source developers completed 246 real issues in repositories they knew well, and were 19% slower with early-2025 AI assistance than without while believing they had been about 20% faster — a result METR itself has since flagged as historical, describing tools and workflows that have changed. Separately it publishes a task time-horizon metric: the length of task, measured by how long skilled humans take, that a model completes autonomously with 50% reliability, reported as doubling roughly every seven months over 2019–2025 with a faster doubling in the most recent period. The metric is the interesting object, not the extrapolation from it: it converts an unbounded question about capability into a quantity with a unit and a stated reliability threshold.

A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Read in · 1

Evidence · 1
Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 2
Assembled narrative · 1

Assembled from 31 blocks · 1 evidence · 28 related

  1. Story
  2. The non-profit research organisation, based in the United States, whose distinctive contribution has been to measure what models do in conditions closer to work than to examinations. Two of its outputs are load-bearing here. In July 2025 it published a randomised controlled trial in which sixteen experienced open-source developers completed 246 real issues in repositories they knew well, and were 19% slower with early-2025 AI assistance than without while believing they had been about 20% faster — a result METR itself has since flagged as historical, describing tools and workflows that have changed. Separately it publishes a task time-horizon metric: the length of task, measured by how long skilled humans take, that a model completes autonomously with 50% reliability, reported as doubling roughly every seven months over 2019–2025 with a faster doubling in the most recent period. The metric is the interesting object, not the extrapolation from it: it converts an unbounded question about capability into a quantity with a unit and a stated reliability threshold.
  3. Knowledge
  4. Machine-generated code
  5. The difficulty of evaluating machine learning systems
  6. METR
  7. Connections
  8. Machine-generated code
  9. The difficulty of evaluating machine learning systems
  10. A randomised trial finds experienced developers slower with AI assistance
  11. Labour-market effects of machine learning
  12. The transformer architecture
  13. METR
  14. A randomised trial finds experienced developers slower with AI assistance
  15. Payroll data shows a relative decline in entry-level employment in AI-exposed occupations
  16. The second International AI Safety Report reports mixed labour findings
  17. The AI Index records both fast benchmark movement and doubts about benchmarks
  18. Introducing SWE-bench Verified
  19. CASP — Critical Assessment of Structure Prediction
  20. Construct validity in AI evaluation
  21. Benchmark saturation
  22. METR
  23. European Union Artificial Intelligence Act
  24. A human-parity claim in translation fails under document-level evaluation
  25. A randomised trial finds experienced developers slower with AI assistance
  26. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
  27. A systematic review of 445 benchmarks reports pervasive validity failures
  28. The Commission proposes deferring the high-risk obligations
  29. The AI Index records both fast benchmark movement and doubts about benchmarks
  30. Evidence
  31. In a randomised trial, sixteen experienced open-source developers working on 246 real issues in repositories they knew well — about five years of prior experience with the projects on average — took 19% longer to complete tasks when allowed to use AI tools available between February and June 2025, while estimating afterwards that the tools had made them about 20% faster; METR has since labelled the result historical, as not necessarily reflecting current tools or workflows. Separately, METR reports a task time-horizon metric — the human-measured duration of task a model completes autonomously with 50% reliability — as having doubled approximately every seven months over 2019–2025, measured across 170 tasks against over 800 human baselines with thirteen frontier models. V55 verification basis: metr.org was not fetched. The trial figures, the arXiv identifier 2507.09089, the task and participant counts and METR's own historical caveat were carried in search results attributing them to METR; the seven-month doubling, the 170-task suite and the 800-plus human baselines were carried with attribution to METR's March 2025 post. Any specific model's time horizon, and reports of a faster doubling in the most recent period, come from secondary summaries and are NOT asserted in this pack.
Close the narrative
Observed changes · 0

No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

Actions

Read the assembled narrativeContinue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

/atlas?object=ORG_METR&experience=ORG_METR