METR
node · earth
The non-profit research organisation, based in the United States, whose distinctive contribution has been to measure what models do in conditions closer to work than to examinations. Two of its outputs are load-bearing here. In July 2025 it published a randomised controlled trial in which sixteen experienced open-source developers completed 246 real issues in repositories they knew well, and were 19% slower with early-2025 AI assistance than without while believing they had been about 20% faster — a result METR itself has since flagged as historical, describing tools and workflows that have changed. Separately it publishes a task time-horizon metric: the length of task, measured by how long skilled humans take, that a model completes autonomously with 50% reliability, reported as doubling roughly every seven months over 2019–2025 with a faster doubling in the most recent period. The metric is the interesting object, not the extrapolation from it: it converts an unbounded question about capability into a quantity with a unit and a stated reliability threshold.
A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.
Read in · 1
Timeline
No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.
Connections · 2
Assembled narrative · 1
Assembled from 31 blocks · 1 evidence · 28 related
- Story
- The non-profit research organisation, based in the United States, whose distinctive contribution has been to measure what models do in conditions closer to work than to examinations. Two of its outputs are load-bearing here. In July 2025 it published a randomised controlled trial in which sixteen experienced open-source developers completed 246 real issues in repositories they knew well, and were 19% slower with early-2025 AI assistance than without while believing they had been about 20% faster — a result METR itself has since flagged as historical, describing tools and workflows that have changed. Separately it publishes a task time-horizon metric: the length of task, measured by how long skilled humans take, that a model completes autonomously with 50% reliability, reported as doubling roughly every seven months over 2019–2025 with a faster doubling in the most recent period. The metric is the interesting object, not the extrapolation from it: it converts an unbounded question about capability into a quantity with a unit and a stated reliability threshold.
- Knowledge
- Machine-generated code
- The difficulty of evaluating machine learning systems
- METR
- Connections
- Machine-generated code
- The difficulty of evaluating machine learning systems
- A randomised trial finds experienced developers slower with AI assistance
- Labour-market effects of machine learning
- The transformer architecture
- METR
- A randomised trial finds experienced developers slower with AI assistance
- Payroll data shows a relative decline in entry-level employment in AI-exposed occupations
- The second International AI Safety Report reports mixed labour findings
- The AI Index records both fast benchmark movement and doubts about benchmarks
- Introducing SWE-bench Verified
- CASP — Critical Assessment of Structure Prediction
- Construct validity in AI evaluation
- Benchmark saturation
- METR
- European Union Artificial Intelligence Act
- A human-parity claim in translation fails under document-level evaluation
- A randomised trial finds experienced developers slower with AI assistance
- An audit finds the most-cited preference leaderboard is optimised rather than merely measured
- A systematic review of 445 benchmarks reports pervasive validity failures
- The Commission proposes deferring the high-risk obligations
- The AI Index records both fast benchmark movement and doubts about benchmarks
- Evidence
- In a randomised trial, sixteen experienced open-source developers working on 246 real issues in repositories they knew well — about five years of prior experience with the projects on average — took 19% longer to complete tasks when allowed to use AI tools available between February and June 2025, while estimating afterwards that the tools had made them about 20% faster; METR has since labelled the result historical, as not necessarily reflecting current tools or workflows. Separately, METR reports a task time-horizon metric — the human-measured duration of task a model completes autonomously with 50% reliability — as having doubled approximately every seven months over 2019–2025, measured across 170 tasks against over 800 human baselines with thirteen frontier models. V55 verification basis: metr.org was not fetched. The trial figures, the arXiv identifier 2507.09089, the task and participant counts and METR's own historical caveat were carried in search results attributing them to METR; the seven-month doubling, the 170-task suite and the 800-plus human baselines were carried with attribution to METR's March 2025 post. Any specific model's time horizon, and reports of a faster doubling in the most recent period, come from secondary summaries and are NOT asserted in this pack.
Observed changes · 0
No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.
Actions
/atlas?object=ORG_METR&experience=ORG_METR