METR (Model Evaluation & Threat Research) · July 2025

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

METR's Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (July 2025; n=16 developers, 246 tasks, randomized controlled trial) found that developers took 19% longer to finish tasks when AI tools were allowed.

Read the original report ↗Cite2 min read · Summary updated
Expected vs. perceived vs. measured effect of AI on task time% change in speed (negative = slower)
Developers' forecast before study
+24%
Developers' belief after study
+20%
Measured effect
−19%
Source: METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 2025.

Key findings

  1. 01
    Developers took 19% longer to complete issues when AI tools were allowed.Each of 246 real issues was randomly assigned to allow or disallow AI; tools were mainly Cursor Pro with Claude 3.5/3.7 Sonnet.
  2. 02
    Before the study, developers expected AI to speed them up by 24%.The measured result ran in the opposite direction from what participants predicted.
  3. 03
    After experiencing the slowdown, developers still believed AI had sped them up by 20%.Self-reported productivity gains were not a reliable guide to measured effects in this setting.
  4. 04
    The 19% slowdown had a confidence interval of +2% to +39%.METR reported this interval in its February 2026 follow-up note; the study had just enough statistical power to rule out zero effect.
  5. 05
    Tasks averaged about two hours each, in repositories averaging 22k+ GitHub stars and 1M+ lines of code.Developers were paid $150 an hour and self-reported implementation time while recording their screens.

By the numbers

19%slower with AI allowed (measured)
24%speedup developers expected beforehand
20%speedup developers believed afterward

What it means for you Draft

For executives at $10–100M companies

This is one of the few randomized tests of AI on real work, and it found experienced people got slower, while feeling faster. For a $10–100M company, it is a caution against judging AI tools by how staff say they feel; try to measure time or output on the same kind of work with and without the tool. It is also an early-2025 snapshot, and tools have changed since.

For practitioners

Experts working in code they know deeply may get less from AI than newcomers or people on unfamiliar tasks. If you pilot a tool, randomize or at least compare matched tasks rather than relying on self-reported time savings.

Limitations

High trust.Pre-specified randomized trial with methods and paper published; small sample; METR is a nonprofit that does not sell AI tools.

Only 16 developers, all highly experienced and working in large, mature open-source repositories they knew well; results may not carry to other developers, codebases, or newer tools. Time was self-reported, though screen recordings were collected.

Cite the original

METR (Model Evaluation & Threat Research). "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity." July 10, 2025. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/