Center for AI Safety · October 2025

Remote Labor Index: Measuring AI Automation of Remote Work

The Center for AI Safety and Scale AI's Remote Labor Index (October 2025; 240 real freelance projects, benchmark) found that the best AI agent tested completed only 2.5% of projects at a quality a reasonable client would accept.

Read the original report ↗Cite2 min read · Summary updated
Share of freelance projects completed at client-acceptable quality%
Manus
2.5%
Grok 4
2.1%
Claude Sonnet 4.5
2.1%
GPT-5
1.7%
ChatGPT agent
1.3%
Gemini 2.5 Pro
0.8%
Source: Center for AI Safety and Scale AI, Remote Labor Index, 2025, p. 9.

Key findings

  1. 01
    The best-performing agent, Manus, reached an automation rate of 2.5%; Grok 4 and Claude Sonnet 4.5 each reached 2.1%, GPT-5 1.7%, ChatGPT agent 1.3%, and Gemini 2.5 Pro 0.8%.A project counted as automated only if evaluators judged the AI deliverable would be accepted by a reasonable client as the commissioned work. (p. 9)
  2. 02
    The 240 projects span 23 Upwork work categories and represent over 6,000 hours of human work valued at over $140,000.Categories include game development, product design, architecture, data analysis and video animation, which is broader than the software-heavy mix of most AI benchmarks. (p. 4–5)
  3. 03
    Projects took human freelancers a mean of 28.9 hours (median 11.5) and cost a mean of $632.60 (median $200).The authors say this is more than twice as long as tasks in prior agent benchmarks and close to the length of typical Upwork jobs. (p. 5)
  4. 04
    45.6% of AI deliverables were rated poor quality, 35.7% incomplete, 17.6% had corrupted or unusable files, and 14.8% had inconsistencies.Categories overlap. Examples include videos far shorter than requested and files in the wrong format. (p. 11)
  5. 05
    AI did comparably well on a narrow set of work: audio editing and mixing, image generation, report writing, and code for interactive data visualizations.These are the kinds of self-contained, easy-to-check outputs where current agents came closest to human deliverables.

By the numbers

2.5%best agent automation rate (Manus)
240real freelance projects in 23 categories
$140K+total value of the human-completed projects

What it means for you Draft

For executives at $10–100M companies

When agents were handed whole, real freelance jobs end to end, they almost never produced something a paying client would accept. For a $10–100M company, this suggests agents are not yet a substitute for outsourcing entire projects; the practical gains still look like helping people with parts of the work. The models tested are from 2025, so this is a baseline to recheck, not a permanent ceiling.

For practitioners

Treat agent output on multi-step, multi-file deliverables as a draft that needs review, and expect the most common failures to be quality, incompleteness and broken files. Narrow, self-contained tasks such as audio cleanup, images, short reports and interactive data visualizations are where agents came closest.

Limitations

High trust.Real paid projects, disclosed grading with 94.4% inter-annotator agreement; co-author Scale AI sells AI training data and evaluation services.

Tests agents working alone end to end, not people working with AI, so it understates the value of AI assistance. Covers only remote, computer-based work, and results reflect the specific models and scaffolds tested in late 2025.

Cite the original

Center for AI Safety. "Remote Labor Index: Measuring AI Automation of Remote Work." October 30, 2025. https://www.remotelabor.ai/