Remote Labor Index: Measuring AI Automation of Remote Work
The Center for AI Safety and Scale AI's Remote Labor Index (October 2025; 240 real freelance projects, benchmark) found that the best AI agent tested completed only 2.5% of projects at a quality a reasonable client would accept.
Key findings
- 01The best-performing agent, Manus, reached an automation rate of 2.5%; Grok 4 and Claude Sonnet 4.5 each reached 2.1%, GPT-5 1.7%, ChatGPT agent 1.3%, and Gemini 2.5 Pro 0.8%.A project counted as automated only if evaluators judged the AI deliverable would be accepted by a reasonable client as the commissioned work. (p. 9)
- 02The 240 projects span 23 Upwork work categories and represent over 6,000 hours of human work valued at over $140,000.Categories include game development, product design, architecture, data analysis and video animation, which is broader than the software-heavy mix of most AI benchmarks. (p. 4–5)
- 03Projects took human freelancers a mean of 28.9 hours (median 11.5) and cost a mean of $632.60 (median $200).The authors say this is more than twice as long as tasks in prior agent benchmarks and close to the length of typical Upwork jobs. (p. 5)
- 0445.6% of AI deliverables were rated poor quality, 35.7% incomplete, 17.6% had corrupted or unusable files, and 14.8% had inconsistencies.Categories overlap. Examples include videos far shorter than requested and files in the wrong format. (p. 11)
- 05AI did comparably well on a narrow set of work: audio editing and mixing, image generation, report writing, and code for interactive data visualizations.These are the kinds of self-contained, easy-to-check outputs where current agents came closest to human deliverables.
By the numbers
What it means for you Draft
When agents were handed whole, real freelance jobs end to end, they almost never produced something a paying client would accept. For a $10–100M company, this suggests agents are not yet a substitute for outsourcing entire projects; the practical gains still look like helping people with parts of the work. The models tested are from 2025, so this is a baseline to recheck, not a permanent ceiling.
Treat agent output on multi-step, multi-file deliverables as a draft that needs review, and expect the most common failures to be quality, incompleteness and broken files. Narrow, self-contained tasks such as audio cleanup, images, short reports and interactive data visualizations are where agents came closest.
Limitations
High trust.Real paid projects, disclosed grading with 94.4% inter-annotator agreement; co-author Scale AI sells AI training data and evaluation services.
Tests agents working alone end to end, not people working with AI, so it understates the value of AI assistance. Covers only remote, computer-based work, and results reflect the specific models and scaffolds tested in late 2025.
Center for AI Safety. "Remote Labor Index: Measuring AI Automation of Remote Work." October 30, 2025. https://www.remotelabor.ai/


