GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
OpenAI's GDPval (September 2025; 220 expert-graded tasks across 44 occupations, benchmark) found that the best model tested, Claude Opus 4.1, produced work rated as good as or better than an industry expert's on 47.6% of tasks.
Key findings
- 01Claude Opus 4.1's deliverables were rated better than or as good as the human expert's on 47.6% of the 220 gold-subset tasks.Graders were occupational experts comparing AI and human work blind. Claude did best on formatting and presentation; GPT-5 did best on accuracy. (p. 6)
- 02Among OpenAI models, the win rate against experts rose from 12.5% for GPT-4o to 39.0% for GPT-5.OpenAI describes the improvement across its model releases as roughly linear over time. (p. 15)
- 03When an expert tries GPT-5, reviews the output and redoes the work if it falls short, the process was estimated to be 1.12x to 1.39x faster and 1.18x to 1.63x cheaper than the expert alone.Raw model inference is roughly 90x faster than the expert, but review time and failed attempts erase most of that gap. The analysis does not count the cost of catastrophic mistakes. (p. 15)
- 04About 29% of expert ratings of GPT-5's losing deliverables called them bad or catastrophic, with about 3% catastrophic.The most common rating for a GPT-5 failure was 'acceptable but subpar'. (p. 15)
- 05The 44 occupations come from the 9 sectors contributing most to U.S. GDP and collectively earn $3 trillion a year.Tasks are one-shot and fully specified, with the context supplied up front, rather than interactive work that unfolds over days. (p. 2)
By the numbers
What it means for you Draft
On well-defined, self-contained knowledge tasks, the best 2025 models came close to experienced professionals about half the time. The same study shows that once someone has to check the output and redo the misses, the time and cost savings shrink to tens of percent, not multiples. For a $10–100M firm, that points to careful use on specific, checkable deliverables, with a person still accountable for the result.
Pick tasks with clear specifications and easy-to-check outputs, and budget real review time, because failures are often formatting or instruction-following misses and a small share are serious. Measure your own win rate on a sample of your real work before assuming savings.
Limitations
Medium trust.Transparent method, expert blind grading and open 220-task set; OpenAI sells the AI models it is evaluating.
Built and run by OpenAI, which sells the models being tested. Tasks are one-shot and fully specified, so they miss the context-gathering, iteration and coordination in real jobs; the speed and cost scenarios are modeled estimates, not observed workplace outcomes.
About the publisher: Makes ChatGPT; studies its own customers.
OpenAI. "GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks." September 25, 2025. https://openai.com/index/gdpval/


