OpenAI · September 2025

GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

OpenAI's GDPval (September 2025; 220 expert-graded tasks across 44 occupations, benchmark) found that the best model tested, Claude Opus 4.1, produced work rated as good as or better than an industry expert's on 47.6% of tasks.

Read the original report ↗Cite2 min read · Summary updated
Model win rate against industry experts, OpenAI models (GDPval gold subset)%
GPT-5
39.0%
o3
35.2%
o4-mini
29.1%
GPT-4o
12.5%
Source: OpenAI, GDPval, 2025, p. 15 (Table 2).

Key findings

  1. 01
    Claude Opus 4.1's deliverables were rated better than or as good as the human expert's on 47.6% of the 220 gold-subset tasks.Graders were occupational experts comparing AI and human work blind. Claude did best on formatting and presentation; GPT-5 did best on accuracy. (p. 6)
  2. 02
    Among OpenAI models, the win rate against experts rose from 12.5% for GPT-4o to 39.0% for GPT-5.OpenAI describes the improvement across its model releases as roughly linear over time. (p. 15)
  3. 03
    When an expert tries GPT-5, reviews the output and redoes the work if it falls short, the process was estimated to be 1.12x to 1.39x faster and 1.18x to 1.63x cheaper than the expert alone.Raw model inference is roughly 90x faster than the expert, but review time and failed attempts erase most of that gap. The analysis does not count the cost of catastrophic mistakes. (p. 15)
  4. 04
    About 29% of expert ratings of GPT-5's losing deliverables called them bad or catastrophic, with about 3% catastrophic.The most common rating for a GPT-5 failure was 'acceptable but subpar'. (p. 15)
  5. 05
    The 44 occupations come from the 9 sectors contributing most to U.S. GDP and collectively earn $3 trillion a year.Tasks are one-shot and fully specified, with the context supplied up front, rather than interactive work that unfolds over days. (p. 2)

By the numbers

47.6%best model rated as good as or better than experts (Claude Opus 4.1)
39.0%GPT-5 win rate vs. experts, up from 12.5% for GPT-4o
1.12–1.39xestimated speedup when an expert reviews and fixes GPT-5 output

What it means for you Draft

For executives at $10–100M companies

On well-defined, self-contained knowledge tasks, the best 2025 models came close to experienced professionals about half the time. The same study shows that once someone has to check the output and redo the misses, the time and cost savings shrink to tens of percent, not multiples. For a $10–100M firm, that points to careful use on specific, checkable deliverables, with a person still accountable for the result.

For practitioners

Pick tasks with clear specifications and easy-to-check outputs, and budget real review time, because failures are often formatting or instruction-following misses and a small share are serious. Measure your own win rate on a sample of your real work before assuming savings.

Limitations

Medium trust.Transparent method, expert blind grading and open 220-task set; OpenAI sells the AI models it is evaluating.

Built and run by OpenAI, which sells the models being tested. Tasks are one-shot and fully specified, so they miss the context-gathering, iteration and coordination in real jobs; the speed and cost scenarios are modeled estimates, not observed workplace outcomes.

About the publisher: Makes ChatGPT; studies its own customers.

Cite the original

OpenAI. "GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks." September 25, 2025. https://openai.com/index/gdpval/