GDPval
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
Real work deliverables (legal briefs, spreadsheets, slide decks, CAD, schedules, video edits) drawn from 44 occupations in the 9 US sectors that contribute most to GDP, written and reviewed by professionals averaging 14 years of experience. A model receives the request plus reference files and produces the deliverable in one shot; occupational experts blindly compare it with the human expert's work and the score is the share of comparisons rated better than or as good as the expert. The main public yardstick for economically valuable knowledge work.
- Released
- 2025-09
- Maintainer
- OpenAI
- Status
- active
- Contamination
- medium
- Metric
- win-or-tie rate vs industry experts (percent, ↑)
- Tasks
- 220
- Domains
- general-assistant knowledge instruction-following
- human
- no measured baseline
Notes. Announced 2025-09-25; arXiv 2510.04374 verified (Patwardhan et al., posted 2025-10-05). Grading is pairwise by human experts (3 samples x 3 graders per task); an experimental automated grader is offered at evals.openai.com but is not a substitute. 50% is expert parity by construction. Artificial Analysis runs the same 220 tasks agentically as GDPval-AA with Elo scoring, which is a different metric and is not recorded here. Vendor-reported 2025-12 and later numbers use newer ChatGPT tools not available to earlier models.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| GPT-5.2 Pro | OpenAI | 74.1% | self-reported | split: gold tools Wins-or-ties, from the appendix table of the GPT-5.2 announcement. | |
| GPT-5.2 Thinking | OpenAI | 70.9% | self-reported | split: gold reasoning_effort: heavy tools Wins-or-ties with expert human judges; run with ChatGPT spreadsheet/presentation tools not available to GPT-5. | |
| Claude Opus 4.1 | Anthropic | 47.6% | paper | split: gold Best model in the paper; sampled via the Claude UI with file-creation features. Run by OpenAI, not Anthropic. | |
| GPT-5 (high) | OpenAI | 39% | paper | split: gold reasoning_effort: high tools Table 2; web search and code interpreter enabled. OpenAI's 2025-12 post cites 38.8% for the same model. | |
| o3 | OpenAI | 35.2% | paper | split: gold Table 2 of the paper; expert pairwise grading. | |
| GPT-4o | OpenAI | 12.5% | paper | split: gold Table 2 of the paper; expert pairwise grading, 220 gold tasks. |