GDPval

GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Real work deliverables (legal briefs, spreadsheets, slide decks, CAD, schedules, video edits) drawn from 44 occupations in the 9 US sectors that contribute most to GDP, written and reviewed by professionals averaging 14 years of experience. A model receives the request plus reference files and produces the deliverable in one shot; occupational experts blindly compare it with the human expert's work and the score is the share of comparisons rated better than or as good as the expert. The main public yardstick for economically valuable knowledge work.

74.1%
GPT-5.2 Pro
self-reported
2025-10-05 → 2025-12-11: 12.5% → 74.1%
Released
2025-09
Maintainer
OpenAI
Status
active
Contamination
medium
Metric
win-or-tie rate vs industry experts (percent, ↑)
Tasks
220
Domains
general-assistant knowledge instruction-following
human
no measured baseline

Notes. Announced 2025-09-25; arXiv 2510.04374 verified (Patwardhan et al., posted 2025-10-05). Grading is pairwise by human experts (3 samples x 3 graders per task); an experimental automated grader is offered at evals.openai.com but is not a substitute. 50% is expert parity by construction. Artificial Analysis runs the same 220 tasks agentically as GDPval-AA with Elo scoring, which is a different metric and is not recorded here. Vendor-reported 2025-12 and later numbers use newer ChatGPT tools not available to earlier models.

Full ledger

SystemDeveloperScoreDateSourceConditions
GPT-5.2 ProOpenAI74.1%self-reportedsplit: gold tools
Wins-or-ties, from the appendix table of the GPT-5.2 announcement.
GPT-5.2 ThinkingOpenAI70.9%self-reportedsplit: gold reasoning_effort: heavy tools
Wins-or-ties with expert human judges; run with ChatGPT spreadsheet/presentation tools not available to GPT-5.
Claude Opus 4.1Anthropic47.6%papersplit: gold
Best model in the paper; sampled via the Claude UI with file-creation features. Run by OpenAI, not Anthropic.
GPT-5 (high)OpenAI39%papersplit: gold reasoning_effort: high tools
Table 2; web search and code interpreter enabled. OpenAI's 2025-12 post cites 38.8% for the same model.
o3OpenAI35.2%papersplit: gold
Table 2 of the paper; expert pairwise grading.
GPT-4oOpenAI12.5%papersplit: gold
Table 2 of the paper; expert pairwise grading, 220 gold tasks.