GDPval

GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

取自美国 GDP 贡献最大的 9 个行业、44 个职业的真实工作交付物(法律文书、电子表格、幻灯片、CAD、排班、视频剪辑),由平均 14 年经验的专业人士编写与审核。模型一次性根据任务描述和参考文件生成交付物,由同职业专家盲评与人类专家作品比较,得分为被评为优于或不逊于专家的比例。是衡量有经济价值的知识工作的主要公开标尺。

74.1%
GPT-5.2 Pro
厂商自报
2025-10-05 → 2025-12-11: 12.5% → 74.1%
发布
2025-09
维护者
OpenAI
状态
active
污染风险
medium
指标
win-or-tie rate vs industry experts (percent, ↑)
题量
220
领域
general-assistant knowledge instruction-following
人类
无实测基线

备注. Announced 2025-09-25; arXiv 2510.04374 verified (Patwardhan et al., posted 2025-10-05). Grading is pairwise by human experts (3 samples x 3 graders per task); an experimental automated grader is offered at evals.openai.com but is not a substitute. 50% is expert parity by construction. Artificial Analysis runs the same 220 tasks agentically as GDPval-AA with Elo scoring, which is a different metric and is not recorded here. Vendor-reported 2025-12 and later numbers use newer ChatGPT tools not available to earlier models.

完整账本

系统开发者分数日期来源条件
GPT-5.2 ProOpenAI74.1%厂商自报split: gold tools
Wins-or-ties, from the appendix table of the GPT-5.2 announcement.
GPT-5.2 ThinkingOpenAI70.9%厂商自报split: gold reasoning_effort: heavy tools
Wins-or-ties with expert human judges; run with ChatGPT spreadsheet/presentation tools not available to GPT-5.
Claude Opus 4.1Anthropic47.6%论文split: gold
Best model in the paper; sampled via the Claude UI with file-creation features. Run by OpenAI, not Anthropic.
GPT-5 (high)OpenAI39%论文split: gold reasoning_effort: high tools
Table 2; web search and code interpreter enabled. OpenAI's 2025-12 post cites 38.8% for the same model.
o3OpenAI35.2%论文split: gold
Table 2 of the paper; expert pairwise grading.
GPT-4oOpenAI12.5%论文split: gold
Table 2 of the paper; expert pairwise grading, 220 gold tasks.