GDPval
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
取自美国 GDP 贡献最大的 9 个行业、44 个职业的真实工作交付物(法律文书、电子表格、幻灯片、CAD、排班、视频剪辑),由平均 14 年经验的专业人士编写与审核。模型一次性根据任务描述和参考文件生成交付物,由同职业专家盲评与人类专家作品比较,得分为被评为优于或不逊于专家的比例。是衡量有经济价值的知识工作的主要公开标尺。
- 发布
- 2025-09
- 维护者
- OpenAI
- 状态
- active
- 污染风险
- medium
- 指标
- win-or-tie rate vs industry experts (percent, ↑)
- 题量
- 220
- 领域
- general-assistant knowledge instruction-following
- 人类
- 无实测基线
备注. Announced 2025-09-25; arXiv 2510.04374 verified (Patwardhan et al., posted 2025-10-05). Grading is pairwise by human experts (3 samples x 3 graders per task); an experimental automated grader is offered at evals.openai.com but is not a substitute. 50% is expert parity by construction. Artificial Analysis runs the same 220 tasks agentically as GDPval-AA with Elo scoring, which is a different metric and is not recorded here. Vendor-reported 2025-12 and later numbers use newer ChatGPT tools not available to earlier models.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| GPT-5.2 Pro | OpenAI | 74.1% | 厂商自报 | split: gold tools Wins-or-ties, from the appendix table of the GPT-5.2 announcement. | |
| GPT-5.2 Thinking | OpenAI | 70.9% | 厂商自报 | split: gold reasoning_effort: heavy tools Wins-or-ties with expert human judges; run with ChatGPT spreadsheet/presentation tools not available to GPT-5. | |
| Claude Opus 4.1 | Anthropic | 47.6% | 论文 | split: gold Best model in the paper; sampled via the Claude UI with file-creation features. Run by OpenAI, not Anthropic. | |
| GPT-5 (high) | OpenAI | 39% | 论文 | split: gold reasoning_effort: high tools Table 2; web search and code interpreter enabled. OpenAI's 2025-12 post cites 38.8% for the same model. | |
| o3 | OpenAI | 35.2% | 论文 | split: gold Table 2 of the paper; expert pairwise grading. | |
| GPT-4o | OpenAI | 12.5% | 论文 | split: gold Table 2 of the paper; expert pairwise grading, 220 gold tasks. |