GPQA Diamond

Graduate-Level Google-Proof Q&A Benchmark, Diamond subset

198 道生物、物理和化学领域的四选一多项选择题,由领域博士编写并经验证:专家意见一致,而可无限制访问网络的熟练非专家却会答错。Diamond 子集仅保留两位专家均答对且多数非专家答错的题目。是前沿科学推理的标准基准;在 n=198 时噪声下限约为正负 3 分。

96% · 人类 69.7%
GPT-6 Astra
聚合站
2023-11-20 → 2026-09-03: 39.4% → 96%
发布
2023-11
维护者
NYU / Cohere / Anthropic (Rein et al.)
状态
saturating
污染风险
medium
指标
accuracy (percent, ↑)
题量
198
领域
science reasoning knowledge
人类
69.7% domain PhD experts (per-question, in-domain)

备注. Answer options are public; the maintainers ask that the dataset never be posted in plain text. Scores above ~92% are within noise of each other.

完整账本

系统开发者分数日期来源条件
GPT-6 AstraOpenAI96%聚合站
GPT-6 Astra (max)OpenAI95.8%独立复现
Epoch AI Benchmarking Hub run started 2026-08-30 (stderr 1.3); OpenAI's launch table reports 96.0, which benchlm.ai republishes.
Gemini 3.1 ProGoogle DeepMind94.3%聚合站
Gemini 3 ProGoogle DeepMind91.9%厂商自报tools: no
Gemini 3 launch post; Epoch's independent run of gemini-3-pro-preview scored 92.6 the next day.
o3-2025-04-16 (high)OpenAI81.8%独立复现
Epoch AI Benchmarking Hub; OpenAI's own o3 launch figure was 83.3.
o1-2024-12-17 (high)OpenAI76.8%独立复现
Epoch AI Benchmarking Hub run (data export); gpt-4o-2024-08-06 scored 49.2 and claude-3-5-sonnet-20241022 55.3 in January 2025 runs.
GPT-4OpenAI39.4%论文shots: 0
Table 1 of the GPQA paper.