GPQA Diamond
Graduate-Level Google-Proof Q&A Benchmark, Diamond subset
198 道生物、物理和化学领域的四选一多项选择题,由领域博士编写并经验证:专家意见一致,而可无限制访问网络的熟练非专家却会答错。Diamond 子集仅保留两位专家均答对且多数非专家答错的题目。是前沿科学推理的标准基准;在 n=198 时噪声下限约为正负 3 分。
- 发布
- 2023-11
- 维护者
- NYU / Cohere / Anthropic (Rein et al.)
- 状态
- 污染风险
- medium
- 指标
- accuracy (percent, ↑)
- 题量
- 198
- 领域
- science reasoning knowledge
- 人类
- 69.7% domain PhD experts (per-question, in-domain)
备注. Answer options are public; the maintainers ask that the dataset never be posted in plain text. Scores above ~92% are within noise of each other.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | 96% | 聚合站 | — | |
| GPT-6 Astra (max) | OpenAI | 95.8% | 独立复现 | Epoch AI Benchmarking Hub run started 2026-08-30 (stderr 1.3); OpenAI's launch table reports 96.0, which benchlm.ai republishes. | |
| Gemini 3.1 Pro | Google DeepMind | 94.3% | 聚合站 | — | |
| Gemini 3 Pro | Google DeepMind | 91.9% | 厂商自报 | tools: no Gemini 3 launch post; Epoch's independent run of gemini-3-pro-preview scored 92.6 the next day. | |
| o3-2025-04-16 (high) | OpenAI | 81.8% | 独立复现 | Epoch AI Benchmarking Hub; OpenAI's own o3 launch figure was 83.3. | |
| o1-2024-12-17 (high) | OpenAI | 76.8% | 独立复现 | Epoch AI Benchmarking Hub run (data export); gpt-4o-2024-08-06 scored 49.2 and claude-3-5-sonnet-20241022 55.3 in January 2025 runs. | |
| GPT-4 | OpenAI | 39.4% | 论文 | shots: 0 Table 1 of the GPQA paper. |