GPQA Diamond
Graduate-Level Google-Proof Q&A Benchmark, Diamond subset
198 four-option multiple-choice questions in biology, physics and chemistry written by domain PhDs and validated so that experts agree while skilled non-experts with unrestricted web access fail. The Diamond subset keeps only questions both experts answered correctly and most non-experts missed. The standard frontier science-reasoning benchmark; noise floor is roughly plus or minus 3 points at n=198.
- Released
- 2023-11
- Maintainer
- NYU / Cohere / Anthropic (Rein et al.)
- Status
- Contamination
- medium
- Metric
- accuracy (percent, ↑)
- Tasks
- 198
- Domains
- science reasoning knowledge
- human
- 69.7% domain PhD experts (per-question, in-domain)
Notes. Answer options are public; the maintainers ask that the dataset never be posted in plain text. Scores above ~92% are within noise of each other.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | 96% | aggregator | — | |
| GPT-6 Astra (max) | OpenAI | 95.8% | independent | Epoch AI Benchmarking Hub run started 2026-08-30 (stderr 1.3); OpenAI's launch table reports 96.0, which benchlm.ai republishes. | |
| Gemini 3.1 Pro | Google DeepMind | 94.3% | aggregator | — | |
| Gemini 3 Pro | Google DeepMind | 91.9% | self-reported | tools: no Gemini 3 launch post; Epoch's independent run of gemini-3-pro-preview scored 92.6 the next day. | |
| o3-2025-04-16 (high) | OpenAI | 81.8% | independent | Epoch AI Benchmarking Hub; OpenAI's own o3 launch figure was 83.3. | |
| o1-2024-12-17 (high) | OpenAI | 76.8% | independent | Epoch AI Benchmarking Hub run (data export); gpt-4o-2024-08-06 scored 49.2 and claude-3-5-sonnet-20241022 55.3 in January 2025 runs. | |
| GPT-4 | OpenAI | 39.4% | paper | shots: 0 Table 1 of the GPQA paper. |