GPQA Diamond

Graduate-Level Google-Proof Q&A Benchmark, Diamond subset

198 four-option multiple-choice questions in biology, physics and chemistry written by domain PhDs and validated so that experts agree while skilled non-experts with unrestricted web access fail. The Diamond subset keeps only questions both experts answered correctly and most non-experts missed. The standard frontier science-reasoning benchmark; noise floor is roughly plus or minus 3 points at n=198.

96% · human 69.7%
GPT-6 Astra
aggregator
2023-11-20 → 2026-09-03: 39.4% → 96%
Released
2023-11
Maintainer
NYU / Cohere / Anthropic (Rein et al.)
Status
saturating
Contamination
medium
Metric
accuracy (percent, ↑)
Tasks
198
Domains
science reasoning knowledge
human
69.7% domain PhD experts (per-question, in-domain)

Notes. Answer options are public; the maintainers ask that the dataset never be posted in plain text. Scores above ~92% are within noise of each other.

Full ledger

SystemDeveloperScoreDateSourceConditions
GPT-6 AstraOpenAI96%aggregator
GPT-6 Astra (max)OpenAI95.8%independent
Epoch AI Benchmarking Hub run started 2026-08-30 (stderr 1.3); OpenAI's launch table reports 96.0, which benchlm.ai republishes.
Gemini 3.1 ProGoogle DeepMind94.3%aggregator
Gemini 3 ProGoogle DeepMind91.9%self-reportedtools: no
Gemini 3 launch post; Epoch's independent run of gemini-3-pro-preview scored 92.6 the next day.
o3-2025-04-16 (high)OpenAI81.8%independent
Epoch AI Benchmarking Hub; OpenAI's own o3 launch figure was 83.3.
o1-2024-12-17 (high)OpenAI76.8%independent
Epoch AI Benchmarking Hub run (data export); gpt-4o-2024-08-06 scored 49.2 and claude-3-5-sonnet-20241022 55.3 in January 2025 runs.
GPT-4OpenAI39.4%papershots: 0
Table 1 of the GPQA paper.