HealthBench

HealthBench: Evaluating Large Language Models Towards Improved Human Health

5,000 realistic multi-turn, multilingual health conversations between a model and a layperson or clinician, spanning seven themes such as emergency referrals, context seeking and global health. Each conversation carries a physician-written rubric (48,562 criteria from 262 physicians in 60 countries); a GPT-4.1 grader checks which criteria the final response meets and the score is points earned over the maximum. Hard (1,000) and Consensus (3,671) subsets isolate unsaturated and physician-validated criteria. The reference evaluation for medical helpfulness and safety.

59.9%
o3
paper
2025-05-13 → 2025-08-07: 32.3% → 46.2%
Released
2025-05
Maintainer
OpenAI
Status
active
Contamination
medium
Metric
rubric score (percent, ↑)
Tasks
5,000
Domains
knowledge safety instruction-following
human
no measured baseline

Notes. Announced 2025-05-12; arXiv 2505.08775 verified (Arora et al., posted 2025-05-13). Scores are 0-1 in the paper and shown here as percent. Physician responses without model help scored below the September 2024 models, and physicians could not improve on o3/GPT-4.1 responses, so no human baseline is recorded. OpenAI's 2026 launch posts report a 'HealthBench Professional (length-adjusted)' variant graded by GPT-5.4; it is not documented publicly and is not tracked here. Conversations are public with a canary string.

Full ledger

SystemDeveloperScoreDateSourceConditions
o3OpenAI59.9%paper
Best model at release; mean of 16 runs (Table 7). Paper states o3 also tops HealthBench Hard at 32%.
GPT-4.1OpenAI47.8%paper
Mean of 16 runs (Table 7).
GPT-5 (thinking)OpenAI46.2%self-reportedsplit: hard reasoning_effort: high
HealthBench Hard only; the announcement gives no full-set number.
GPT-4o (Aug 2024)OpenAI32.3%paper
Mean of 16 runs (Table 7).