HealthBench
HealthBench: Evaluating Large Language Models Towards Improved Human Health
5,000 realistic multi-turn, multilingual health conversations between a model and a layperson or clinician, spanning seven themes such as emergency referrals, context seeking and global health. Each conversation carries a physician-written rubric (48,562 criteria from 262 physicians in 60 countries); a GPT-4.1 grader checks which criteria the final response meets and the score is points earned over the maximum. Hard (1,000) and Consensus (3,671) subsets isolate unsaturated and physician-validated criteria. The reference evaluation for medical helpfulness and safety.
- Released
- 2025-05
- Maintainer
- OpenAI
- Status
- active
- Contamination
- medium
- Metric
- rubric score (percent, ↑)
- Tasks
- 5,000
- Domains
- knowledge safety instruction-following
- human
- no measured baseline
Notes. Announced 2025-05-12; arXiv 2505.08775 verified (Arora et al., posted 2025-05-13). Scores are 0-1 in the paper and shown here as percent. Physician responses without model help scored below the September 2024 models, and physicians could not improve on o3/GPT-4.1 responses, so no human baseline is recorded. OpenAI's 2026 launch posts report a 'HealthBench Professional (length-adjusted)' variant graded by GPT-5.4; it is not documented publicly and is not tracked here. Conversations are public with a canary string.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| o3 | OpenAI | 59.9% | paper | Best model at release; mean of 16 runs (Table 7). Paper states o3 also tops HealthBench Hard at 32%. | |
| GPT-4.1 | OpenAI | 47.8% | paper | Mean of 16 runs (Table 7). | |
| GPT-5 (thinking) | OpenAI | 46.2% | self-reported | split: hard reasoning_effort: high HealthBench Hard only; the announcement gives no full-set number. | |
| GPT-4o (Aug 2024) | OpenAI | 32.3% | paper | Mean of 16 runs (Table 7). |