HealthBench

HealthBench: Evaluating Large Language Models Towards Improved Human Health

5,000 段模型与普通用户或临床人员之间的真实多轮、多语言健康对话,涵盖急诊转诊、追问上下文、全球健康等七类主题。每段对话附医生撰写的评分细则(来自 60 国 262 名医生的 48,562 条标准);GPT-4.1 评分器判断末轮回复满足哪些标准,得分为所得分值占满分的比例。Hard(1,000)与 Consensus(3,671)子集分别隔离未饱和与医生共识标准。医疗有用性与安全性的参考评估。

59.9%
o3
论文
2025-05-13 → 2025-08-07: 32.3% → 46.2%
发布
2025-05
维护者
OpenAI
状态
active
污染风险
medium
指标
rubric score (percent, ↑)
题量
5,000
领域
knowledge safety instruction-following
人类
无实测基线

备注. Announced 2025-05-12; arXiv 2505.08775 verified (Arora et al., posted 2025-05-13). Scores are 0-1 in the paper and shown here as percent. Physician responses without model help scored below the September 2024 models, and physicians could not improve on o3/GPT-4.1 responses, so no human baseline is recorded. OpenAI's 2026 launch posts report a 'HealthBench Professional (length-adjusted)' variant graded by GPT-5.4; it is not documented publicly and is not tracked here. Conversations are public with a canary string.

完整账本

系统开发者分数日期来源条件
o3OpenAI59.9%论文
Best model at release; mean of 16 runs (Table 7). Paper states o3 also tops HealthBench Hard at 32%.
GPT-4.1OpenAI47.8%论文
Mean of 16 runs (Table 7).
GPT-5 (thinking)OpenAI46.2%厂商自报split: hard reasoning_effort: high
HealthBench Hard only; the announcement gives no full-set number.
GPT-4o (Aug 2024)OpenAI32.3%论文
Mean of 16 runs (Table 7).