HealthBench
HealthBench: Evaluating Large Language Models Towards Improved Human Health
5,000 段模型与普通用户或临床人员之间的真实多轮、多语言健康对话,涵盖急诊转诊、追问上下文、全球健康等七类主题。每段对话附医生撰写的评分细则(来自 60 国 262 名医生的 48,562 条标准);GPT-4.1 评分器判断末轮回复满足哪些标准,得分为所得分值占满分的比例。Hard(1,000)与 Consensus(3,671)子集分别隔离未饱和与医生共识标准。医疗有用性与安全性的参考评估。
- 发布
- 2025-05
- 维护者
- OpenAI
- 状态
- active
- 污染风险
- medium
- 指标
- rubric score (percent, ↑)
- 题量
- 5,000
- 领域
- knowledge safety instruction-following
- 人类
- 无实测基线
备注. Announced 2025-05-12; arXiv 2505.08775 verified (Arora et al., posted 2025-05-13). Scores are 0-1 in the paper and shown here as percent. Physician responses without model help scored below the September 2024 models, and physicians could not improve on o3/GPT-4.1 responses, so no human baseline is recorded. OpenAI's 2026 launch posts report a 'HealthBench Professional (length-adjusted)' variant graded by GPT-5.4; it is not documented publicly and is not tracked here. Conversations are public with a canary string.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| o3 | OpenAI | 59.9% | 论文 | Best model at release; mean of 16 runs (Table 7). Paper states o3 also tops HealthBench Hard at 32%. | |
| GPT-4.1 | OpenAI | 47.8% | 论文 | Mean of 16 runs (Table 7). | |
| GPT-5 (thinking) | OpenAI | 46.2% | 厂商自报 | split: hard reasoning_effort: high HealthBench Hard only; the announcement gives no full-set number. | |
| GPT-4o (Aug 2024) | OpenAI | 32.3% | 论文 | Mean of 16 runs (Table 7). |