TruthfulQA
TruthfulQA: Measuring How Models Mimic Human Falsehoods
817 个问题,覆盖 38 个类别(健康、法律、金融、政治、阴谋论),精心设计使部分人类会因常见误解而答错。模型评分方式为自由生成(由微调评判模型或人类评分者判定真实且有信息量)或多项选择变体 MC1 和 MC2(真实答案上的概率质量)。它揭示了模仿性谬误的"逆向缩放",至今仍是事实性与迎合性研究的参考点,尽管题目已被广泛用于训练。
- 发布
- 2021-09
- 维护者
- University of Oxford / OpenAI (Lin, Hilton, Evans)
- 状态
- saturated
- 污染风险
- high
- 指标
- MC2 accuracy (percent, ↑)
- 题量
- 817
- 领域
- factuality safety
- 人类
- 94% one human participant answering the generation task (94% true, 87% true and informative)
备注. The 94% human figure is a single participant on the generation task and is not comparable with MC2. The benchmark was retired from the Open LLM Leaderboard in mid-2024 because of saturation and contamination, and frontier vendors no longer report it, so the ledger holds only the paper baseline.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| GPT-3-175B (helpful prompt) | OpenAI (evaluated by Lin et al.) | 58% | 论文 | split: generation Human-judged % true answers (Fig. 4); the single human participant scored 94%. MC1/MC2 numbers for later models are not comparable with this row. |