TruthfulQA

TruthfulQA: Measuring How Models Mimic Human Falsehoods

817 个问题,覆盖 38 个类别(健康、法律、金融、政治、阴谋论),精心设计使部分人类会因常见误解而答错。模型评分方式为自由生成(由微调评判模型或人类评分者判定真实且有信息量)或多项选择变体 MC1 和 MC2(真实答案上的概率质量)。它揭示了模仿性谬误的"逆向缩放",至今仍是事实性与迎合性研究的参考点,尽管题目已被广泛用于训练。

58% · 人类 94%
GPT-3-175B (helpful prompt)
论文
发布
2021-09
维护者
University of Oxford / OpenAI (Lin, Hilton, Evans)
状态
saturated
污染风险
high
指标
MC2 accuracy (percent, ↑)
题量
817
领域
factuality safety
人类
94% one human participant answering the generation task (94% true, 87% true and informative)

备注. The 94% human figure is a single participant on the generation task and is not comparable with MC2. The benchmark was retired from the Open LLM Leaderboard in mid-2024 because of saturation and contamination, and frontier vendors no longer report it, so the ledger holds only the paper baseline.

完整账本

系统开发者分数日期来源条件
GPT-3-175B (helpful prompt)OpenAI (evaluated by Lin et al.)58%论文split: generation
Human-judged % true answers (Fig. 4); the single human participant scored 94%. MC1/MC2 numbers for later models are not comparable with this row.