TruthfulQA
TruthfulQA: Measuring How Models Mimic Human Falsehoods
817 questions across 38 categories (health, law, finance, politics, conspiracies) crafted so that some humans would answer falsely because of common misconceptions. Models are scored either on free-form generation (judged truthful and informative by a fine-tuned judge or human raters) or on multiple-choice variants MC1 and MC2 (probability mass on true answers). It exposed the 'inverse scaling' of imitative falsehoods and remains a reference point for factuality and sycophancy work, though the questions are widely trained on.
- Released
- 2021-09
- Maintainer
- University of Oxford / OpenAI (Lin, Hilton, Evans)
- Status
- saturated
- Contamination
- high
- Metric
- MC2 accuracy (percent, ↑)
- Tasks
- 817
- Domains
- factuality safety
- human
- 94% one human participant answering the generation task (94% true, 87% true and informative)
Notes. The 94% human figure is a single participant on the generation task and is not comparable with MC2. The benchmark was retired from the Open LLM Leaderboard in mid-2024 because of saturation and contamination, and frontier vendors no longer report it, so the ledger holds only the paper baseline.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| GPT-3-175B (helpful prompt) | OpenAI (evaluated by Lin et al.) | 58% | paper | split: generation Human-judged % true answers (Fig. 4); the single human participant scored 94%. MC1/MC2 numbers for later models are not comparable with this row. |