TruthfulQA

TruthfulQA: Measuring How Models Mimic Human Falsehoods

817 questions across 38 categories (health, law, finance, politics, conspiracies) crafted so that some humans would answer falsely because of common misconceptions. Models are scored either on free-form generation (judged truthful and informative by a fine-tuned judge or human raters) or on multiple-choice variants MC1 and MC2 (probability mass on true answers). It exposed the 'inverse scaling' of imitative falsehoods and remains a reference point for factuality and sycophancy work, though the questions are widely trained on.

58% · human 94%
GPT-3-175B (helpful prompt)
paper
Released
2021-09
Maintainer
University of Oxford / OpenAI (Lin, Hilton, Evans)
Status
saturated
Contamination
high
Metric
MC2 accuracy (percent, ↑)
Tasks
817
Domains
factuality safety
human
94% one human participant answering the generation task (94% true, 87% true and informative)

Notes. The 94% human figure is a single participant on the generation task and is not comparable with MC2. The benchmark was retired from the Open LLM Leaderboard in mid-2024 because of saturation and contamination, and frontier vendors no longer report it, so the ledger holds only the paper baseline.

Full ledger

SystemDeveloperScoreDateSourceConditions
GPT-3-175B (helpful prompt)OpenAI (evaluated by Lin et al.)58%papersplit: generation
Human-judged % true answers (Fig. 4); the single human participant scored 94%. MC1/MC2 numbers for later models are not comparable with this row.