AA-Omniscience
AA-Omniscience: Knowledge and Hallucination Benchmark
6,000 short-answer factual questions generated from authoritative academic and industry sources across 42 topics in six domains (business, humanities, health, law, software engineering, science and engineering). An LLM grader labels each answer correct, partial, incorrect or not attempted; the Omniscience Index (-100 to 100) awards correct answers, subtracts incorrect ones and leaves abstentions neutral, so it measures knowledge calibration rather than raw recall. Accuracy and hallucination rate are reported separately and feed the Artificial Analysis Intelligence Index.
paper · website · leaderboard · dataset
- Released
- 2025-11
- Maintainer
- Artificial Analysis
- Status
- active
- Contamination
- low
- Metric
- Omniscience Index (score, ↑)
- Tasks
- 6,000
- Domains
- knowledge factuality
- human
- no measured baseline
Notes. arXiv 2511.13029 verified (Jackson, Keating, Cameron, Hill-Smith; posted 2025-11-17). Only a public subset of questions is released; the full 6,000 are held privately, hence low contamination risk. Index of 0 means as many correct as incorrect answers. All rows are Artificial Analysis's own runs (independent-evaluation); the current leaderboard shows rounded integer index values for 31 tracked models, while the paper reports one decimal. Methodology: artificialanalysis.ai/methodology/intelligence-benchmarking#aa-omniscience (grader GPT-5.6 Luna medium as of 2026-09).
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| GPT-6 Astra (high) | OpenAI | 44 | independent | reasoning_effort: high Top of the leaderboard on access date (integer-rounded). AA's 2026-09-03 article attributes the jump over GPT-5.6 Sol to hallucination rate falling from 92% to 51% at max effort. | |
| Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Anthropic | 43 | independent | reasoning_effort: max Leaderboard rounds to integers. AA's 2026-09-01 article reports 67.2% accuracy and a 72.6% hallucination rate for this model; ~4% of tokens served by Anthropic's default fallback models. | |
| GPT-6 Astra (xhigh) | OpenAI | 43 | independent | reasoning_effort: xhigh Leaderboard rounds to integers. | |
| Claude 4.1 Opus | Anthropic | 4.8 | paper | Highest score at launch; one of only three models above zero in the paper. |