AA-Omniscience

AA-Omniscience: Knowledge and Hallucination Benchmark

6,000 short-answer factual questions generated from authoritative academic and industry sources across 42 topics in six domains (business, humanities, health, law, software engineering, science and engineering). An LLM grader labels each answer correct, partial, incorrect or not attempted; the Omniscience Index (-100 to 100) awards correct answers, subtracts incorrect ones and leaves abstentions neutral, so it measures knowledge calibration rather than raw recall. Accuracy and hallucination rate are reported separately and feed the Artificial Analysis Intelligence Index.

44
GPT-6 Astra (high)
independent
2025-11-17 → 2026-09-03: 4.8 → 44
Released
2025-11
Maintainer
Artificial Analysis
Status
active
Contamination
low
Metric
Omniscience Index (score, ↑)
Tasks
6,000
Domains
knowledge factuality
human
no measured baseline

Notes. arXiv 2511.13029 verified (Jackson, Keating, Cameron, Hill-Smith; posted 2025-11-17). Only a public subset of questions is released; the full 6,000 are held privately, hence low contamination risk. Index of 0 means as many correct as incorrect answers. All rows are Artificial Analysis's own runs (independent-evaluation); the current leaderboard shows rounded integer index values for 31 tracked models, while the paper reports one decimal. Methodology: artificialanalysis.ai/methodology/intelligence-benchmarking#aa-omniscience (grader GPT-5.6 Luna medium as of 2026-09).

Full ledger

SystemDeveloperScoreDateSourceConditions
GPT-6 Astra (high)OpenAI44independentreasoning_effort: high
Top of the leaderboard on access date (integer-rounded). AA's 2026-09-03 article attributes the jump over GPT-5.6 Sol to hallucination rate falling from 92% to 51% at max effort.
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic43independentreasoning_effort: max
Leaderboard rounds to integers. AA's 2026-09-01 article reports 67.2% accuracy and a 72.6% hallucination rate for this model; ~4% of tokens served by Anthropic's default fallback models.
GPT-6 Astra (xhigh)OpenAI43independentreasoning_effort: xhigh
Leaderboard rounds to integers.
Claude 4.1 OpusAnthropic4.8paper
Highest score at launch; one of only three models above zero in the paper.