AA-Omniscience

AA-Omniscience: Knowledge and Hallucination Benchmark

6,000 道由权威学术与行业来源生成的短答事实题,覆盖商业、人文社科、健康、法律、软件工程、科学工程数学六大领域 42 个主题。LLM 评分器将答案标为正确、部分正确、错误或未作答;Omniscience Index(-100 至 100)对正确加分、错误扣分、弃答不计,衡量知识校准而非单纯记忆。准确率与幻觉率另行报告并计入 Artificial Analysis Intelligence Index。

44
GPT-6 Astra (high)
独立复现
2025-11-17 → 2026-09-03: 4.8 → 44
发布
2025-11
维护者
Artificial Analysis
状态
active
污染风险
low
指标
Omniscience Index (score, ↑)
题量
6,000
领域
knowledge factuality
人类
无实测基线

备注. arXiv 2511.13029 verified (Jackson, Keating, Cameron, Hill-Smith; posted 2025-11-17). Only a public subset of questions is released; the full 6,000 are held privately, hence low contamination risk. Index of 0 means as many correct as incorrect answers. All rows are Artificial Analysis's own runs (independent-evaluation); the current leaderboard shows rounded integer index values for 31 tracked models, while the paper reports one decimal. Methodology: artificialanalysis.ai/methodology/intelligence-benchmarking#aa-omniscience (grader GPT-5.6 Luna medium as of 2026-09).

完整账本

系统开发者分数日期来源条件
GPT-6 Astra (high)OpenAI44独立复现reasoning_effort: high
Top of the leaderboard on access date (integer-rounded). AA's 2026-09-03 article attributes the jump over GPT-5.6 Sol to hallucination rate falling from 92% to 51% at max effort.
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic43独立复现reasoning_effort: max
Leaderboard rounds to integers. AA's 2026-09-01 article reports 67.2% accuracy and a 72.6% hallucination rate for this model; ~4% of tokens served by Anthropic's default fallback models.
GPT-6 Astra (xhigh)OpenAI43独立复现reasoning_effort: xhigh
Leaderboard rounds to integers.
Claude 4.1 OpusAnthropic4.8论文
Highest score at launch; one of only three models above zero in the paper.