AA-Omniscience
AA-Omniscience: Knowledge and Hallucination Benchmark
6,000 道由权威学术与行业来源生成的短答事实题,覆盖商业、人文社科、健康、法律、软件工程、科学工程数学六大领域 42 个主题。LLM 评分器将答案标为正确、部分正确、错误或未作答;Omniscience Index(-100 至 100)对正确加分、错误扣分、弃答不计,衡量知识校准而非单纯记忆。准确率与幻觉率另行报告并计入 Artificial Analysis Intelligence Index。
- 发布
- 2025-11
- 维护者
- Artificial Analysis
- 状态
- active
- 污染风险
- low
- 指标
- Omniscience Index (score, ↑)
- 题量
- 6,000
- 领域
- knowledge factuality
- 人类
- 无实测基线
备注. arXiv 2511.13029 verified (Jackson, Keating, Cameron, Hill-Smith; posted 2025-11-17). Only a public subset of questions is released; the full 6,000 are held privately, hence low contamination risk. Index of 0 means as many correct as incorrect answers. All rows are Artificial Analysis's own runs (independent-evaluation); the current leaderboard shows rounded integer index values for 31 tracked models, while the paper reports one decimal. Methodology: artificialanalysis.ai/methodology/intelligence-benchmarking#aa-omniscience (grader GPT-5.6 Luna medium as of 2026-09).
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| GPT-6 Astra (high) | OpenAI | 44 | 独立复现 | reasoning_effort: high Top of the leaderboard on access date (integer-rounded). AA's 2026-09-03 article attributes the jump over GPT-5.6 Sol to hallucination rate falling from 92% to 51% at max effort. | |
| Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Anthropic | 43 | 独立复现 | reasoning_effort: max Leaderboard rounds to integers. AA's 2026-09-01 article reports 67.2% accuracy and a 72.6% hallucination rate for this model; ~4% of tokens served by Anthropic's default fallback models. | |
| GPT-6 Astra (xhigh) | OpenAI | 43 | 独立复现 | reasoning_effort: xhigh Leaderboard rounds to integers. | |
| Claude 4.1 Opus | Anthropic | 4.8 | 论文 | Highest score at launch; one of only three models above zero in the paper. |