GAIA
GAIA: a benchmark for General AI Assistants
466 个真实世界问题(166 个公开验证、300 个私有测试),需要网页浏览、文件处理、多模态阅读和工具使用才能得出简短、无歧义的答案;归一化后按准精确匹配评分。分三个难度等级,Level 3 需要长工具使用链。问题对人类来说很容易,因此与人类准确率的差距是通用助手的主要信号。
- 发布
- 2023-11
- 维护者
- Meta FAIR / Hugging Face / AutoGPT (Mialon et al.)
- 状态
- 污染风险
- medium
- 指标
- accuracy (percent, ↑)
- 题量
- 466
- 领域
- general-assistant tool-use web reasoning multimodal
- 人类
- 92% human annotators (validation phase, average over levels)
备注. Test answers are private but the leaderboard is self-submitted and dominated by multi-model ensembles; top scores are within a few points of the 92% human figure. HAL (hal.cs.princeton.edu/gaia) re-runs open scaffolds on the validation set for independent numbers.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| CustomGPT.ai Research Lab v44 | CustomGPT.ai | 93.36% | 官方榜单 | split: test scaffold: CustomGPT.ai enterprise agent tools Top of the test leaderboard as of access date; ensemble of Claude, Gemini and GPT models (self-submitted, auto-scored). | |
| Co-Sight Pro v1.0.1 | ZTE-AICloud | 93.02% | 官方榜单 | split: test scaffold: Co-Sight Pro tools Ensemble of ZTE Nebula LLM, Gemini 3.1 Pro, GPT 5.5 and Claude Opus 4.7 (self-submitted, auto-scored). | |
| Nemotron-ToolOrchestra | NVIDIA | 87.38% | 官方榜单 | split: test scaffold: ToolOrchestra (Nemotron-ToolOrchestrator-8B orchestrating GPT-5, Claude Opus 4.1, Qwen2.5-Math-72B) tools | |
| HAL Generalist Agent + Claude Sonnet 4.5 | Anthropic | 74.55% | 独立复现 | split: validation scaffold: HAL Generalist Agent tools HAL re-run on the validation set, 2 runs, $178 total cost. Month-only date on the leaderboard; 1st used. | |
| GPT4 + plugins | OpenAI | 15% | 论文 | tools Headline figure from the abstract; Table 4 gives 30.3 / 9.7 / 0 by level. |