GAIA

GAIA: a benchmark for General AI Assistants

466 个真实世界问题(166 个公开验证、300 个私有测试),需要网页浏览、文件处理、多模态阅读和工具使用才能得出简短、无歧义的答案;归一化后按准精确匹配评分。分三个难度等级,Level 3 需要长工具使用链。问题对人类来说很容易,因此与人类准确率的差距是通用助手的主要信号。

93.36% · 人类 92%
CustomGPT.ai Research Lab v44
官方榜单
2023-11-21 → 2026-06-03: 15% → 93.36%
发布
2023-11
维护者
Meta FAIR / Hugging Face / AutoGPT (Mialon et al.)
状态
saturating
污染风险
medium
指标
accuracy (percent, ↑)
题量
466
领域
general-assistant tool-use web reasoning multimodal
人类
92% human annotators (validation phase, average over levels)

备注. Test answers are private but the leaderboard is self-submitted and dominated by multi-model ensembles; top scores are within a few points of the 92% human figure. HAL (hal.cs.princeton.edu/gaia) re-runs open scaffolds on the validation set for independent numbers.

完整账本

系统开发者分数日期来源条件
CustomGPT.ai Research Lab v44CustomGPT.ai93.36%官方榜单split: test scaffold: CustomGPT.ai enterprise agent tools
Top of the test leaderboard as of access date; ensemble of Claude, Gemini and GPT models (self-submitted, auto-scored).
Co-Sight Pro v1.0.1ZTE-AICloud93.02%官方榜单split: test scaffold: Co-Sight Pro tools
Ensemble of ZTE Nebula LLM, Gemini 3.1 Pro, GPT 5.5 and Claude Opus 4.7 (self-submitted, auto-scored).
Nemotron-ToolOrchestraNVIDIA87.38%官方榜单split: test scaffold: ToolOrchestra (Nemotron-ToolOrchestrator-8B orchestrating GPT-5, Claude Opus 4.1, Qwen2.5-Math-72B) tools
HAL Generalist Agent + Claude Sonnet 4.5Anthropic74.55%独立复现split: validation scaffold: HAL Generalist Agent tools
HAL re-run on the validation set, 2 runs, $178 total cost. Month-only date on the leaderboard; 1st used.
GPT4 + pluginsOpenAI15%论文tools
Headline figure from the abstract; Table 4 gives 30.3 / 9.7 / 0 by level.