GAIA

GAIA: a benchmark for General AI Assistants

466 real-world questions (166 public validation, 300 private test) that need web browsing, file handling, multimodal reading and tool use to reach a short, unambiguous answer; scored by quasi-exact match after normalisation. Three difficulty levels, with Level 3 requiring long tool-use chains. Questions are easy for humans, so the gap to human accuracy is the headline signal for general assistants.

93.36% · human 92%
CustomGPT.ai Research Lab v44
official
2023-11-21 → 2026-06-03: 15% → 93.36%
Released
2023-11
Maintainer
Meta FAIR / Hugging Face / AutoGPT (Mialon et al.)
Status
saturating
Contamination
medium
Metric
accuracy (percent, ↑)
Tasks
466
Domains
general-assistant tool-use web reasoning multimodal
human
92% human annotators (validation phase, average over levels)

Notes. Test answers are private but the leaderboard is self-submitted and dominated by multi-model ensembles; top scores are within a few points of the 92% human figure. HAL (hal.cs.princeton.edu/gaia) re-runs open scaffolds on the validation set for independent numbers.

Full ledger

SystemDeveloperScoreDateSourceConditions
CustomGPT.ai Research Lab v44CustomGPT.ai93.36%officialsplit: test scaffold: CustomGPT.ai enterprise agent tools
Top of the test leaderboard as of access date; ensemble of Claude, Gemini and GPT models (self-submitted, auto-scored).
Co-Sight Pro v1.0.1ZTE-AICloud93.02%officialsplit: test scaffold: Co-Sight Pro tools
Ensemble of ZTE Nebula LLM, Gemini 3.1 Pro, GPT 5.5 and Claude Opus 4.7 (self-submitted, auto-scored).
Nemotron-ToolOrchestraNVIDIA87.38%officialsplit: test scaffold: ToolOrchestra (Nemotron-ToolOrchestrator-8B orchestrating GPT-5, Claude Opus 4.1, Qwen2.5-Math-72B) tools
HAL Generalist Agent + Claude Sonnet 4.5Anthropic74.55%independentsplit: validation scaffold: HAL Generalist Agent tools
HAL re-run on the validation set, 2 runs, $178 total cost. Month-only date on the leaderboard; 1st used.
GPT4 + pluginsOpenAI15%papertools
Headline figure from the abstract; Table 4 gives 30.3 / 9.7 / 0 by level.