GAIA
GAIA: a benchmark for General AI Assistants
466 real-world questions (166 public validation, 300 private test) that need web browsing, file handling, multimodal reading and tool use to reach a short, unambiguous answer; scored by quasi-exact match after normalisation. Three difficulty levels, with Level 3 requiring long tool-use chains. Questions are easy for humans, so the gap to human accuracy is the headline signal for general assistants.
paper · website · leaderboard · dataset · code
- Released
- 2023-11
- Maintainer
- Meta FAIR / Hugging Face / AutoGPT (Mialon et al.)
- Status
- Contamination
- medium
- Metric
- accuracy (percent, ↑)
- Tasks
- 466
- Domains
- general-assistant tool-use web reasoning multimodal
- human
- 92% human annotators (validation phase, average over levels)
Notes. Test answers are private but the leaderboard is self-submitted and dominated by multi-model ensembles; top scores are within a few points of the 92% human figure. HAL (hal.cs.princeton.edu/gaia) re-runs open scaffolds on the validation set for independent numbers.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| CustomGPT.ai Research Lab v44 | CustomGPT.ai | 93.36% | official | split: test scaffold: CustomGPT.ai enterprise agent tools Top of the test leaderboard as of access date; ensemble of Claude, Gemini and GPT models (self-submitted, auto-scored). | |
| Co-Sight Pro v1.0.1 | ZTE-AICloud | 93.02% | official | split: test scaffold: Co-Sight Pro tools Ensemble of ZTE Nebula LLM, Gemini 3.1 Pro, GPT 5.5 and Claude Opus 4.7 (self-submitted, auto-scored). | |
| Nemotron-ToolOrchestra | NVIDIA | 87.38% | official | split: test scaffold: ToolOrchestra (Nemotron-ToolOrchestrator-8B orchestrating GPT-5, Claude Opus 4.1, Qwen2.5-Math-72B) tools | |
| HAL Generalist Agent + Claude Sonnet 4.5 | Anthropic | 74.55% | independent | split: validation scaffold: HAL Generalist Agent tools HAL re-run on the validation set, 2 runs, $178 total cost. Month-only date on the leaderboard; 1st used. | |
| GPT4 + plugins | OpenAI | 15% | paper | tools Headline figure from the abstract; Table 4 gives 30.3 / 9.7 / 0 by level. |