AgentBench
AgentBench: Evaluating LLMs as Agents
Eight interactive environments (operating system shell, SQL database, knowledge graph, digital card game, lateral-thinking puzzles, ALFWorld household, WebShop, Mind2Web browsing) in which an LLM acts over multiple turns. Each environment has its own success metric and the overall score is a weighted average normalised across environments, so it measures breadth of agentic decision-making rather than one skill. An early standard for comparing API and open-weight models as agents.
paper · website · leaderboard · dataset · code
- Released
- 2023-08
- Maintainer
- Tsinghua University KEG / Zhipu AI (Liu et al.)
- Status
- retired
- Contamination
- high
- Metric
- overall score (weighted average) (score, ↑)
- Tasks
- —
- Domains
- tool-use reasoning web code
- human
- no measured baseline
Notes. The original v0.2 leaderboard stopped being updated after the paper era; in October 2025 the repository was repurposed as 'AgentBench FC', a function-calling variant on five containerised tasks with its own leaderboard, and scores are not comparable to the original overall score. Marked retired because the maintainers no longer accept results for the original benchmark.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| gpt-4 (0613) | OpenAI | 4.01 | paper | Overall AgentBench score (Table 3); weighted average, not a percentage. | |
| claude-3 opus | Anthropic | 3.11 | paper | Overall score in Table 3 of arXiv v3 (October 2025 revision, which added later models). | |
| gpt-3.5-turbo (0613) | OpenAI | 2.32 | paper | Overall AgentBench score (Table 3). |