AgentBench

AgentBench: Evaluating LLMs as Agents

Eight interactive environments (operating system shell, SQL database, knowledge graph, digital card game, lateral-thinking puzzles, ALFWorld household, WebShop, Mind2Web browsing) in which an LLM acts over multiple turns. Each environment has its own success metric and the overall score is a weighted average normalised across environments, so it measures breadth of agentic decision-making rather than one skill. An early standard for comparing API and open-weight models as agents.

4.01
gpt-4 (0613)
paper
2023-08-07 → 2025-10-04: 2.32 → 3.11
Released
2023-08
Maintainer
Tsinghua University KEG / Zhipu AI (Liu et al.)
Status
retired
Contamination
high
Metric
overall score (weighted average) (score, ↑)
Tasks
Domains
tool-use reasoning web code
human
no measured baseline

Notes. The original v0.2 leaderboard stopped being updated after the paper era; in October 2025 the repository was repurposed as 'AgentBench FC', a function-calling variant on five containerised tasks with its own leaderboard, and scores are not comparable to the original overall score. Marked retired because the maintainers no longer accept results for the original benchmark.

Full ledger

SystemDeveloperScoreDateSourceConditions
gpt-4 (0613)OpenAI4.01paper
Overall AgentBench score (Table 3); weighted average, not a percentage.
claude-3 opusAnthropic3.11paper
Overall score in Table 3 of arXiv v3 (October 2025 revision, which added later models).
gpt-3.5-turbo (0613)OpenAI2.32paper
Overall AgentBench score (Table 3).