AgentBench

AgentBench: Evaluating LLMs as Agents

包含八个交互式环境(操作系统 shell、SQL 数据库、知识图谱、数字卡牌游戏、横向思维谜题、ALFWorld 家居、WebShop、Mind2Web 网页浏览),LLM 需在其中进行多轮行动。每个环境有自己的成功指标,总分为跨环境归一化后的加权平均,因此衡量的是智能体决策的广度而非单一技能。是早期比较 API 模型与开放权重模型智能体能力的标准之一。

4.01
gpt-4 (0613)
论文
2023-08-07 → 2025-10-04: 2.32 → 3.11
发布
2023-08
维护者
Tsinghua University KEG / Zhipu AI (Liu et al.)
状态
retired
污染风险
high
指标
overall score (weighted average) (score, ↑)
题量
领域
tool-use reasoning web code
人类
无实测基线

备注. The original v0.2 leaderboard stopped being updated after the paper era; in October 2025 the repository was repurposed as 'AgentBench FC', a function-calling variant on five containerised tasks with its own leaderboard, and scores are not comparable to the original overall score. Marked retired because the maintainers no longer accept results for the original benchmark.

完整账本

系统开发者分数日期来源条件
gpt-4 (0613)OpenAI4.01论文
Overall AgentBench score (Table 3); weighted average, not a percentage.
claude-3 opusAnthropic3.11论文
Overall score in Table 3 of arXiv v3 (October 2025 revision, which added later models).
gpt-3.5-turbo (0613)OpenAI2.32论文
Overall AgentBench score (Table 3).