BFCL

Berkeley Function Calling Leaderboard (V4)

评估模型根据自然语言请求生成函数(工具)调用的准确度。V1 引入对 Python、Java、JavaScript 和 REST 中单个、多个及并行调用的抽象语法树匹配;V2 增加用户贡献的实时函数;V3 增加多轮、多步和状态跟踪场景;V4 增加网页搜索与记忆类智能体任务及格式敏感性检查。总分为各子类别准确率的无权重平均,并同时报告成本与延迟。是工具调用能力的事实标准。

77.47%
Claude-Opus-4-5-20251101 (FC)
官方榜单
2026-04-12 → 2026-04-12: 31.9% → 77.47%
发布
2024-02
维护者
UC Berkeley Gorilla team (Patil, Mao et al.)
状态
active
污染风险
medium
指标
overall accuracy (percent, ↑)
题量
领域
tool-use
人类
无实测基线

备注. The 'paper' link is the launch blog (February 2024); the ICML 2025 paper describes V4, and the related Gorilla paper is arXiv 2305.15334. Overall scores changed with each version (V1 to V4), so only compare rows from the same leaderboard version. The official leaderboard was last updated 2026-04-12 and models released after that are absent.

完整账本

系统开发者分数日期来源条件
Claude-Opus-4-5-20251101 (FC)Anthropic77.47%官方榜单split: FC
Rank 1 of 109 on BFCL V4 (multi-turn 68.38%); total run cost $86.55. Newer models (Opus 5, GPT-6 Astra) have not been added since the 2026-04-12 update.
Gemini-3-Pro-Preview (Prompt)Google72.51%官方榜单
BFCL V4 overall; FC mode scored 68.14. Snapshot date.
o3-2025-04-16 (Prompt)OpenAI63.05%官方榜单
BFCL V4 overall, prompt (non-FC) mode; snapshot date.
GPT-4.1-2025-04-14 (FC)OpenAI53.96%官方榜单split: FC
BFCL V4 overall; snapshot date.
Llama-3.3-70B-Instruct (FC)Meta31.9%官方榜单split: FC
BFCL V4 overall accuracy, leaderboard snapshot last updated 2026-04-12 (commit f7cf735); date is the snapshot date, not the model's.