BFCL

Berkeley Function Calling Leaderboard (V4)

Evaluates how accurately a model produces function (tool) calls from natural-language requests. V1 introduced abstract-syntax-tree matching of single, multiple and parallel calls in Python, Java, JavaScript and REST; V2 added user-contributed live functions; V3 added multi-turn, multi-step and state-tracking scenarios; V4 added web-search and memory agentic tasks and format-sensitivity checks. The overall score is an unweighted average of sub-category accuracies, reported with cost and latency. The de-facto standard for tool-calling ability.

77.47%
Claude-Opus-4-5-20251101 (FC)
official
2026-04-12 → 2026-04-12: 31.9% → 77.47%
Released
2024-02
Maintainer
UC Berkeley Gorilla team (Patil, Mao et al.)
Status
active
Contamination
medium
Metric
overall accuracy (percent, ↑)
Tasks
Domains
tool-use
human
no measured baseline

Notes. The 'paper' link is the launch blog (February 2024); the ICML 2025 paper describes V4, and the related Gorilla paper is arXiv 2305.15334. Overall scores changed with each version (V1 to V4), so only compare rows from the same leaderboard version. The official leaderboard was last updated 2026-04-12 and models released after that are absent.

Full ledger

SystemDeveloperScoreDateSourceConditions
Claude-Opus-4-5-20251101 (FC)Anthropic77.47%officialsplit: FC
Rank 1 of 109 on BFCL V4 (multi-turn 68.38%); total run cost $86.55. Newer models (Opus 5, GPT-6 Astra) have not been added since the 2026-04-12 update.
Gemini-3-Pro-Preview (Prompt)Google72.51%official
BFCL V4 overall; FC mode scored 68.14. Snapshot date.
o3-2025-04-16 (Prompt)OpenAI63.05%official
BFCL V4 overall, prompt (non-FC) mode; snapshot date.
GPT-4.1-2025-04-14 (FC)OpenAI53.96%officialsplit: FC
BFCL V4 overall; snapshot date.
Llama-3.3-70B-Instruct (FC)Meta31.9%officialsplit: FC
BFCL V4 overall accuracy, leaderboard snapshot last updated 2026-04-12 (commit f7cf735); date is the snapshot date, not the model's.