BFCL
Berkeley Function Calling Leaderboard (V4)
Evaluates how accurately a model produces function (tool) calls from natural-language requests. V1 introduced abstract-syntax-tree matching of single, multiple and parallel calls in Python, Java, JavaScript and REST; V2 added user-contributed live functions; V3 added multi-turn, multi-step and state-tracking scenarios; V4 added web-search and memory agentic tasks and format-sensitivity checks. The overall score is an unweighted average of sub-category accuracies, reported with cost and latency. The de-facto standard for tool-calling ability.
paper · website · leaderboard · dataset · code
- Released
- 2024-02
- Maintainer
- UC Berkeley Gorilla team (Patil, Mao et al.)
- Status
- active
- Contamination
- medium
- Metric
- overall accuracy (percent, ↑)
- Tasks
- —
- Domains
- tool-use
- human
- no measured baseline
Notes. The 'paper' link is the launch blog (February 2024); the ICML 2025 paper describes V4, and the related Gorilla paper is arXiv 2305.15334. Overall scores changed with each version (V1 to V4), so only compare rows from the same leaderboard version. The official leaderboard was last updated 2026-04-12 and models released after that are absent.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| Claude-Opus-4-5-20251101 (FC) | Anthropic | 77.47% | official | split: FC Rank 1 of 109 on BFCL V4 (multi-turn 68.38%); total run cost $86.55. Newer models (Opus 5, GPT-6 Astra) have not been added since the 2026-04-12 update. | |
| Gemini-3-Pro-Preview (Prompt) | 72.51% | official | BFCL V4 overall; FC mode scored 68.14. Snapshot date. | ||
| o3-2025-04-16 (Prompt) | OpenAI | 63.05% | official | BFCL V4 overall, prompt (non-FC) mode; snapshot date. | |
| GPT-4.1-2025-04-14 (FC) | OpenAI | 53.96% | official | split: FC BFCL V4 overall; snapshot date. | |
| Llama-3.3-70B-Instruct (FC) | Meta | 31.9% | official | split: FC BFCL V4 overall accuracy, leaderboard snapshot last updated 2026-04-12 (commit f7cf735); date is the snapshot date, not the model's. |