LMArena Text
LMArena (Chatbot Arena) Text leaderboard
Live crowdsourced human-preference ranking of chat models. Visitors type any prompt, receive answers from two anonymous models side by side and vote; millions of pairwise votes are fit with a Bradley-Terry model (originally online Elo) to produce an Arena score with bootstrap confidence intervals, plus style-controlled and category views. Because prompts are fresh and judged by real users it is hard to contaminate and captures perceived helpfulness, but it rewards persuasive formatting and depends on which models are sampled.
paper · website · leaderboard · dataset · code
- Released
- 2023-05
- Maintainer
- LMArena (Arena; formerly LMSYS / UC Berkeley)
- Status
- active
- Contamination
- low
- Metric
- Arena score (Bradley-Terry Elo scale) (elo, ↑)
- Tasks
- —
- Domains
- human-preference general-assistant
- human
- no measured baseline
Notes. Launched 2023-05-03 as Chatbot Arena (lmsys.org blog); paper arXiv 2403.04132 verified. Ledger rows are the top text model at dated snapshots taken from LMSYS blog tables and the maintainer's leaderboard data files (elo_results_YYYYMMDD.pkl in the lmarena-ai/arena-leaderboard Space, whose last committed snapshot is 2025-08-29). Scores are only comparable within one snapshot: the scale drifts as models are added and the rating system moved from online Elo to Bradley-Terry in December 2023. The live site (lmarena.ai / arena.ai) was unreachable from this environment on 2026-09-04, so no 2026 snapshot is recorded.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| gemini-2.5-pro | 1466 | official | elo_results_20250829.pkl text/full table; 35,586 battles; rounded from 1466.2. Under style control the same snapshot ranks gemini-2.5-pro 1456, gpt-5-high 1447, claude-opus-4-1 thinking 1447. | ||
| gemini-exp-1206 | 1374 | official | elo_results_20250105.pkl text/full table; 18,068 battles. Rounded from 1374.2. | ||
| gpt-4o-2024-05-13 | OpenAI | 1287 | official | elo_results_20240706.pkl text/full table; 55,826 battles. Rounded from 1287.4. | |
| GPT-4 | OpenAI | 1227 | official | Online Elo over 42K votes (Apr 24 - Jun 19, 2023). | |
| GPT-4-Turbo | OpenAI | 1217 | official | First Bradley-Terry (MLE) leaderboard; 7,007 votes for this model, 130K total. |