LMArena Text
LMArena (Chatbot Arena) Text leaderboard
实时众包的人类偏好排行榜。访问者输入任意提示,并排收到两个匿名模型的回答后投票;数百万成对投票以 Bradley-Terry 模型(最初为在线 Elo)拟合得到 Arena 分数及自举置信区间,另有风格控制与分类视图。提示新鲜且由真实用户评判,难以污染并反映感知有用性,但会奖励有说服力的排版并受模型采样策略影响。
- 发布
- 2023-05
- 维护者
- LMArena (Arena; formerly LMSYS / UC Berkeley)
- 状态
- active
- 污染风险
- low
- 指标
- Arena score (Bradley-Terry Elo scale) (elo, ↑)
- 题量
- —
- 领域
- human-preference general-assistant
- 人类
- 无实测基线
备注. Launched 2023-05-03 as Chatbot Arena (lmsys.org blog); paper arXiv 2403.04132 verified. Ledger rows are the top text model at dated snapshots taken from LMSYS blog tables and the maintainer's leaderboard data files (elo_results_YYYYMMDD.pkl in the lmarena-ai/arena-leaderboard Space, whose last committed snapshot is 2025-08-29). Scores are only comparable within one snapshot: the scale drifts as models are added and the rating system moved from online Elo to Bradley-Terry in December 2023. The live site (lmarena.ai / arena.ai) was unreachable from this environment on 2026-09-04, so no 2026 snapshot is recorded.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| gemini-2.5-pro | 1466 | 官方榜单 | elo_results_20250829.pkl text/full table; 35,586 battles; rounded from 1466.2. Under style control the same snapshot ranks gemini-2.5-pro 1456, gpt-5-high 1447, claude-opus-4-1 thinking 1447. | ||
| gemini-exp-1206 | 1374 | 官方榜单 | elo_results_20250105.pkl text/full table; 18,068 battles. Rounded from 1374.2. | ||
| gpt-4o-2024-05-13 | OpenAI | 1287 | 官方榜单 | elo_results_20240706.pkl text/full table; 55,826 battles. Rounded from 1287.4. | |
| GPT-4 | OpenAI | 1227 | 官方榜单 | Online Elo over 42K votes (Apr 24 - Jun 19, 2023). | |
| GPT-4-Turbo | OpenAI | 1217 | 官方榜单 | First Bradley-Terry (MLE) leaderboard; 7,007 votes for this model, 130K total. |