LMArena Text

LMArena (Chatbot Arena) Text leaderboard

实时众包的人类偏好排行榜。访问者输入任意提示,并排收到两个匿名模型的回答后投票;数百万成对投票以 Bradley-Terry 模型(最初为在线 Elo)拟合得到 Arena 分数及自举置信区间,另有风格控制与分类视图。提示新鲜且由真实用户评判,难以污染并反映感知有用性,但会奖励有说服力的排版并受模型采样策略影响。

1466
gemini-2.5-pro
官方榜单
2023-06-22 → 2025-08-29: 1227 → 1466
发布
2023-05
维护者
LMArena (Arena; formerly LMSYS / UC Berkeley)
状态
active
污染风险
low
指标
Arena score (Bradley-Terry Elo scale) (elo, ↑)
题量
领域
human-preference general-assistant
人类
无实测基线

备注. Launched 2023-05-03 as Chatbot Arena (lmsys.org blog); paper arXiv 2403.04132 verified. Ledger rows are the top text model at dated snapshots taken from LMSYS blog tables and the maintainer's leaderboard data files (elo_results_YYYYMMDD.pkl in the lmarena-ai/arena-leaderboard Space, whose last committed snapshot is 2025-08-29). Scores are only comparable within one snapshot: the scale drifts as models are added and the rating system moved from online Elo to Bradley-Terry in December 2023. The live site (lmarena.ai / arena.ai) was unreachable from this environment on 2026-09-04, so no 2026 snapshot is recorded.

完整账本

系统开发者分数日期来源条件
gemini-2.5-proGoogle1466官方榜单
elo_results_20250829.pkl text/full table; 35,586 battles; rounded from 1466.2. Under style control the same snapshot ranks gemini-2.5-pro 1456, gpt-5-high 1447, claude-opus-4-1 thinking 1447.
gemini-exp-1206Google1374官方榜单
elo_results_20250105.pkl text/full table; 18,068 battles. Rounded from 1374.2.
gpt-4o-2024-05-13OpenAI1287官方榜单
elo_results_20240706.pkl text/full table; 55,826 battles. Rounded from 1287.4.
GPT-4OpenAI1227官方榜单
Online Elo over 42K votes (Apr 24 - Jun 19, 2023).
GPT-4-TurboOpenAI1217官方榜单
First Bradley-Terry (MLE) leaderboard; 7,007 votes for this model, 130K total.