RULER

RULER: What's the Real Context Size of Your Long-Context Language Models?

合成长上下文套件,包含四类共 13 个任务(多种大海捞针检索变体、多跳变量追踪、高频词聚合以及长输入问答),可按 4K 到 128K 或更长的 token 长度配置生成。以各长度下跨任务平均准确率评分;模型的"有效上下文"为其超过固定阈值(Llama-2-7B 在 4K 下的 85.6%)的最长长度。它揭示了宣称的与可用的上下文窗口之间的差距。

95.1%
Jamba-1.5-large
官方榜单
2024-04-09 → 2025-05-14: 81.2% → 90.6%
发布
2024-04
维护者
NVIDIA (Hsieh et al.)
状态
saturating
污染风险
low
指标
accuracy (average over 13 tasks) (percent, ↑)
题量
13
领域
long-context
人类
无实测基线

备注. Inputs are generated on the fly, so there is nothing to memorize, but the tasks are synthetic and easier than realistic long-document work; strong models now score above 90% at 128K. The maintainers' table mixes their own runs with author-reported numbers (marked with an asterisk), and closed models (GPT-4, Gemini 1.5 Pro) were last run in 2024.

完整账本

系统开发者分数日期来源条件
Jamba-1.5-largeAI21 Labs95.1%官方榜单split: 128K
README main table, author-reported from the Jamba-1.5 report (arXiv 2408.12570); 4K-128K average 96.0, ranked 1st on both weighted averages. Date is the Jamba report's arXiv date.
Gemini-1.5-proGoogle DeepMind94.4%论文split: 128K
Paper Table 3; 4K-128K average 95.8, effective length >128K (top of the paper's table).
Qwen2.5-14B-Instruct-1MAlibaba92.2%官方榜单split: 128K
README main table, author-reported from the Qwen2.5-1M report (arXiv 2501.15383); 4K-128K average 95.7. Date is that report's arXiv date.
Qwen3-235B-A22BAlibaba90.6%官方榜单split: 128K
README main table, author-reported from the Qwen3 technical report (arXiv 2505.09388); 4K-128K average 95.0. Date is that report's arXiv date.
GPT-4 (gpt-4-1106-preview)OpenAI81.2%论文split: 128K
Paper Table 3; 4K-128K average 91.6, effective length 64K. Llama-2-7B threshold at 4K is 85.6.