RULER
RULER: What's the Real Context Size of Your Long-Context Language Models?
合成长上下文套件,包含四类共 13 个任务(多种大海捞针检索变体、多跳变量追踪、高频词聚合以及长输入问答),可按 4K 到 128K 或更长的 token 长度配置生成。以各长度下跨任务平均准确率评分;模型的"有效上下文"为其超过固定阈值(Llama-2-7B 在 4K 下的 85.6%)的最长长度。它揭示了宣称的与可用的上下文窗口之间的差距。
- 发布
- 2024-04
- 维护者
- NVIDIA (Hsieh et al.)
- 状态
- 污染风险
- low
- 指标
- accuracy (average over 13 tasks) (percent, ↑)
- 题量
- 13
- 领域
- long-context
- 人类
- 无实测基线
备注. Inputs are generated on the fly, so there is nothing to memorize, but the tasks are synthetic and easier than realistic long-document work; strong models now score above 90% at 128K. The maintainers' table mixes their own runs with author-reported numbers (marked with an asterisk), and closed models (GPT-4, Gemini 1.5 Pro) were last run in 2024.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| Jamba-1.5-large | AI21 Labs | 95.1% | 官方榜单 | split: 128K README main table, author-reported from the Jamba-1.5 report (arXiv 2408.12570); 4K-128K average 96.0, ranked 1st on both weighted averages. Date is the Jamba report's arXiv date. | |
| Gemini-1.5-pro | Google DeepMind | 94.4% | 论文 | split: 128K Paper Table 3; 4K-128K average 95.8, effective length >128K (top of the paper's table). | |
| Qwen2.5-14B-Instruct-1M | Alibaba | 92.2% | 官方榜单 | split: 128K README main table, author-reported from the Qwen2.5-1M report (arXiv 2501.15383); 4K-128K average 95.7. Date is that report's arXiv date. | |
| Qwen3-235B-A22B | Alibaba | 90.6% | 官方榜单 | split: 128K README main table, author-reported from the Qwen3 technical report (arXiv 2505.09388); 4K-128K average 95.0. Date is that report's arXiv date. | |
| GPT-4 (gpt-4-1106-preview) | OpenAI | 81.2% | 论文 | split: 128K Paper Table 3; 4K-128K average 91.6, effective length 64K. Llama-2-7B threshold at 4K is 85.6. |