RULER
RULER: What's the Real Context Size of Your Long-Context Language Models?
A synthetic long-context suite of 13 tasks in four categories (needle-in-a-haystack retrieval variants, multi-hop variable tracing, aggregation of frequent words, and question answering over long inputs) generated at configurable lengths from 4K to 128K tokens or more. Scored as accuracy averaged across tasks at each length; a model's 'effective context' is the longest length at which it beats a fixed threshold (Llama-2-7B at 4K, 85.6%). It exposes the gap between claimed and usable context windows.
paper · website · leaderboard · code
- Released
- 2024-04
- Maintainer
- NVIDIA (Hsieh et al.)
- Status
- Contamination
- low
- Metric
- accuracy (average over 13 tasks) (percent, ↑)
- Tasks
- 13
- Domains
- long-context
- human
- no measured baseline
Notes. Inputs are generated on the fly, so there is nothing to memorize, but the tasks are synthetic and easier than realistic long-document work; strong models now score above 90% at 128K. The maintainers' table mixes their own runs with author-reported numbers (marked with an asterisk), and closed models (GPT-4, Gemini 1.5 Pro) were last run in 2024.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| Jamba-1.5-large | AI21 Labs | 95.1% | official | split: 128K README main table, author-reported from the Jamba-1.5 report (arXiv 2408.12570); 4K-128K average 96.0, ranked 1st on both weighted averages. Date is the Jamba report's arXiv date. | |
| Gemini-1.5-pro | Google DeepMind | 94.4% | paper | split: 128K Paper Table 3; 4K-128K average 95.8, effective length >128K (top of the paper's table). | |
| Qwen2.5-14B-Instruct-1M | Alibaba | 92.2% | official | split: 128K README main table, author-reported from the Qwen2.5-1M report (arXiv 2501.15383); 4K-128K average 95.7. Date is that report's arXiv date. | |
| Qwen3-235B-A22B | Alibaba | 90.6% | official | split: 128K README main table, author-reported from the Qwen3 technical report (arXiv 2505.09388); 4K-128K average 95.0. Date is that report's arXiv date. | |
| GPT-4 (gpt-4-1106-preview) | OpenAI | 81.2% | paper | split: 128K Paper Table 3; 4K-128K average 91.6, effective length 64K. Llama-2-7B threshold at 4K is 85.6. |