AIME 2025
American Invitational Mathematics Examination 2025 (I and II)
收录 2025 年 AIME I(2025 年 2 月 6 日)与 AIME II(2025 年 2 月 12 日)共 30 道题,每题答案为 000 到 999 的整数。得分为精确答对题目的比例,通常对多次采样取平均(平均 pass@1)。由于试题每年全新编写,对 2025 年 2 月之前训练的模型而言是一次干净的竞赛数学推理测试;使用代码解释器的成绩与无工具运行不可比。
- 发布
- 2025-02
- 维护者
- Mathematical Association of America (exam); evaluated by MathArena and model developers
- 状态
- saturated
- 污染风险
- high
- 指标
- accuracy (percent, ↑)
- 题量
- 30
- 领域
- math reasoning
- 人类
- 无实测基线
备注. Scores are almost always reported on I+II combined (30 problems); a single exam has 15. Problems and solutions have been public since February 2025, so models released later have likely seen them. Several frontier systems report 100% (with or without tools), so the exam no longer separates them; MathArena marks AIME 2025 as deprecated in favour of newer competitions.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| GPT-5.2 (high) | OpenAI | 100% | 独立复现 | tools: no pass_k: 1 MathArena AIME 2025 table, 100.00% +/- 0.00 for both high and xhigh effort. Date is model release; MathArena marks the competition deprecated. | |
| GPT-5 (no tools) | OpenAI | 94.6% | 厂商自报 | tools: no With thinking. GPT-5 with python scored 99.6 in the same chart; tool runs are not comparable. | |
| o3 (high) | OpenAI | 88.9% | 厂商自报 | tools: no GPT-5 launch chart 'AIME 2025 (no tools)'; MathArena independently measured o3 (high) at 89.2. | |
| DeepSeek-R1 | DeepSeek | 70% | 独立复现 | tools: no pass_k: 1 MathArena AIME 2025 table (I+II, 30 problems), 70.00% +/- 8.20. MathArena does not show a run date; date used is AIME II exam day (R1 was released 2025-01-20, before the exam). |