BenchCAD
BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD
17,900 个经执行验证的 CadQuery 程序,覆盖 106 个工业零件族(齿轮、弹簧、钻头、管件等),约半数锚定真实 ISO/DIN/EN/ASME/IEC 规范表。主任务 Vision2Code 给出四视图渲染,要求生成 CadQuery 程序并重新执行,以 IoU-score(体素 IoU 乘执行率)评分;配套的 Vision QA、Code QA 和 Code Edit 任务分别隔离感知、参数抽象和程序合成能力。智能体变体提供 Python 沙箱用于渲染、测量和迭代。是各厂商引用的 AI for hardware 标尺。
- 发布
- 2026-05
- 维护者
- BenchCAD team (Rice University, University of Virginia, UC San Diego; Zhang, Liu, Chen et al.)
- 状态
- active
- 污染风险
- medium
- 指标
- Vision2Code IoU-score (voxel IoU x exec rate) (score, ↑)
- 题量
- 17,900
- 领域
- multimodal code reasoning
- 人类
- 无实测基线
备注. Scores are 0-1 (IoU-score), not percentages. The official leaderboard (assets/data/leaderboard.json) marks self-reported vendor rows with an asterisk and re-grades submissions itself; OpenAI's GPT-6 Astra table quotes BenchCAD as 95.9% and notes Anthropic's Claude scores used three modifications to the eval, so launch-table percentages are on a different scale from the site's IoU-score. Dataset (CC-BY-4.0) is public on Hugging Face; a BenchCAD 2.0 with explicit parametric designs is in development at BenchCAD-org/benchcad-2.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| Claude Fable 5.1 (max, Python tools) | Anthropic | 0.843 | 厂商自报 | split: vision2code-tools tools reasoning_effort: max Self-reported voxel IoU from the Fable 5.1 / Mythos 5.1 system card (fig. 8.14.2.A) on a random 1,000-file subset, averaged over five runs; republished on the official leaderboard. No-tools 0.437. Highest with-tools figure as of access date. | |
| Grok 4.6 (xhigh, Python sandbox) | SpaceXAI | 0.8055 | 官方榜单 | split: vision2code-tools tools reasoning_effort: xhigh BenchCAD team's own agentic run on its scorer (run arranged by SpaceXAI); no-tools IoU-score 0.3638. Leaderboard gives 2026-08 only, 1st used. | |
| GPT-5.6 Sol (max) | OpenAI | 0.706 | 厂商自报 | split: vision2code tools: no reasoning_effort: max Self-reported (asterisk) from OpenAI's GPT-5.6 launch table, republished on the official leaderboard, not re-graded; harness undisclosed. With-tools 0.834. Leaderboard gives 2026-07 only, 1st used. | |
| Gemini 3.1 Pro (thinking) | 0.289 | 官方榜单 | split: vision2code tools: no Re-graded by the BenchCAD team on the full split; best no-tools row among models it ran itself. Leaderboard gives 'tested 2026-06' only, 1st used. | ||
| Claude Opus 4.7 (max) | Anthropic | 0.2692 | 官方榜单 | split: vision2code tools: no reasoning_effort: max Re-graded by the BenchCAD team; exec rate 96.5%. Leaderboard gives 'tested 2026-06' only, 1st used. |