BenchCAD

BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD

17,900 个经执行验证的 CadQuery 程序,覆盖 106 个工业零件族(齿轮、弹簧、钻头、管件等),约半数锚定真实 ISO/DIN/EN/ASME/IEC 规范表。主任务 Vision2Code 给出四视图渲染,要求生成 CadQuery 程序并重新执行,以 IoU-score(体素 IoU 乘执行率)评分;配套的 Vision QA、Code QA 和 Code Edit 任务分别隔离感知、参数抽象和程序合成能力。智能体变体提供 Python 沙箱用于渲染、测量和迭代。是各厂商引用的 AI for hardware 标尺。

0.843
Claude Fable 5.1 (max, Python tools)
厂商自报
2026-06-01 → 2026-09-01: 0.2692 → 0.843
发布
2026-05
维护者
BenchCAD team (Rice University, University of Virginia, UC San Diego; Zhang, Liu, Chen et al.)
状态
active
污染风险
medium
指标
Vision2Code IoU-score (voxel IoU x exec rate) (score, ↑)
题量
17,900
领域
multimodal code reasoning
人类
无实测基线

备注. Scores are 0-1 (IoU-score), not percentages. The official leaderboard (assets/data/leaderboard.json) marks self-reported vendor rows with an asterisk and re-grades submissions itself; OpenAI's GPT-6 Astra table quotes BenchCAD as 95.9% and notes Anthropic's Claude scores used three modifications to the eval, so launch-table percentages are on a different scale from the site's IoU-score. Dataset (CC-BY-4.0) is public on Hugging Face; a BenchCAD 2.0 with explicit parametric designs is in development at BenchCAD-org/benchcad-2.

完整账本

系统开发者分数日期来源条件
Claude Fable 5.1 (max, Python tools)Anthropic0.843厂商自报split: vision2code-tools tools reasoning_effort: max
Self-reported voxel IoU from the Fable 5.1 / Mythos 5.1 system card (fig. 8.14.2.A) on a random 1,000-file subset, averaged over five runs; republished on the official leaderboard. No-tools 0.437. Highest with-tools figure as of access date.
Grok 4.6 (xhigh, Python sandbox)SpaceXAI0.8055官方榜单split: vision2code-tools tools reasoning_effort: xhigh
BenchCAD team's own agentic run on its scorer (run arranged by SpaceXAI); no-tools IoU-score 0.3638. Leaderboard gives 2026-08 only, 1st used.
GPT-5.6 Sol (max)OpenAI0.706厂商自报split: vision2code tools: no reasoning_effort: max
Self-reported (asterisk) from OpenAI's GPT-5.6 launch table, republished on the official leaderboard, not re-graded; harness undisclosed. With-tools 0.834. Leaderboard gives 2026-07 only, 1st used.
Gemini 3.1 Pro (thinking)Google0.289官方榜单split: vision2code tools: no
Re-graded by the BenchCAD team on the full split; best no-tools row among models it ran itself. Leaderboard gives 'tested 2026-06' only, 1st used.
Claude Opus 4.7 (max)Anthropic0.2692官方榜单split: vision2code tools: no reasoning_effort: max
Re-graded by the BenchCAD team; exec rate 96.5%. Leaderboard gives 'tested 2026-06' only, 1st used.