BenchCAD

BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD

17,900 execution-verified CadQuery programs across 106 industrial part families (gears, springs, drills, fittings), about half anchored to real ISO/DIN/EN/ASME/IEC specification tables. The prime task, Vision2Code, shows four orthographic renders and asks for a CadQuery program that is re-executed and scored by IoU-score (voxel IoU times execution rate); matched Vision QA, Code QA and Code Edit tasks isolate perception, parametric abstraction and program synthesis. An agentic variant adds a Python sandbox to render, measure and iterate. It is the yardstick vendors cite for AI-for-hardware.

0.843
Claude Fable 5.1 (max, Python tools)
self-reported
2026-06-01 → 2026-09-01: 0.2692 → 0.843
Released
2026-05
Maintainer
BenchCAD team (Rice University, University of Virginia, UC San Diego; Zhang, Liu, Chen et al.)
Status
active
Contamination
medium
Metric
Vision2Code IoU-score (voxel IoU x exec rate) (score, ↑)
Tasks
17,900
Domains
multimodal code reasoning
human
no measured baseline

Notes. Scores are 0-1 (IoU-score), not percentages. The official leaderboard (assets/data/leaderboard.json) marks self-reported vendor rows with an asterisk and re-grades submissions itself; OpenAI's GPT-6 Astra table quotes BenchCAD as 95.9% and notes Anthropic's Claude scores used three modifications to the eval, so launch-table percentages are on a different scale from the site's IoU-score. Dataset (CC-BY-4.0) is public on Hugging Face; a BenchCAD 2.0 with explicit parametric designs is in development at BenchCAD-org/benchcad-2.

Full ledger

SystemDeveloperScoreDateSourceConditions
Claude Fable 5.1 (max, Python tools)Anthropic0.843self-reportedsplit: vision2code-tools tools reasoning_effort: max
Self-reported voxel IoU from the Fable 5.1 / Mythos 5.1 system card (fig. 8.14.2.A) on a random 1,000-file subset, averaged over five runs; republished on the official leaderboard. No-tools 0.437. Highest with-tools figure as of access date.
Grok 4.6 (xhigh, Python sandbox)SpaceXAI0.8055officialsplit: vision2code-tools tools reasoning_effort: xhigh
BenchCAD team's own agentic run on its scorer (run arranged by SpaceXAI); no-tools IoU-score 0.3638. Leaderboard gives 2026-08 only, 1st used.
GPT-5.6 Sol (max)OpenAI0.706self-reportedsplit: vision2code tools: no reasoning_effort: max
Self-reported (asterisk) from OpenAI's GPT-5.6 launch table, republished on the official leaderboard, not re-graded; harness undisclosed. With-tools 0.834. Leaderboard gives 2026-07 only, 1st used.
Gemini 3.1 Pro (thinking)Google0.289officialsplit: vision2code tools: no
Re-graded by the BenchCAD team on the full split; best no-tools row among models it ran itself. Leaderboard gives 'tested 2026-06' only, 1st used.
Claude Opus 4.7 (max)Anthropic0.2692officialsplit: vision2code tools: no reasoning_effort: max
Re-graded by the BenchCAD team; exec rate 96.5%. Leaderboard gives 'tested 2026-06' only, 1st used.