Benchmarks / Parametric CAD Bench V3
Parametric CAD Bench V3 (CAD Bench V3)
Built by gNucleus AI · released 21 Sep 2026
100 FreeCAD modeling tasks from gNucleus AI: 30 parts to create from a text spec, 30 parts to create and then edit, and 40 parts to build from an engineering drawing. A coding agent works in a terminal container, and a deterministic geometry scorer grades the FreeCAD model that it produces. V3 (September 2026) is a new series, separate from the frozen V1 and V2 boards.
Frontier
56.9%
Overall score
GPT-6 Astra (max) with Codex 0.154.0 · harness: Codex 0.154.0
21 Sep 2026 · Source: gNucleus AI (benchmark maintainers)
Harbor snapshot of 2026-09-21. Site shows 56.87% ± 6.85; we computed the interval from the ±. Create 52.26, create-and-edit 44.37, image-to-CAD 69.70. Scored 100/100, 5 perfect tasks, $317.81 for the cohort. The maintainers say the top three are not statistically separated.
The best result is at 20–60% of the ceiling, or at 50–100% of human parity.
What it measures
Mean continuous task reward over the 100 tasks, shown on a 0 to 100 scale. The verifier (gnucleus-freecad-validator 0.6.0, FreeCAD 1.1.0) compares the produced model with a held-back reference model and the spec: an oriented-bounding-box size gate, per-component comparisons, geometry weights, and spatial alignment. A failed or unscored trial counts as zero. Each official row is one fresh run with one attempt per task, no network except the model provider, and a public Harbor job that anyone can inspect. A 95% confidence interval is computed over the 100 per-task rewards.
Overall score: Mean continuous task reward across the 100 tasks, on a 0 to 100 scale; failed or unscored trials count as zero. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- cli
- Grading
- automated-tests
- Tasks
- 100
- Contamination
- Reference geometry stays in a separate verifier image. Agents run with no network except the model provider. In V2 the maintainers saw agents try to find or download the reference CAD or the grader, which is one reason for the stricter V3 isolation.
- Reuse
- The V3 task suite is distributed through Harbor Hub (gnucleus-ai/cad-bench@v3); we found no license text for it. The reference CAD models are held back in a separate verifier environment. The verifier (freecad-validator) is Apache 2.0. The V1 and V2 results archives on Hugging Face are Apache 2.0. (cite-only)
Limits to keep in mind
- The top three rows (56.87%, 56.75%, 50.86%) have overlapping 95% confidence intervals, so the maintainers say the order among them is not statistically decisive on 100 tasks. Source
- Every row is a model plus an agent harness. The same Gemini 3.8 Flash model scored 36.19% with mini-swe-agent and 27.56% with Antigravity, so rows are not clean model rankings. Source
- All rows so far are the maintainers' own runs of frontier vendors' models. Six create-and-edit tasks scored zero for every row. Source
- gNucleus AI sells CAD generation models and agents, so the maintainer has a commercial stake in the field it benchmarks. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Overall score | Perfect tasks | Date | Source |
|---|---|---|---|---|
| Claude Opus 5 (max) with Claude Code 2.1.270 harness: Claude Code 2.1.270 | 50.9% | 6% | 21 Sep 2026 | gNucleus AI · primary |
| Claude Fable 5.1 (max) with Claude Code 2.1.270 harness: Claude Code 2.1.270 | 56.8% | 7% | 21 Sep 2026 | gNucleus AI · primary |
| GPT-6 Astra (max) with Codex 0.154.0 · frontier harness: Codex 0.154.0 | 56.9% | 5% | 21 Sep 2026 | gNucleus AI · primary |
Timeline
- 21 Sep 2026 — gNucleus AI releases Parametric CAD Bench V3; GPT-6 Astra with Codex leads at 56.87%. Source
Where it sits in the atlas
Work ladder: Direct evidence for Architecture and engineering. On the work ladder it counts as 7% of tasks fully completed (Perfect tasks). How the ladder works
Work it measures (O*NET work activities): Create visual designs or displays; Design materials or devices; Evaluate designs, specifications, or other technical data.
Go to the source
- Website cadbench.ai
- Full leaderboard cadbench.ai
- Code github.com
- Announcement cadbench.ai
- Dataset hub.harborframework.com
Last checked 24 Sep 2026 against 10 primary sources. See an error? Tell us.
How to cite
Credit the original work first: Parametric CAD Bench V3 by gNucleus AI.
Then, if you used this page:
Can Agents Work. "Parametric CAD Bench V3: frontier results and sources." https://canagentswork.com/benchmarks/parametric-cad-bench-v3/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-parametric-cad-bench-v3,
title = {{Parametric CAD Bench V3: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/parametric-cad-bench-v3/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}