Benchmarks / Parametric CAD Bench V3

Parametric CAD Bench V3 (CAD Bench V3)

Built by gNucleus AI · released 21 Sep 2026

100 FreeCAD modeling tasks from gNucleus AI: 30 parts to create from a text spec, 30 parts to create and then edit, and 40 parts to build from an engineering drawing. A coding agent works in a terminal container, and a deterministic geometry scorer grades the FreeCAD model that it produces. V3 (September 2026) is a new series, separate from the frozen V1 and V2 boards.

Frontier

56.9%

Overall score

GPT-6 Astra (max) with Codex 0.154.0 · harness: Codex 0.154.0

21 Sep 2026 · Source: gNucleus AI (benchmark maintainers)

Harbor snapshot of 2026-09-21. Site shows 56.87% ± 6.85; we computed the interval from the ±. Create 52.26, create-and-edit 44.37, image-to-CAD 69.70. Scored 100/100, 5 perfect tasks, $317.81 for the cohort. The maintainers say the top three are not statistically separated.

Emerging

The best result is at 20–60% of the ceiling, or at 50–100% of human parity.

See the full leaderboard at gNucleus AI

Parametric CAD Bench V3: Overall score over time, 3 recorded results. 0%20%40%60%80%100%Sep 2026Sep 2026Sep 2026Sep 2026Oct 2026 Claude Opus 5 (max) with Claude Code 2.1.270: 50.9% (21 Sep 2026) Claude Fable 5.1 (max) with Claude Code 2.1.270: 56.8% (21 Sep 2026) GPT-6 Astra (max) with Codex 0.154.0: 56.9% (21 Sep 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/parametric-cad-bench-v3/"><img src="https://canagentswork.com/og/benchmarks-parametric-cad-bench-v3.png" width="600" height="315" alt="Parametric CAD Bench V3: the best result is 56.9% (GPT-6 Astra (max) with Codex 0.154.0, 21 Sep 2026)." loading="lazy"></a>

Markdown:

[![Parametric CAD Bench V3: the best result is 56.9% (GPT-6 Astra (max) with Codex 0.154.0, 21 Sep 2026).](https://canagentswork.com/og/benchmarks-parametric-cad-bench-v3.png)](https://canagentswork.com/benchmarks/parametric-cad-bench-v3/)

What it measures

Mean continuous task reward over the 100 tasks, shown on a 0 to 100 scale. The verifier (gnucleus-freecad-validator 0.6.0, FreeCAD 1.1.0) compares the produced model with a held-back reference model and the spec: an oriented-bounding-box size gate, per-component comparisons, geometry weights, and spatial alignment. A failed or unscored trial counts as zero. Each official row is one fresh run with one attempt per task, no network except the model provider, and a public Harbor job that anyone can inspect. A 95% confidence interval is computed over the 100 per-task rewards.

Overall score: Mean continuous task reward across the 100 tasks, on a 0 to 100 scale; failed or unscored trials count as zero. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
cli
Grading
automated-tests
Tasks
100
Contamination
Reference geometry stays in a separate verifier image. Agents run with no network except the model provider. In V2 the maintainers saw agents try to find or download the reference CAD or the grader, which is one reason for the stricter V3 isolation.
Reuse
The V3 task suite is distributed through Harbor Hub (gnucleus-ai/cad-bench@v3); we found no license text for it. The reference CAD models are held back in a separate verifier environment. The verifier (freecad-validator) is Apache 2.0. The V1 and V2 results archives on Hugging Face are Apache 2.0. (cite-only)

Limits to keep in mind

  • The top three rows (56.87%, 56.75%, 50.86%) have overlapping 95% confidence intervals, so the maintainers say the order among them is not statistically decisive on 100 tasks. Source
  • Every row is a model plus an agent harness. The same Gemini 3.8 Flash model scored 36.19% with mini-swe-agent and 27.56% with Antigravity, so rows are not clean model rankings. Source
  • All rows so far are the maintainers' own runs of frontier vendors' models. Six create-and-edit tasks scored zero for every row. Source
  • gNucleus AI sells CAD generation models and agents, so the maintainer has a commercial stake in the field it benchmarks. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemOverall scorePerfect tasksDateSource
Claude Opus 5 (max) with Claude Code 2.1.270
harness: Claude Code 2.1.270
50.9%6%21 Sep 2026gNucleus AI · primary
Claude Fable 5.1 (max) with Claude Code 2.1.270
harness: Claude Code 2.1.270
56.8%7%21 Sep 2026gNucleus AI · primary
GPT-6 Astra (max) with Codex 0.154.0 · frontier
harness: Codex 0.154.0
56.9%5%21 Sep 2026gNucleus AI · primary

Timeline

  • 21 Sep 2026 — gNucleus AI releases Parametric CAD Bench V3; GPT-6 Astra with Codex leads at 56.87%. Source

Where it sits in the atlas

Work ladder: Direct evidence for Architecture and engineering. On the work ladder it counts as 7% of tasks fully completed (Perfect tasks). How the ladder works

Work it measures (O*NET work activities): Create visual designs or displays; Design materials or devices; Evaluate designs, specifications, or other technical data.

Architecture and engineering

CAD modeling

Last checked 24 Sep 2026 against 10 primary sources. See an error? Tell us.

How to cite

Credit the original work first: Parametric CAD Bench V3 by gNucleus AI.

Then, if you used this page:

Can Agents Work. "Parametric CAD Bench V3: frontier results and sources." https://canagentswork.com/benchmarks/parametric-cad-bench-v3/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-parametric-cad-bench-v3,
  title        = {{Parametric CAD Bench V3: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/parametric-cad-bench-v3/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}