Research · · 12 min read · Seldon Research

How good are agents actually at CAD?

White concept sketch of a low sports car on a deep indigo field.

AI has transformed how quickly we can build software. We believe the same acceleration is coming to engineering and design as frontier models and especially computer-use agents advance. These industries are responsible for a significant share of the worldwide economy but still constrained by slow, manual workflows. CADBench measures how close agents are to creating reliable, editable designs that can be used in the real world. It is the beginning of a broader effort at Seldon to evaluate and create training data for computer-use agents across PCB design, architecture, chip design, simulation software, visual design and other professional tools.

tasks
105
unsolved tasks
68.1%
best pass rate
30.0%

Models

CADBench measures whether frontier agents can execute native, long-horizon mechanical-design work inside Autodesk Fusion and whether the resulting geometry, feature history and constraints survive deterministic verification.

RankModelMean verifier score
01
AnthropicFable 5
57.86
02
AnthropicOpus 5
56.95
03
GoogleGemini 3.8 Flash
54.71
04
OpenAIGPT-5.6 Sol
51.22
05
GoogleGemini 3.7 Flash
50.15
06
OpenAIGPT-5.6 Terra
25.84
07
MetaMuse Spark 1.3
20.66
08
Moonshot AIKimi K3
17.17
09
AlibabaQwen 3.8 Max
17.04
10
xAIGrok 4.6
16.73
11
MetaMuse Spark 1.2
13.44
12
Inkling
6.75

Tasks

TASK
MODEL
Loading task…
Verifier23.24 / 100
Final body geometry
16.67%7 checks
Geometry checks1 / 7 passed
Analytic face types
Body is valid
Bounding dimensions
Centroid alignment
Physical properties
Surface distance
Body topology
Requested construction
49.52%19 checks
Sketch4 / 7 passed
Circle count
Circle diameters
Fully constrained
Global center
Feature is healthy
Feature is present
Sketch support
Extrusion6 / 9 passed
Extrusion direction
Extrusion distance
Distance extent
Feature is healthy
Boolean operation
Feature is present
References new sketch
Selected profile count
Taper angle
Feature sequence0 / 3 passed
Exact feature order
New features are healthy
New features are unsuppressed
EXPECTED FEATURE SEQUENCE

Sketch → Extrude → Sketch → Extrude → Construction plane → Sketch → Extrude → Sketch → Extrude → Sketch → Extrude → Extrude → Sketch → Extrude → Sketch → Extrude

ACTUAL FEATURE SEQUENCE

Sketch → Extrude → Sketch → Sketch

5.77Mreported tokens
29m 06swall time
TASK 1Task promptGPT-5.6 Sol · xhigh
VERIFIER SCORE23.24 / 100

Action-compressed recording. The video advances only when a new UI state was captured. Model thinking, API latency and unchanged time are omitted; this run took 29m 06s in real time.

Every CADBench task was authored by a human domain expert and grounded in a real CAD workflow from robotic-arm design to complex assembly. Every verifier is written by hand to model the underlying workflow. Scores combine intermediate process reward with outcome rewards, so a high score requires both sound construction and correct form-factor.

Analysis

Higher score · lower cost020406080$0.50$1.00$3.00$10.00$30.00Cost per task · log scaleMean verifier scoreFable 5Opus 5Gemini 3.8GPT-5.6 SolGemini 3.7GPT-5.6 TerraMuse Spark 1.3Kimi K3Qwen 3.8 MaxGrok 4.6Muse Spark 1.2InklingFable 5 — $28.96 · 57.86 mean verifier scoreOpus 5 — $24.71 · 56.95 mean verifier scoreGemini 3.8 Flash — $2.77 · 54.71 mean verifier scoreGPT-5.6 Sol — $3.93 · 51.22 mean verifier scoreGemini 3.7 Flash — $0.53 · 50.15 mean verifier scoreGPT-5.6 Terra — $2.22 · 25.84 mean verifier scoreMuse Spark 1.3 — $1.48 · 20.66 mean verifier scoreKimi K3 — $2.82 · 17.17 mean verifier scoreQwen 3.8 Max — $3.19 · 17.04 mean verifier scoreGrok 4.6 — $2.21 · 16.73 mean verifier scoreMuse Spark 1.2 — $0.67 · 13.44 mean verifier scoreInkling — $0.39 · 6.75 mean verifier score
ModelTotal CostCost / task
AnthropicFable 5$1,998.00$28.96
AnthropicOpus 5$1,705.00$24.71
GoogleGemini 3.7 Flash$121.00$1.15
OpenAIGPT-5.6 Sol$1,129.00$10.75
xAIGrok 4.6$559.00$5.32
MetaMuse Spark 1.2$74.00$0.70
OpenAIGPT-5.6 Terra$153.00$2.22
Moonshot AIKimi K3$752.00$7.16
AlibabaQwen 3.8 Max$728.00$6.93
Inkling$105.00$1.00
GoogleGemini 3.8 Flash$290.72$2.77
MetaMuse Spark 1.3$155.33$1.48
The cost axis uses a logarithmic scale so lower-cost runs remain distinguishable beside the Claude models.

Harness choice

We ran Gemini 3.7 Flash with five harnesses to select the best one.

HarnessPass RateMean verifier scoreReliabilityCostCost / taskTokens / taskTurns / task
ALE-Claw Selected24.6%50.1594.2%$98.59$0.942.45M155
OpenCode14.5%28.8792.8%$127.61$1.220.46M53
Antigravity15.9%41.3569.2%$271.26$2.581.94M220
Codex harness0.0%15.3278.3%$617.67$5.886.04M50
Prime Agent0.0%17.9878.1%$638.19$6.0835.83M224

Methodology

COST LIMITS

500 turns per run

Every model is run three times per task. We cap the models at a maximum of 500 turns.

MODEL CONFIG

Twelve frontier models

Run at xhigh reasoning through their official API or AI Gateway, with exponential backoff and up to three attempts per request.

ENVIRONMENT

Autodesk Fusion · Windows 11

Accessed through computer use on a sandbox VM.

Stay posted if new models are added