ALE-Claw
Every leaderboard run uses ALE-Claw, the harness used in Agent’s Last Exam.
Research · · 12 min read · Seldon Research

AI has transformed how quickly we can build software. We believe the same acceleration is coming to engineering and design as frontier models and especially computer-use agents advance. These industries are responsible for a significant share of the worldwide economy but still constrained by slow, manual workflows. CADBench measures how close agents are to creating reliable, editable designs that can be used in the real world. It is the beginning of a broader effort at Seldon to evaluate and create training data for computer-use agents across PCB design, architecture, chip design, simulation software, visual design and other professional tools.
CADBench measures whether frontier agents can execute native, long-horizon mechanical-design work inside Autodesk Fusion and whether the resulting geometry, feature history and constraints survive deterministic verification.
Fable 5
Opus 5
Gemini 3.8 Flash
Gemini 3.7 Flash
Muse Spark 1.3
Kimi K3
Qwen 3.8 Max
Muse Spark 1.2Sketch → Extrude → Sketch → Extrude → Construction plane → Sketch → Extrude → Sketch → Extrude → Sketch → Extrude → Extrude → Sketch → Extrude → Sketch → Extrude
Sketch → Extrude → Sketch → Sketch
Every CADBench task was authored by a human domain expert and grounded in a real CAD workflow from robotic-arm design to complex assembly. Every verifier is written by hand to model the underlying workflow. Scores combine intermediate process reward with outcome rewards, so a high score requires both sound construction and correct form-factor.
| Model | Total Cost | Cost / task |
|---|---|---|
| $1,998.00 | $28.96 | |
| $1,705.00 | $24.71 | |
| $121.00 | $1.15 | |
| $1,129.00 | $10.75 | |
| $559.00 | $5.32 | |
| $74.00 | $0.70 | |
| $153.00 | $2.22 | |
| $752.00 | $7.16 | |
| $728.00 | $6.93 | |
| Inkling | $105.00 | $1.00 |
| $290.72 | $2.77 | |
| $155.33 | $1.48 |
We ran Gemini 3.7 Flash with five harnesses to select the best one.
| Harness | Pass Rate | Mean verifier score | Reliability | Cost | Cost / task | Tokens / task | Turns / task |
|---|---|---|---|---|---|---|---|
| ALE-Claw Selected | 24.6% | 50.15 | 94.2% | $98.59 | $0.94 | 2.45M | 155 |
| OpenCode | 14.5% | 28.87 | 92.8% | $127.61 | $1.22 | 0.46M | 53 |
| Antigravity | 15.9% | 41.35 | 69.2% | $271.26 | $2.58 | 1.94M | 220 |
| Codex harness | 0.0% | 15.32 | 78.3% | $617.67 | $5.88 | 6.04M | 50 |
| Prime Agent | 0.0% | 17.98 | 78.1% | $638.19 | $6.08 | 35.83M | 224 |
Every leaderboard run uses ALE-Claw, the harness used in Agent’s Last Exam.
Every model is run three times per task. We cap the models at a maximum of 500 turns.
Run at xhigh reasoning through their official API or AI Gateway, with exponential backoff and up to three attempts per request.
Accessed through computer use on a sandbox VM.