Brief
CADWorld benchmark tests AI agents on long-horizon FreeCAD design tasks
A benchmark called CADWorld evaluates computer-use agents on long-horizon mechanical CAD tasks in FreeCAD, with success measured by executable checks on saved artifacts. The strongest of seven agents solved 17.5% of tasks, against an 87.0% expert reference pass.
A new benchmark called CADWorld tests computer-use agents on long-horizon mechanical design in FreeCAD. It includes 200 tasks across 11 workflow categories, from sketching and part modeling to assembly, CAM, FEM, measurement, mesh processing, and technical drawing. Agents work only through screenshots and GUI actions, and success is judged by executable checks on saved FreeCAD artifacts and auxiliary outputs.
Across seven current agents, the strongest reached 17.5% success, while an expert reference pass was 87.0%. The authors report that weaker agents often fail before producing a valid artifact, whereas stronger agents more often fail on structural, geometric, and construction-process requirements.
Our reading
Our reading is that the benchmark exposes a gap between general GUI competence and reliable execution of persistent engineering workflows, but these are early results from a single benchmark and not a settled finding.
Source details and supporting facts
Each line is stated by the page named above it.
Stated by arXiv
- CADWorld contains 200 tasks spanning 11 mechanical-CAD workflow categories.
- Across seven current agents on the full benchmark, the strongest agent achieves 17.5% success, compared with an 87.0% expert reference pass.
- Agents operate through screenshots and GUI actions, while success is determined by task-specific executable checks over saved FreeCAD artifacts and auxiliary outputs.
- Weaker agents often fail before producing a valid artifact, whereas stronger agents increasingly fail on structural, geometric, and construction-process requirements.
Sources
- arXivText stored 16 September 2026
How this story was checked. Written from the 1 page listed above, stored 16 September 2026; claims checked against that stored text on 16 September 2026.
What that means
- 4 of 4 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
- Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
- The check reads stored text only: no claim rests on a fresh look that did not happen.
- Where the reporting was silent, the text says so instead of filling the gap.