π€ AI Summary
This study addresses the difficulty general-purpose agents face in reliably executing complex workflows involving geometric-physical constraints and cross-software dependencies within specialized engineering environments. We construct the first engineering benchmark covering the complete design closed loop, encompassing six domains, 26 software tools, and over a thousand expert-level tasks, to evaluate the end-to-end operational capabilities of frontier models. Methodologically, we integrate GUI and CLI interfaces while deploying domain-specific validator suites, and propose an artifact-centric unified verification framework that enables continuous quantitative scoring of geometric validity, physical feasibility, and specification compliance, overcoming the limitations of traditional binary evaluation. Experiments reveal significant capability gaps: the strongest model achieves only 44.3 points, with multi-software coordination success rates as low as 3.6%, establishing a rigorous baseline for measuring progress in engineering automation.
π Abstract
Autonomous agents have made rapid progress in general-purpose computer use, but reliable automation of professional industrial engineering remains out of reach, as engineering workflows demand reasoning over geometric and physical constraints and dependencies preserved across software and design stages. We present EngiWorld, the first benchmark structured around the complete design loop: 1,301 expert-curated tasks spanning 6 engineering domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, with both GUI and CLI interfaces and 6 task types ranging from software-selection to open-ended tasks. We further introduce an artifact-centric evaluation methodology built on a unified domain-verifier suite, which programmatically checks the geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts, and scores quantitative design tasks continuously by specification attainment rather than binary success. Evaluation of seven frontier models reveals a substantial capability gap: the strongest model achieves an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed. EngiWorld provides the first rigorous foundation for measuring progress toward agents that operate professional engineering software end to end.