EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the difficulty general-purpose agents face in reliably executing complex workflows involving geometric-physical constraints and cross-software dependencies within specialized engineering environments. We construct the first engineering benchmark covering the complete design closed loop, encompassing six domains, 26 software tools, and over a thousand expert-level tasks, to evaluate the end-to-end operational capabilities of frontier models. Methodologically, we integrate GUI and CLI interfaces while deploying domain-specific validator suites, and propose an artifact-centric unified verification framework that enables continuous quantitative scoring of geometric validity, physical feasibility, and specification compliance, overcoming the limitations of traditional binary evaluation. Experiments reveal significant capability gaps: the strongest model achieves only 44.3 points, with multi-software coordination success rates as low as 3.6%, establishing a rigorous baseline for measuring progress in engineering automation.
πŸ“ Abstract
Autonomous agents have made rapid progress in general-purpose computer use, but reliable automation of professional industrial engineering remains out of reach, as engineering workflows demand reasoning over geometric and physical constraints and dependencies preserved across software and design stages. We present EngiWorld, the first benchmark structured around the complete design loop: 1,301 expert-curated tasks spanning 6 engineering domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, with both GUI and CLI interfaces and 6 task types ranging from software-selection to open-ended tasks. We further introduce an artifact-centric evaluation methodology built on a unified domain-verifier suite, which programmatically checks the geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts, and scores quantitative design tasks continuously by specification attainment rather than binary success. Evaluation of seven frontier models reveals a substantial capability gap: the strongest model achieves an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed. EngiWorld provides the first rigorous foundation for measuring progress toward agents that operate professional engineering software end to end.
Problem

Research questions and friction points this paper is trying to address.

autonomous agents
engineering automation
benchmark evaluation
frontier models
professional software
Innovation

Methods, ideas, or system contributions that make the work stand out.

Autonomous Agents
Engineering Benchmark
Artifact-centric Evaluation
Domain Verifier
Professional Software Automation
πŸ”Ž Similar Papers
Hongcheng Gao
Hongcheng Gao
University of Chinese Academy of Sciences
Natural Language ProcessingLarge Language ModelsVision Language Models
H
Hailong Qu
Chongqing University
Y
Yu Lei
UCAS
H
Henghui Sun
Shandong University
H
Haoyang Li
BUPT
Y
Yipeng Wei
UCAS
N
Naihao Xue
Fudan University
Xiaohan Yu
Xiaohan Yu
Macquarie University
computer visionsmart farmingultra-fine-grained visual categorization
Zhuo Tao
Zhuo Tao
Institute of Computing Technology, Chinese Academy of Sciences
multi-modal learningvision-and-language
Y
Yihe Zang
Xi’an Jiaotong University
Y
Yajiao Wang
UCAS
J
Jingyi Tang
Peking University
Y
Yi Li
UCAS
J
Jingjing Zhou
UCAS
J
Jie Luo
Tsinghua University
Bohan Zeng
Bohan Zeng
PhD student, Peking University
Data-Centric AIComputer VisionDiffusion Model3D
C
Chengyu Shen
Peking University
Hao Jiang
Hao Jiang
Zhejiang University
C
Chong Chen
Tsinghua University
Bowen Qu
Bowen Qu
Peking University, Ex: Rhymes.ai Aria Team
Multimodal learningVision-Language ModelsComputer Vision
O
Olive Huang
Peking University
Zeqiang Wang
Zeqiang Wang
University of Surrey
Medical AINatural Language Processing