🤖 AI Summary
This study addresses the deficiency of multimodal large language models in dynamically reasoning about system states, constraints, and intervention consequences within engineering domains. To this end, it constructs a four-tier progressive evaluation benchmark comprising 3,229 questions across seven domains, unifying the representation of objects, relations, and constraints. By introducing atomized scoring criteria alongside visual ablation and supervised fine-tuning experiments, the work compares open-source and closed-source model performance. As the first four-level evaluation framework, it reveals that strong perception does not entail strong reasoning, exposing a significant gap between single-task proficiency and effective intervention. Experiments demonstrate that the best open-source model lags behind closed-source counterparts by 14.7%, removing visual evidence causes severe performance degradation, and fine-tuning improves only basic grounding capabilities without enhancing higher-order diagnostic and interventional reasoning.
📝 Abstract
Multimodal engineering benchmarks largely evaluate static understanding, such as recognizing components, interpreting diagrams, or answering technical questions. This leaves a missing middle between engineering perception and full design generation: whether a model can use an understood system state to reason about relations, constraints, and the consequences of design changes. We introduce \textsc{EngIntervene}, a benchmark for this capability. It contains 3,229 questions across seven engineering domains and organizes evaluation into four levels: state grounding (T1), relational and mechanistic reasoning (T2), constraint-aware diagnosis (T3), and intervention reasoning (T4), which asks whether a proposed modification achieves its target while preserving required constraints. The tasks instantiate a unified engineering state representation spanning objects, relations, constraints, and design objectives, and T2--T4 are scored against structured reference answers with atomic criteria. Across open- and closed-weight multimodal models, stronger grounding does not reliably translate into better diagnosis or intervention, and the best open-weight model trails the best closed model by 14.7 percentage points on the T2--T4 average. Removing or shuffling visual evidence consistently degrades performance, while benchmark-specific supervised fine-tuning improves T1 but not T2--T4. T4 further exposes a large gap between satisfying individual revision criteria and producing a fully valid intervention. Engineering reasoning thus requires not only recovering the current state, but also reliably using it to reason about constraints and post-intervention consequences. Code and benchmark artifacts are available at https://github.com/changcv2021/EngIntervene