EngIntervene: Benchmarking Multimodal Engineering State Understanding and Design Intervention Reasoning

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the deficiency of multimodal large language models in dynamically reasoning about system states, constraints, and intervention consequences within engineering domains. To this end, it constructs a four-tier progressive evaluation benchmark comprising 3,229 questions across seven domains, unifying the representation of objects, relations, and constraints. By introducing atomized scoring criteria alongside visual ablation and supervised fine-tuning experiments, the work compares open-source and closed-source model performance. As the first four-level evaluation framework, it reveals that strong perception does not entail strong reasoning, exposing a significant gap between single-task proficiency and effective intervention. Experiments demonstrate that the best open-source model lags behind closed-source counterparts by 14.7%, removing visual evidence causes severe performance degradation, and fine-tuning improves only basic grounding capabilities without enhancing higher-order diagnostic and interventional reasoning.
📝 Abstract
Multimodal engineering benchmarks largely evaluate static understanding, such as recognizing components, interpreting diagrams, or answering technical questions. This leaves a missing middle between engineering perception and full design generation: whether a model can use an understood system state to reason about relations, constraints, and the consequences of design changes. We introduce \textsc{EngIntervene}, a benchmark for this capability. It contains 3,229 questions across seven engineering domains and organizes evaluation into four levels: state grounding (T1), relational and mechanistic reasoning (T2), constraint-aware diagnosis (T3), and intervention reasoning (T4), which asks whether a proposed modification achieves its target while preserving required constraints. The tasks instantiate a unified engineering state representation spanning objects, relations, constraints, and design objectives, and T2--T4 are scored against structured reference answers with atomic criteria. Across open- and closed-weight multimodal models, stronger grounding does not reliably translate into better diagnosis or intervention, and the best open-weight model trails the best closed model by 14.7 percentage points on the T2--T4 average. Removing or shuffling visual evidence consistently degrades performance, while benchmark-specific supervised fine-tuning improves T1 but not T2--T4. T4 further exposes a large gap between satisfying individual revision criteria and producing a fully valid intervention. Engineering reasoning thus requires not only recovering the current state, but also reliably using it to reason about constraints and post-intervention consequences. Code and benchmark artifacts are available at https://github.com/changcv2021/EngIntervene
Problem

Research questions and friction points this paper is trying to address.

multimodal engineering benchmark
design intervention reasoning
engineering state understanding
constraint-aware diagnosis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Benchmark
Engineering State Representation
Intervention Reasoning
Constraint-aware Diagnosis
Structured Evaluation
J
Jinchang Zhang
Intelligent Vision and Sensing (IVS) Lab, Indiana University Bloomington, Bloomington, IN, USA
Y
Yingda Tao
Intelligent Vision and Sensing (IVS) Lab, Indiana University Bloomington, Bloomington, IN, USA
Jiakai Lin
Jiakai Lin
University of Georgia
Computer Vision
Guoyu Lu
Guoyu Lu
SUNY Binghamton
RoboticsComputer VisionMachine Learning