🤖 AI Summary
This study addresses the absence of evaluation benchmarks for large language models generating consistent multi-view SysML diagrams by introducing SEMAADB, a dataset comprising 3,000 engineering contexts and 15,000 associated diagrams. As the first large-scale, human-validated benchmark for multi-view SysML consistency, it fills a critical gap in automated modeling evaluation. Methodologically, the work leverages large language models to perform diagram repair and cross-diagram updates, incorporating a semantic consistency verification mechanism. Experimental results demonstrate that the best-performing model achieves a semantic repair rate of 64.3% and a cross-diagram propagation F1 score of 80.7%. These findings reveal that maintaining multi-view semantic consistency remains a fundamental challenge for current approaches.
📝 Abstract
Systems engineers use several diagrams to describe the structure and behavior of systems. Engineers create these diagrams together to make sure that they use the same elements and remain consistent with one another. Large language models can generate diagrams as text or code, which makes it possible to create system diagrams automatically. However, their ability to generate coherent sets of diagrams is not well understood, and existing datasets and benchmarks do not directly measure this ability at scale. We introduce SEMAADB (Systems Engineering Modeling Assistant with AI Dataset and Benchmark), a dataset of 3,000 engineering contexts and 15,000 diagrams. Each context contains five connected SysML views: Requirement, Block Definition, Activity, State Machine, and Sequence. Here, a view is a diagram that presents one aspect of a system. We checked the diagram sets for consistency and valid rendering. A set of 100 contexts is also human-verified and forms the benchmark test set. We evaluate three language models on two tasks. In diagram repair, the strongest model repairs 64.3% of semantic errors . In cross-diagram update the best propagation F1 is 80.7% when a model applies one change across related diagrams. The results show that syntax repair is nearly solved, but semantic repair and consistency across diagrams are still challenging tasks for models. SEMAADB therefore provides both a large diagram resource and a set of benchmarks for measuring coherent multi-diagram SysML generation.