🤖 AI Summary
This study addresses the weak adaptability, low accuracy, and lack of autonomous reliability assurance of large language models (LLMs) in knowledge graph (KG)-driven enterprise modeling. It systematically evaluates ChatGPT-4o’s capability to generate enterprise knowledge graphs. Employing a mixed-methods approach—integrating expert surveys with controlled experiments—the work quantitatively and qualitatively identifies, for the first time, LLMs’ capability boundaries and the pronounced degradation of reliability with increasing modeling complexity. Results show that while LLMs consistently perform well on structured subtasks, their accuracy drops significantly in complex semantic modeling. Human expert oversight proves critical for ensuring output completeness and correctness. The study proposes a “human–AI collaborative verification” paradigm, offering a reproducible methodological framework and practical pathway for KG-augmented enterprise modeling.
📝 Abstract
The role of large language models (LLMs) in enterprise modeling has recently started to shift from academic research to that of industrial applications. Thereby, LLMs represent a further building block for the machine-supported generation of enterprise models. In this paper we employ a knowledge graph-based approach for enterprise modeling and investigate the potential benefits of LLMs in this context. In addition, the findings of an expert survey and ChatGPT-4o-based experiments demonstrate that LLM-based model generations exhibit minimal variability, yet remain constrained to specific tasks, with reliability declining for more intricate tasks. The survey results further suggest that the supervision and intervention of human modeling experts are essential to ensure the accuracy and integrity of the generated models.