🤖 AI Summary
This study addresses the limitations of existing approaches in semantic accuracy, fragmented evaluation, and practical applicability by systematically reviewing research on large language model (LLM)-based automatic generation of BPMN process models from natural language. Through an analysis of architectural evolution, prompt engineering, intermediate representations, and iterative refinement mechanisms, it elucidates the paradigm shift from traditional NLP to LLM-driven methods. The work proposes a novel pathway integrating retrieval-augmented generation (RAG), interactive modeling, and standardized evaluation, thereby clarifying both the potential and inherent limitations of LLMs in process modeling. Furthermore, it offers a structured roadmap to guide future research in this emerging domain.
📝 Abstract
Recent advances in Generative Artificial Intelligence, particularly Large Language Models (LLMs), have stimulated growing interest in automating or assisting Business Process Modeling tasks using natural language. Several approaches have been proposed to transform textual process descriptions into BPMN and related workflow models. However, the extent to which these approaches effectively support complex process modeling in organizational settings remains unclear. This article presents a literature review of AI-driven methods for transforming natural language into BPMN process models, with a particular focus on the role of LLMs. Following a structured review strategy, relevant studies were identified and analyzed to classify existing approaches, examine how LLMs are integrated into text-to-model pipelines, and investigate the evaluation practices used to assess generated models. The analysis reveals a clear shift from rule-based and traditional NLP pipelines toward LLM-based architectures that rely on prompt engineering, intermediate representations, and iterative refinement mechanisms. While these approaches significantly expand the capabilities of automated process model generation, the literature also exposes persistent challenges related to semantic correctness, evaluation fragmentation, reproducibility, and limited validation in real-world organizational contexts. Based on these findings, this review identifies key research gaps and discusses promising directions for future research, including the integration of contextual knowledge through Retrieval-Augmented Generation (RAG), its integration with LLMs, the development of interactive modeling architectures, and the need for more comprehensive and standardized evaluation frameworks.