Score
Designs and implements automated evaluation frameworks and tools that compute and aggregate multiple syntactic, pragmatic, and semantic quality metrics for BPMN process models. These systems produce multi-metric assessments, diagnostic scores and reports, and machine-readable signals (e.g., reward or fitness values) used to compare, select, or guide the generation and analysis of models (including fidelity and variability diagnostics).
Existing visualization tools for compliance checking lack systematic characterization of analytical tasks, hindering rigorous effectiveness evaluation. This paper introduces the first multidimensional task taxonomy specifically designed for compliance checking, modeling core trace-to-model alignment tasks in process mining along six dimensions: objective, method, constraint type, data characteristics, data target, and cardinality. Crucially, this taxonomy explicitly links the semantic requirements of compliance checking with established visual analytics design principles—thereby bridging the semantic gap between process mining and visual analytics. It provides a reusable theoretical framework to rigorously define visualization purposes, evaluate tool effectiveness, and support co-design of analysis systems. As a result, the interpretability and practical utility of complex compliance analysis outcomes are significantly enhanced.
Existing BPMN+DMN process models lack semantic-level automated verification; mainstream tools support only syntactic validation, while behavioral errors require manual execution and debugging, and model transformations remain opaque. Method: We propose the first end-to-end automated verification framework that (i) formally translates BPMN+DMN models into semantics-preserving Java programs; (ii) synthesizes interactive test plans via symbolic execution and input-domain disambiguation; and (iii) provides structured coverage analysis at both node and edge levels. Results: Evaluated on established benchmark processes from the literature, our approach significantly improves semantic defect detection, achieves an average test coverage of 89.3%, and accelerates verification by over 20× compared to manual methods.
This work addresses the challenge of behavioral inconsistency in automatically generated BPMN models due to semantic ambiguity in natural language process descriptions. It proposes the first closed-loop diagnosis and repair framework that operates without requiring ground-truth BPMN annotations. By analyzing the distribution of key performance indicators (KPIs) across multiple model generations, the approach identifies behavioral variations and employs model-based diagnostic techniques to pinpoint gateway logic discrepancies. These discrepancies are traced back to specific source text fragments, which are then refined through an evidence-driven textual revision process. Evaluated on clinical guidelines for diabetic kidney disease management, the method significantly reduces behavioral variability in regenerated models and enhances the semantic stability of executable process models, establishing an end-to-end mapping from behavioral inconsistency to targeted textual correction.
This work addresses the fragility of XML parsing and low editing success rates in natural-language-driven BPMN modeling. We propose an LLM-based process modeling approach leveraging a domain-specific JSON representation. Our core contributions are: (1) a lightweight, semantically explicit BPMN JSON Schema that replaces XML as the LLM’s input/output interface; (2) a quality evaluation framework combining graph edit distance (GED) and normalized GED to quantify structural fidelity, alongside a binary success metric to assess editing reliability. Experiments show that our method achieves process generation similarity comparable to XML-based baselines, while significantly improving editing success rate, inference speed, and robustness. The implementation is open-sourced, establishing a more efficient and resilient paradigm for LLM-powered business process modeling.
This study addresses the lack of systematic evaluation of large language models’ (LLMs) capability to generate BPMN business process models and the neglect of multidimensional quality criteria in existing research. To bridge this gap, we propose BEF4LLM, the first four-dimensional evaluation framework specifically designed for BPMN modeling, which assesses LLM performance along syntactic, pragmatic, semantic, and validity dimensions using standardized metrics and benchmarks against human experts. Experimental results demonstrate that LLMs excel in syntactic and pragmatic aspects, while their semantic quality, though slightly inferior to that of human experts, remains closely comparable. These findings substantiate the practical potential of LLMs in real-world business process modeling tasks.
This study addresses a critical gap in current LLM-driven BPMN modeling tools, which largely overlook human factors and fail to meet the authentic needs of domain experts. Employing a mixed-methods approach that integrates focus groups with standardized usability questionnaires (e.g., CUQ), this work presents the first systematic evaluation of user experience and acceptance of LLM-based process modeling collaborators. The findings reveal a significant tension between moderate usability (mean score: 67.2/100) and low trust (only 48.8%), identifying reliability as a key bottleneck. Key limitations include insufficient output quality and the absence of deep follow-up questioning mechanisms. To address these challenges, the study advocates for integrating human-centered evaluation with automated benchmarking and proposes five practical application scenarios, establishing a new paradigm for trustworthy and effective AI-assisted process modeling.
This study addresses the challenges of automatically reconstructing BPMN models from unstructured natural language descriptions, including specification heterogeneity, multilingual inputs, and the absence of ground-truth references. To overcome these issues, the authors propose a large language model (LLM)-driven, multi-stage automation pipeline that integrates multilingual translation, SpiffWorkflow-based execution validation, and LLM-guided iterative repair to generate high-quality, executable BPMN 2.0 XML ground-truth corpora. A novel multidimensional similarity evaluation framework—combining structural alignment, type distribution, and semantic embeddings—is introduced to enable fully automated, large-scale BPMN generation and refinement without manual intervention. Evaluated on 750 public process diagrams, the approach successfully constructs 387 validated models with an average reconstruction similarity exceeding 0.75, including approximately 50 near-perfect reconstructions differing only in element naming.
This study addresses the lack of a systematic synthesis and cross-directional integration of formal grammars in business process management (BPM). Through a systematic literature review of 34 core studies, it identifies and integrates seven distinct application areas of formal grammars in BPM for the first time, revealing their largely isolated development. Leveraging theoretical foundations such as the Bunge-Wand-Weber ontology, process algebras, graph grammars, attribute grammars, and grammar inference, the work constructs a comprehensive taxonomy that clarifies the role of formal grammars across the entire BPM lifecycle—including process design, modeling, execution, verification, and mining. Furthermore, it articulates five corpus-based open challenges, laying the groundwork for a unified syntactic theory and its deeper integration into BPM research and practice.
This study addresses the limitations of existing approaches in semantic accuracy, fragmented evaluation, and practical applicability by systematically reviewing research on large language model (LLM)-based automatic generation of BPMN process models from natural language. Through an analysis of architectural evolution, prompt engineering, intermediate representations, and iterative refinement mechanisms, it elucidates the paradigm shift from traditional NLP to LLM-driven methods. The work proposes a novel pathway integrating retrieval-augmented generation (RAG), interactive modeling, and standardized evaluation, thereby clarifying both the potential and inherent limitations of LLMs in process modeling. Furthermore, it offers a structured roadmap to guide future research in this emerging domain.
This study addresses the limitations of existing autonomous business process execution approaches, which predominantly focus on control-flow constraints and struggle to support compliance-aware decision-making under multifaceted requirements involving data-aware and temporal conditions. To overcome this gap, the work introduces a unified multi-perspective framework that formally integrates data and time constraints through a numeric planning-based modeling approach. This enables efficient what-if analysis and optimal continuation recommendations for partially executed processes. Experimental results demonstrate that the proposed method not only ensures regulatory compliance but also exhibits strong scalability, substantially enhancing the effectiveness and practicality of autonomous decision-making in AI-augmented business process management systems.