Score
Designs and implements methods to group collections of process models (including local process models) into coherent clusters based on structural and/or behavioral similarity, by defining similarity metrics and clustering algorithms. Builds analyses and tools to partition model sets, reduce redundancy, and evaluate cluster quality for organization, selection, or summarization of models.
Highly variable event logs yield overly complex and poorly interpretable process models via automated discovery, while existing trace clustering methods largely neglect the probabilistic nature of activities and transitions, failing to capture real execution dynamics. This paper proposes a model-driven stochastic trace clustering method: grounded in stochastic process models, it introduces an entropy-based correlation measure derived from direct-follows probabilities and jointly optimizes trace assignment via structural alignment and generative likelihood. An efficient iterative algorithm ensures linear scalability. To our knowledge, this is the first approach to unify stochastic modeling with model-driven optimization in trace clustering, significantly enhancing control-flow pattern clarity and clustering quality. Extensive evaluation on multiple real-world datasets demonstrates superior behavioral representation accuracy and clustering stability over state-of-the-art methods, and reveals systematic effects of stochasticity on clustering performance ranking.
This work addresses the challenge in local process model (LPM) discovery where analysts are often overwhelmed by an explosion of high-scoring yet highly redundant models exhibiting similar structures or behaviors. To mitigate this issue, the authors propose a grouping framework that leverages both model similarity and contextual information from event logs to cluster LPMs, followed by selecting representative models from each cluster to form a concise, non-redundant sample set. Moving beyond conventional approaches that rely solely on scoring metrics for model selection, the proposed method significantly reduces redundancy while preserving—or even enhancing—coverage of the original process behavior. Experimental evaluation on multiple real-life event logs demonstrates that this framework improves both the efficiency and diversity of process understanding.
This work addresses the critical challenge of effectively comparing structural differences between two entity resolution (ER) clustering results in the absence of ground-truth labels. It proposes the Case Count Metric System (CCMS), which introduces and operationalizes, for the first time, a quantitative framework for four types of cluster transformations—preservation, merging, splitting, and overlapping—without requiring labeled data. By leveraging a cluster-set transformation analysis algorithm, CCMS enables fine-grained, unsupervised comparison of ER outcomes. Integrated with interactive analysis and visualization capabilities, the system has been successfully deployed in both academic and industrial settings, significantly enhancing the interpretability and efficiency of ER method evaluation and tuning.
SLURM logs in HPC scientific workflows lack explicit case identifiers, hindering direct application of process mining. Method: This paper proposes an automatic job-correlation method based on implicit job dependency modeling—parsing SLURM logs and jointly leveraging spatiotemporal job feature matching and graph-structured modeling to achieve end-to-end clustering of unannotated jobs. Contribution/Results: We introduce the first systematic preprocessing framework for process mining on HPC logs, integrating algorithms such as Heuristics Miner to support process discovery and bottleneck diagnosis. Evaluated on real-world HPC cluster logs, our approach significantly improves workflow traceability, accurately identifies I/O- and scheduler-related performance bottlenecks, and enables high-fidelity reconstruction of end-to-end process models.
Business process optimization remains challenging due to fragmented methodologies across process mining, predictive process monitoring, and process-aware recommendation—each operating in isolation without a unified theoretical foundation or integration framework. Method: This paper proposes a closed-loop optimization framework that systematically integrates Alpha algorithm/Inductive Miner for process discovery, LSTM/Transformer for runtime prediction, collaborative filtering/graph neural networks for action recommendation, and explainable AI (XAI) for interpretability—enabling automated bottleneck identification, anomaly forecasting, and prescriptive optimization from event logs. Contribution/Results: We establish the first unified conceptual boundary, evolutionary taxonomy, and synergy paradigm across the three domains; construct a comprehensive classification schema covering 120+ studies; clarify application scopes and standardized evaluation benchmarks; and deliver an industrially actionable methodology selection guide with validated deployment pathways.
This work addresses the limitations of traditional case notions in object-centric process mining, which either oversimplify coordination semantics through flattening or become overly complex due to resource objects. To overcome this, the study introduces entity-relationship (ER) modeling into case conceptualization for the first time. By identifying primary entities as process anchors and distinguishing secondary coordinating entities from resource entities, cases are automatically generated as connected components over the induced relationship graph. This approach yields a semantically sound and structurally concise partitioning that supports transitive closure propagation. Evaluated on an OCEL event log comprising 1,000 events, the method successfully produced 40 coherent and semantically clear cases, effectively enabling downstream process discovery and conformance analysis.
This work addresses the challenge of effectively integrating process mining results into early-stage requirements engineering by proposing an automated modeling approach tailored to Use Case Maps (UCMs) within the ITU-T URN standard. By extending the PM4Py library, the authors develop the first process mining pipeline that treats UCMs as first-class outputs, supporting configurable actor mapping and nested hierarchical decomposition. The method enables high-fidelity bidirectional interoperability with the jUCMNav tool. Empirical evaluation on both public and synthetic event logs demonstrates its capability to accurately represent behavioral models across multiple abstraction levels, thereby advancing process mining as a practical enabler for model-driven requirements engineering.
Existing clustering comparison methods struggle to effectively assess the agreement between clustering results containing overlapping clusters and outliers and ground-truth labels, often yielding misleading evaluations due to structural biases. This work presents the first systematic approach to measuring clustering similarity tailored for such complex scenarios, integrating set-matching principles with information-theoretic concepts. The proposed measure is rigorously defined and satisfies several desirable theoretical properties. Comprehensive theoretical analysis and extensive experiments demonstrate its superior robustness and fairness, significantly mitigating the evaluation bias inherent in conventional metrics when confronted with overlapping structures and outliers. This method thus provides a reliable tool for the quantitative comparison of complex clustering outcomes.
Identifying, evaluating, and validating clustering structures in Bayesian mixture models has long been hindered by the complexity of the posterior distribution. This work proposes CliPS, a novel approach that reformulates mixture models as point processes and introduces a low-dimensional parametric functional mapping to transform MCMC samples. By leveraging the separability between the point process representation and the functional mapping, CliPS simultaneously accomplishes cluster identification, assessment of solution quality, and structural validation within the posterior space. The method effectively extracts distinguishable clustering patterns by isolating interpretable features from complex posterior geometries. Extensive experiments on both simulated and real-world datasets demonstrate that CliPS reliably recovers well-separated cluster distributions, confirming its effectiveness and broad applicability across diverse data scenarios.
This study addresses the lack of rigorous statistical assessment for the reliability of output structures in complex clustering pipelines that involve multiple data-dependent stages such as anomaly detection, feature selection, and clustering. To bridge this gap, the work systematically applies selective inference to the entire clustering analysis workflow, establishing a statistical framework that enables valid significance testing of final cluster assignments. The proposed method rigorously controls the type I error rate at any pre-specified nominal level and demonstrates strong empirical performance on both synthetic and real-world datasets. By doing so, it provides a principled and reliable foundation for statistical inference in multi-stage, data-driven clustering procedures.