Score
Designs and implements algorithms and pipelines to discover, mine, and generalize recurring lexical and sequence-based patterns — including temporal and usage patterns — from text, token sequences, or event logs, producing motifs, templates, and association rules. Analyzes pattern frequency, ordering, and temporal relationships and outputs generalized pattern representations for further analysis or automated processing.
The absence of automated, systematic approaches for modeling pattern recognition in conceptual modeling hinders improvements in knowledge representation and modeling quality. Method: This paper introduces frequent subgraph mining—systematically applied to conceptual modeling for the first time—and proposes a cross-language (OntoUML/ArchiMate), multi-criteria structural pattern discovery framework. It integrates a gSpan variant with graph editing, graph isomorphism testing, and pattern abstraction techniques to build an extensible, exploratory analysis tool. Contribution/Results: Evaluated on two authoritative datasets, the method successfully identifies highly reusable structural patterns. It demonstrates effectiveness in assessing modeling practices, supporting language evolution, and optimizing model quality—thereby filling a critical research gap in automated pattern mining for conceptual modeling.
This study addresses the challenges of redundancy and limited interpretability in mining frequent interaction patterns from spatiotemporal event data. To this end, it proposes modeling events as labeled nodes and representing their spatiotemporal precedence relationships through directed acyclic graphs (DAGs). The work introduces, for the first time, frequent closed embedded sub-DAGs as a compact, non-redundant, and semantically meaningful representation of such patterns. The authors design and implement the DigDag algorithm to efficiently mine these substructures. Experimental results demonstrate that, under identical parameter settings, DigDag significantly outperforms SLEUTH and CSTPM in computational efficiency, while the discovered patterns exhibit clear practical relevance in qualitative analysis.
This work addresses the challenge of effectively analyzing massive, heterogeneous high-performance computing (HPC) logs, which hinders fault diagnosis and performance optimization. The authors propose a scalable log analysis workflow that uniquely integrates frequent pattern mining based on finite-state automata with job-level log correlation. By leveraging the Aho–Corasick automaton for efficient pattern storage and matching, and incorporating system hierarchy and message priority information, the approach enables automated detection and clustering of errors and anomalous events. Experiments on an exascale-class supercomputing system demonstrate that the method accurately identifies characteristic error sequences, reveals distinct failure patterns across different applications, and supports real-time, interpretable monitoring to enhance system resilience.
In grammar reverse-engineering of legacy parsers, insufficient input samples often lead to incomplete grammar coverage—particularly missing edge cases or deprecated features. Method: This paper proposes an automated input generation approach based on dynamic symbolic execution (DSE), the first to apply DSE to grammar mining. We design a three-stage decoupled input generation framework and an iterative expansion strategy to effectively mitigate DSE’s inherent limitations in handling structured inputs. Crucially, our method requires no prior input samples and systematically triggers deep parser behaviors. Results: Evaluated on 11 real-world benchmarks, our generated grammars achieve precision and recall comparable to state-of-the-art methods, while significantly improving detection of subtle semantic features and historical edge-case usages.
This work addresses the critical yet underexplored problem of unsupervised log template mining from security incident logs—specifically, leveraging large language models (LLMs) in a zero-shot, fully unsupervised setting without labeled data or manual rules. We propose a lightweight fine-tuning framework that integrates semantic clustering, dynamic template abstraction, and log-structure priors to guide template extraction. Our method avoids reliance on handcrafted heuristics or supervised signals while preserving interpretability and efficiency. Evaluated across multiple real-world security log datasets, it achieves 92.1% template accuracy—outperforming state-of-the-art unsupervised baselines by an average of 11.3%. Moreover, it significantly improves downstream tasks, including alert compression and anomaly detection. By transcending the limitations of conventional clustering- and regex-based approaches, this work establishes a reproducible, generalizable, LLM-driven unsupervised paradigm for log understanding.
This work addresses the high computational cost and poor scalability of maximal frequent episode mining (MaxFEM) under low support thresholds or with long episodes. To overcome these limitations, we propose ParMaxFEM—an algorithm that introduces the first efficient parallelization of MaxFEM, implemented in C++ with multithreaded optimizations and seamlessly integrated into Desbordante, an open-source data profiling platform that treats patterns as first-class citizens for interactive analysis. Experimental results demonstrate that our single-threaded implementation achieves up to 8× speedup over the SPMF baseline, while the 8-core parallel variant attains a maximum acceleration of 35×, substantially enhancing the efficiency and practicality of mining frequent patterns in large-scale sequential data.
This paper formally defines the low-utility sequential pattern mining (LUSPM) problem, addressing a theoretical and algorithmic gap left by existing high-utility SPM methods, which are not directly applicable to low-utility scenarios. To tackle the high computational complexity and lack of dedicated algorithms for LUSPM, we propose: (i) a redefinition of sequence utility, (ii) a novel sequence-utility chain data structure, and (iii) three algorithms—LUSPM_b, LUSPM_s, and LUSPM_e—based on subsequence contraction and expansion operations. We further introduce the concept of a maximum non-containment sequence set and employ multi-level pruning strategies to significantly improve efficiency. Experimental results demonstrate that LUSPM_s and LUSPM_e substantially outperform baseline methods in both runtime and memory consumption, exhibiting strong scalability; among them, LUSPM_e achieves the best overall performance. The proposed framework is particularly suitable for applications requiring identification of infrequent yet critical behaviors, such as intrusion detection and genomic sequence analysis.
Sequence-to-sequence models exhibit significant deficiencies in compositional and systematic generalization—particularly under out-of-distribution conditions. To address this, we propose an unsupervised pseudo-construction mining method that requires no architectural modifications or additional annotations; instead, it automatically extracts variable-slot templates from training data to explicitly model form-meaning pairings and enhance structural recombination capability. Our approach integrates pseudo-constructions into the preprocessing pipeline of the SCAN dataset, thereby improving data efficiency and generalization robustness. Experiments demonstrate state-of-the-art performance on the highly challenging ADD JUMP and AROUND RIGHT splits, achieving accuracies of 47.8% and 20.3%, respectively—surpassing most supervised baselines using only 40% of the training data. This work constitutes the first fully unsupervised, construction-level representation mining framework, establishing a novel paradigm for systematic generalization in low-resource settings.
This study addresses the prevalent yet underexplored issue of refactorable repetitive step subsequences in Behavior-Driven Development (BDD) tests, for which no automated identification and classification methods previously existed. We propose the first end-to-end framework that leverages Sentence-BERT, UMAP, and HDBSCAN to perform semantic clustering on Gherkin corpora, thereby uncovering recurring fragments. These fragments are then annotated manually to train an XGBoost classifier that ranks their refactoring potential and assigns them to one of three established refactoring patterns. Applying our approach across 339 repositories, we identify 692,020 repetitive patterns and release the first large-scale annotated dataset, along with a complete toolchain and evaluation benchmark. Experimental results demonstrate that our XGBoost classifier achieves an F1 score of 0.891, significantly outperforming rule-based baselines and LLM-based judges, with 75% of test scenarios containing high-potential refactoring candidates.