lexical pattern discovery

Designs and implements algorithms and pipelines to discover, mine, and generalize recurring lexical and sequence-based patterns — including temporal and usage patterns — from text, token sequences, or event logs, producing motifs, templates, and association rules. Analyzes pattern frequency, ordering, and temporal relationships and outputs generalized pattern representations for further analysis or automated processing.

lexicalpatterndiscovery

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$190K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Mining Frequent Structures in Conceptual Models

Jun 11, 2024
MF
Mattia Fumagalli
🏛️ Free University of Bozen-Bolzano | University of Twente | University of Catania

The absence of automated, systematic approaches for modeling pattern recognition in conceptual modeling hinders improvements in knowledge representation and modeling quality. Method: This paper introduces frequent subgraph mining—systematically applied to conceptual modeling for the first time—and proposes a cross-language (OntoUML/ArchiMate), multi-criteria structural pattern discovery framework. It integrates a gSpan variant with graph editing, graph isomorphism testing, and pattern abstraction techniques to build an extensible, exploratory analysis tool. Contribution/Results: Evaluated on two authoritative datasets, the method successfully identifies highly reusable structural patterns. It demonstrates effectiveness in assessing modeling practices, supporting language evolution, and optimizing model quality—thereby filling a critical research gap in automated pattern mining for conceptual modeling.

Knowledge RepresentationMature MethodStructural Patterns

This study addresses the challenges of redundancy and limited interpretability in mining frequent interaction patterns from spatiotemporal event data. To this end, it proposes modeling events as labeled nodes and representing their spatiotemporal precedence relationships through directed acyclic graphs (DAGs). The work introduces, for the first time, frequent closed embedded sub-DAGs as a compact, non-redundant, and semantically meaningful representation of such patterns. The authors design and implement the DigDag algorithm to efficiently mine these substructures. Experimental results demonstrate that, under identical parameter settings, DigDag significantly outperforms SLEUTH and CSTPM in computational efficiency, while the discovered patterns exhibit clear practical relevance in qualitative analysis.

embedded subgraphsfrequent closed sub-DAGsinteraction patterns

This work addresses the challenge of effectively analyzing massive, heterogeneous high-performance computing (HPC) logs, which hinders fault diagnosis and performance optimization. The authors propose a scalable log analysis workflow that uniquely integrates frequent pattern mining based on finite-state automata with job-level log correlation. By leveraging the Aho–Corasick automaton for efficient pattern storage and matching, and incorporating system hierarchy and message priority information, the approach enables automated detection and clustering of errors and anomalous events. Experiments on an exascale-class supercomputing system demonstrate that the method accurately identifies characteristic error sequences, reveals distinct failure patterns across different applications, and supports real-time, interpretable monitoring to enhance system resilience.

anomaly detectionHPC logslog analysis

Generating Inputs for Grammar Mining using Dynamic Symbolic Execution

Aug 05, 2025
AP
Andreas Pointner
🏛️ University of Applied Sciences Upper Austria | Johannes Kepler University

In grammar reverse-engineering of legacy parsers, insufficient input samples often lead to incomplete grammar coverage—particularly missing edge cases or deprecated features. Method: This paper proposes an automated input generation approach based on dynamic symbolic execution (DSE), the first to apply DSE to grammar mining. We design a three-stage decoupled input generation framework and an iterative expansion strategy to effectively mitigate DSE’s inherent limitations in handling structured inputs. Crucially, our method requires no prior input samples and systematically triggers deep parser behaviors. Results: Evaluated on 11 real-world benchmarks, our generated grammars achieve precision and recall comparable to state-of-the-art methods, while significantly improving detection of subtle semantic features and historical edge-case usages.

Generating diverse inputs for grammar mining automaticallyImproving grammar coverage by capturing edge casesOvercoming limitations of Dynamic Symbolic Execution for parsers

Using Large Language Models for Template Detection from Security Event Logs

Sep 08, 2024
RV
Risto Vaarandi
🏛️ Tallinn University of Technology | Northern Arizona University

This work addresses the critical yet underexplored problem of unsupervised log template mining from security incident logs—specifically, leveraging large language models (LLMs) in a zero-shot, fully unsupervised setting without labeled data or manual rules. We propose a lightweight fine-tuning framework that integrates semantic clustering, dynamic template abstraction, and log-structure priors to guide template extraction. Our method avoids reliance on handcrafted heuristics or supervised signals while preserving interpretability and efficiency. Evaluated across multiple real-world security log datasets, it achieves 92.1% template accuracy—outperforming state-of-the-art unsupervised baselines by an average of 11.3%. Moreover, it significantly improves downstream tasks, including alert compression and anomaly detection. By transcending the limitations of conventional clustering- and regex-based approaches, this work establishes a reproducible, generalizable, LLM-driven unsupervised paradigm for log understanding.

Applying LLMs for unsupervised template miningDetecting templates from unstructured security event logsImproving event log analysis for cybersecurity monitoring

Latest Papers

What's happening recently
View more

This work addresses the high computational cost and poor scalability of maximal frequent episode mining (MaxFEM) under low support thresholds or with long episodes. To overcome these limitations, we propose ParMaxFEM—an algorithm that introduces the first efficient parallelization of MaxFEM, implemented in C++ with multithreaded optimizations and seamlessly integrated into Desbordante, an open-source data profiling platform that treats patterns as first-class citizens for interactive analysis. Experimental results demonstrate that our single-threaded implementation achieves up to 8× speedup over the SPMF baseline, while the 8-core parallel variant attains a maximum acceleration of 35×, substantially enhancing the efficiency and practicality of mining frequent patterns in large-scale sequential data.

computational efficiencyepisode mininghigh-performance algorithm

Efficient Mining of Low-Utility Sequential Patterns

Oct 11, 2025
JZ
Jian Zhu
🏛️ Guangdong University of Technology | Jinan University | Shantou University | University of Illinois Chicago

This paper formally defines the low-utility sequential pattern mining (LUSPM) problem, addressing a theoretical and algorithmic gap left by existing high-utility SPM methods, which are not directly applicable to low-utility scenarios. To tackle the high computational complexity and lack of dedicated algorithms for LUSPM, we propose: (i) a redefinition of sequence utility, (ii) a novel sequence-utility chain data structure, and (iii) three algorithms—LUSPM_b, LUSPM_s, and LUSPM_e—based on subsequence contraction and expansion operations. We further introduce the concept of a maximum non-containment sequence set and employ multi-level pruning strategies to significantly improve efficiency. Experimental results demonstrate that LUSPM_s and LUSPM_e substantially outperform baseline methods in both runtime and memory consumption, exhibiting strong scalability; among them, LUSPM_e achieves the best overall performance. The proposed framework is particularly suitable for applications requiring identification of infrequent yet critical behaviors, such as intrusion detection and genomic sequence analysis.

Addressing the lack of algorithms for low-utility sequential pattern miningDeveloping novel algorithms to discover complete low-utility sequential patternsRedefining sequence utility and introducing compact data structures for efficiency

Can Constructions "SCAN" Compositionality ?

Sep 24, 2025
GK
Ganesh Katrapati
🏛️ International Institute of Information Technology Hyderabad

Sequence-to-sequence models exhibit significant deficiencies in compositional and systematic generalization—particularly under out-of-distribution conditions. To address this, we propose an unsupervised pseudo-construction mining method that requires no architectural modifications or additional annotations; instead, it automatically extracts variable-slot templates from training data to explicitly model form-meaning pairings and enhance structural recombination capability. Our approach integrates pseudo-constructions into the preprocessing pipeline of the SCAN dataset, thereby improving data efficiency and generalization robustness. Experiments demonstrate state-of-the-art performance on the highly challenging ADD JUMP and AROUND RIGHT splits, achieving accuracies of 47.8% and 20.3%, respectively—surpassing most supervised baselines using only 40% of the training data. This work constitutes the first fully unsupervised, construction-level representation mining framework, establishing a novel paradigm for systematic generalization in low-resource settings.

Models lack internalization of form-meaning pairings for productive recombinationSequence to Sequence models fail at compositionality and systematic generalizationUnsupervised method mines pseudo-constructions to improve out-of-distribution performance

This study addresses the prevalent yet underexplored issue of refactorable repetitive step subsequences in Behavior-Driven Development (BDD) tests, for which no automated identification and classification methods previously existed. We propose the first end-to-end framework that leverages Sentence-BERT, UMAP, and HDBSCAN to perform semantic clustering on Gherkin corpora, thereby uncovering recurring fragments. These fragments are then annotated manually to train an XGBoost classifier that ranks their refactoring potential and assigns them to one of three established refactoring patterns. Applying our approach across 339 repositories, we identify 692,020 repetitive patterns and release the first large-scale annotated dataset, along with a complete toolchain and evaluation benchmark. Experimental results demonstrate that our XGBoost classifier achieves an F1 score of 0.891, significantly outperforming rule-based baselines and LLM-based judges, with 75% of test scenarios containing high-potential refactoring candidates.

Behaviour-Driven Developmentextraction-worthinessrefactoring pattern selection

Hot Scholars

PS

Philip S. Yu

Professor of Computer Science, University of Illinons at Chicago
Data miningDatabasePrivacy
YH

Yupeng Hou

University of California, San Diego
Recommender SystemsLarge Language Models
ZZ

Zibin Zheng

IEEE Fellow, Highly Cited Researcher, Sun Yat-sen University, China
BlockchainSmart ContractServices ComputingSoftware Reliability
CT

Christoph Treude

Associate Professor of Computer Science, Singapore Management University
Software EngineeringEmpirical Software EngineeringHuman-AI InteractionAI for Science
JM

Julian McAuley

Professor, UC San Diego
Recommender SystemsNatural Language ProcessingPersonalizationComputer Music