🤖 AI Summary
This paper formally defines the low-utility sequential pattern mining (LUSPM) problem, addressing a theoretical and algorithmic gap left by existing high-utility SPM methods, which are not directly applicable to low-utility scenarios. To tackle the high computational complexity and lack of dedicated algorithms for LUSPM, we propose: (i) a redefinition of sequence utility, (ii) a novel sequence-utility chain data structure, and (iii) three algorithms—LUSPM_b, LUSPM_s, and LUSPM_e—based on subsequence contraction and expansion operations. We further introduce the concept of a maximum non-containment sequence set and employ multi-level pruning strategies to significantly improve efficiency. Experimental results demonstrate that LUSPM_s and LUSPM_e substantially outperform baseline methods in both runtime and memory consumption, exhibiting strong scalability; among them, LUSPM_e achieves the best overall performance. The proposed framework is particularly suitable for applications requiring identification of infrequent yet critical behaviors, such as intrusion detection and genomic sequence analysis.
📝 Abstract
Discovering valuable insights from rich data is a crucial task for exploratory data analysis. Sequential pattern mining (SPM) has found widespread applications across various domains. In recent years, low-utility sequential pattern mining (LUSPM) has shown strong potential in applications such as intrusion detection and genomic sequence analysis. However, existing research in utility-based SPM focuses on high-utility sequential patterns, and the definitions and strategies used in high-utility SPM cannot be directly applied to LUSPM. Moreover, no algorithms have yet been developed specifically for mining low-utility sequential patterns. To address these problems, we formalize the LUSPM problem, redefine sequence utility, and introduce a compact data structure called the sequence-utility chain to efficiently record utility information. Furthermore, we propose three novel algorithm--LUSPM_b, LUSPM_s, and LUSPM_e--to discover the complete set of low-utility sequential patterns. LUSPM_b serves as an exhaustive baseline, while LUSPM_s and LUSPM_e build upon it, generating subsequences through shrinkage and extension operations, respectively. In addition, we introduce the maximal non-mutually contained sequence set and incorporate multiple pruning strategies, which significantly reduce redundant operations in both LUSPM_s and LUSPM_e. Finally, extensive experimental results demonstrate that both LUSPM_s and LUSPM_e substantially outperform LUSPM_b and exhibit excellent scalability. Notably, LUSPM_e achieves superior efficiency, requiring less runtime and memory consumption than LUSPM_s. Our code is available at https://github.com/Zhidong-Lin/LUSPM.