🤖 AI Summary
This work addresses the problem of efficiently finding maximal exact matches (MEMs) and set-maximal exact matches (SMEMs) in highly repetitive or compressed genomic sequences. The authors propose a compact index structure built upon run-length compressed strings, which for the first time integrates run-length encoding with minimal suffix sets to support efficient MEM/SMEM queries within sublinear space. The approach is specifically optimized for haplotype-aware SMEM retrieval and simulates suffix tree operations to achieve MEM queries in O(ρₚ log m) time with constant-time edge traversal overhead. This significantly enhances the efficiency of sequence alignment directly in the compressed domain, offering a scalable solution for analyzing large, repetitive genomic datasets.
📝 Abstract
We show how to store a text $T [1..n]$ consisting of $ρ_T$ runs in $O (ρ_T + χ)$ space, where $χ$ is the size of the smallest suffixient set for $T$, such that when we are given a pattern $P [1..m]$ consisting of $ρ_P$ runs we can find the maximal exact matches (MEMs) of $P$ with respect to $T$ in $O (ρ_P \log m)$ time plus constant time for each edge we would fully or partially descend in the suffix tree for $T$ while finding those MEMs. We then adapt and optimize our result to finding set-maximal exact matches (SMEMs) of query haplotypes with respect to stored haplotype panels.