MEM-finding with run-length compressed strings

📅 2026-08-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the problem of efficiently finding maximal exact matches (MEMs) and set-maximal exact matches (SMEMs) in highly repetitive or compressed genomic sequences. The authors propose a compact index structure built upon run-length compressed strings, which for the first time integrates run-length encoding with minimal suffix sets to support efficient MEM/SMEM queries within sublinear space. The approach is specifically optimized for haplotype-aware SMEM retrieval and simulates suffix tree operations to achieve MEM queries in O(ρₚ log m) time with constant-time edge traversal overhead. This significantly enhances the efficiency of sequence alignment directly in the compressed domain, offering a scalable solution for analyzing large, repetitive genomic datasets.
📝 Abstract
We show how to store a text $T [1..n]$ consisting of $ρ_T$ runs in $O (ρ_T + χ)$ space, where $χ$ is the size of the smallest suffixient set for $T$, such that when we are given a pattern $P [1..m]$ consisting of $ρ_P$ runs we can find the maximal exact matches (MEMs) of $P$ with respect to $T$ in $O (ρ_P \log m)$ time plus constant time for each edge we would fully or partially descend in the suffix tree for $T$ while finding those MEMs. We then adapt and optimize our result to finding set-maximal exact matches (SMEMs) of query haplotypes with respect to stored haplotype panels.
Problem

Research questions and friction points this paper is trying to address.

maximal exact matches
run-length compressed strings
suffix tree
haplotype panels
set-maximal exact matches
Innovation

Methods, ideas, or system contributions that make the work stand out.

run-length compression
maximal exact matches
suffix tree
haplotype panels
space-efficient indexing
🔎 Similar Papers
No similar papers found.