Mixing FM-indexes and CSAs: backward search over an order-1 rank encoding

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high cache-miss rates of FM-indexes over large alphabets and their scalability gap relative to Compressed Suffix Arrays (CSAs) by proposing a hybrid indexing architecture based on first-order rank encoding. The method maps text into a small, skewed alphabet via context frequencies to perform backward search, and subsequently recovers initial character information using a PsiE array for exact counting and locating. This design effectively integrates the cache efficiency of FM-indexes with the binary search advantages of CSAs. Experimental evaluations demonstrate that the prototype system achieves optimal search performance on medium-sized alphabets and noisy repetitive data. However, the compact variant incurs a 1.7- to 3.9-fold increase in space overhead, and its generalizability to real-world scenarios remains to be verified.
📝 Abstract
FM-indexes and compressed suffix arrays (CSAs) are often treated as interchangeable, but they behave differently as the alphabet grows. An FM-index step costs about one cache miss per level of a wavelet tree, so it gets slower with the alphabet size. A CSA step is a binary search whose range shrinks as characters get rarer. We describe a simple hybrid. Each character of the text is replaced by the rank of its frequency among the characters that follow the previous character. We backward-search on this encoding, which is over a small, skewed alphabet, and recover the one piece of information the encoding loses (the first character of the pattern) with a single CSA-like step on an array we call $\PsiE$. Counting is exact, and locating works with standard suffix-array sampling. A prototype on synthetic repetitive data shows that the hybrid is the fastest of the indexes we tried at intermediate alphabet sizes with 1\% noise, but even its compact version is 1.7 to 3.9 times larger than a compressed run-length CSA or FM-index of the original text, because the encoding and $\PsiE$ together have more runs than the original Burrows--Wheeler transform. Whether that changes on real data, such as parses and minimizer digests, is the main open question.
Problem

Research questions and friction points this paper is trying to address.

FM-index
compressed suffix array
backward search
rank encoding
alphabet size
Innovation

Methods, ideas, or system contributions that make the work stand out.

FM-index
Compressed Suffix Array
Backward Search
Order-1 Rank Encoding
Hybrid Index
🔎 Similar Papers
No similar papers found.