🤖 AI Summary
This study addresses the high cache-miss rates of FM-indexes over large alphabets and their scalability gap relative to Compressed Suffix Arrays (CSAs) by proposing a hybrid indexing architecture based on first-order rank encoding. The method maps text into a small, skewed alphabet via context frequencies to perform backward search, and subsequently recovers initial character information using a PsiE array for exact counting and locating. This design effectively integrates the cache efficiency of FM-indexes with the binary search advantages of CSAs. Experimental evaluations demonstrate that the prototype system achieves optimal search performance on medium-sized alphabets and noisy repetitive data. However, the compact variant incurs a 1.7- to 3.9-fold increase in space overhead, and its generalizability to real-world scenarios remains to be verified.
📝 Abstract
FM-indexes and compressed suffix arrays (CSAs) are often treated as interchangeable, but they behave differently as the alphabet grows. An FM-index step costs about one cache miss per level of a wavelet tree, so it gets slower with the alphabet size. A CSA step is a binary search whose range shrinks as characters get rarer. We describe a simple hybrid. Each character of the text is replaced by the rank of its frequency among the characters that follow the previous character. We backward-search on this encoding, which is over a small, skewed alphabet, and recover the one piece of information the encoding loses (the first character of the pattern) with a single CSA-like step on an array we call $\PsiE$. Counting is exact, and locating works with standard suffix-array sampling. A prototype on synthetic repetitive data shows that the hybrid is the fastest of the indexes we tried at intermediate alphabet sizes with 1\% noise, but even its compact version is 1.7 to 3.9 times larger than a compressed run-length CSA or FM-index of the original text, because the encoding and $\PsiE$ together have more runs than the original Burrows--Wheeler transform. Whether that changes on real data, such as parses and minimizer digests, is the main open question.