🤖 AI Summary
This study addresses the decline in classification accuracy caused by k-mer sharing and the high cost of index construction in large reference databases by proposing KATKA. Built upon a compressed suffix array and a run-length compressed full label array, this method achieves weighted taxonomic classification by statistically evaluating the frequencies of maximal exact matches (MEMs) across genera. Furthermore, it integrates grammar-based compression to substantially reduce both index size and construction overhead. Experimental evaluations on the SILVA database demonstrate that KATKA attains a genus-level accuracy of 93.8%, with an index footprint of merely 1.44 GB that can be constructed within minutes. Additionally, single-sequence queries require only 66 μs. Overall, KATKA exhibits superior comprehensive performance compared to mainstream tools such as Kraken2.
📝 Abstract
Taxonomic classifiers such as Kraken assign each $k$-mer of a reference database to the lowest common ancestor (LCA) of the genomes containing it, but this works less well as databases grow, because more and more $k$-mers are shared across species. Cliffy (Ahmed, Boucher and Langmead, 2025) instead uses variable-length exact matches found with an r-index, and can list approximately the genera containing each match; on 16S rRNA it is more accurate than Kraken~2, but its index is large and expensive to build. We present KATKA, which finds the maximal exact matches (MEMs) of at least a given length in each read with Boyer--Moore--Li on a run-length compressed suffix array, counts the occurrences of each MEM in each genus exactly with a complete, run-length compressed tag array, and gives each genus credit in proportion to those counts. On the SILVA 16S rRNA database, KATKA's default index takes 1.44\,GB and can be built in minutes on a desktop computer; it classifies a read in 66\,$μ$s with one thread and reaches 93.8\% genus-level accuracy, close to what Cliffy reports for its 9\,GB index. Grammar-compressing the runs of the tag array shrinks the index to 1.04\,GB, at 75\,$μ$s per read. On the same machine and reads, it is more accurate than Kraken~2 (79.3\%) and Tagger (81.7 to 92.8\%, depending on how mates that disagree are scored). Indexing minimizer digests instead of the sequences makes the index three times smaller and classification 1.7 times faster, at a cost of 1.3 points of accuracy. KATKA is available at https://github.com/TravisGagie/KATKA.