arXiv Science⌕ Search

arXiv · 2610.10500

Taxonomic Classification with Complete Tag Arrays

Abstract

Taxonomic classifiers such as Kraken assign each $k$-mer of a reference database to the lowest common ancestor (LCA) of the genomes containing it, but this works less well as databases grow, because more and more $k$-mers are shared across species. Cliffy (Ahmed, Boucher and Langmead, 2025) instead uses variable-length exact matches found with an r-index, and can list approximately the genera containing each match; on 16S rRNA it is more accurate than Kraken~2, but its index is large and expensive to build. We present KATKA, which finds the maximal exact matches (MEMs) of at least a given length in each read with Boyer--Moore--Li on a run-length compressed suffix array, counts the occurrences of each MEM in each genus exactly with a complete, run-length compressed tag array, and gives each genus credit in proportion to those counts. On the SILVA 16S rRNA database, KATKA's default index takes 1.44\,GB and can be built in minutes on a desktop computer; it classifies a read in 66\,$μ$s with one thread and reaches 93.8\% genus-level accuracy, close to what Cliffy reports for its 9\,GB index. Grammar-compressing the runs of the tag array shrinks the index to 1.04\,GB, at 75\,$μ$s per read. On the same machine and reads, it is more accurate than Kraken~2 (79.3\%) and Tagger (81.7 to 92.8\%, depending on how mates that disagree are scored). Indexing minimizer digests instead of the sequences makes the index three times smaller and classification 1.7 times faster, at a cost of 1.3 points of accuracy. KATKA is available at https://github.com/TravisGagie/KATKA.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Travis Gagie, Gonzalo Navarro. 2026-10-07. Taxonomic Classification with Complete Tag Arrays. https://arxiv.org/abs/2610.10500

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Almost Optimal Constant-Round Approximation of Dominating Set in Graph Classes with Excluded Minors

For every fixed proper minor-closed class $\mathscr C$ and every $ε>0$, we give a deterministic LOCAL algorithm that returns a dominating set of size at most $(2a(\mathscr C)+1+ε)γ_f(G)$ on every $G\in\mathscr C$. Here $a(\mathscr C)$ is the supremum edge-to-vertex ratio in $\mathscr C$, and $γ_f(G)$ is the fractional domination number. The class also admits a deterministic $(1+ε)$-approximation for fractional dominating set and a randomized algorithm that always returns a dominating set and has expected size at most $(1+ε)γ(G)$. In each case, the number of rounds depends only on $\mathscr C$ and $ε$. None of these algorithms requires the number of vertices or the maximum degree as part of the input. For planar graphs, this gives the deterministic guarantee $(7+ε)γ_f(G)$. Together with the lower bound of Hilke, Lenzen and Suomela, it determines the infimum of the deterministic constant-round approximation ratios for planar minimum dominating set as $7$, settling a question that had remained open since their work. The corresponding infima, measured against the integral optimum, are $7$ for graphs of Euler genus at most any fixed $g\ge0$, $2t-3$ for $K_t$-minor-free graphs with $3\le t\le9$, and $2r+1$ for graphs of treewidth or pathwidth at most any fixed $r\ge1$. We also prove that, for every integer $r\ge1$, no deterministic constant-round LOCAL algorithm achieves an approximation ratio below $2r+1$ on the $r$-th powers of paths, even when every vertex knows the number of vertices. This gives a new proof that the limiting constants are optimal for planar graphs, graphs of bounded treewidth or pathwidth, and $K_t$-minor-free graphs with $3\le t\le9$. For triangle-free planar graphs, the corresponding infimum is $5$.

cs.DS↗

Truly Sub-$3^n$ Min-Sum Subset Convolution and Join Ordering

We present a deterministic reduction from min-sum subset convolution to min-plus matrix product. We show that if the min-plus product of two $D\times D$ matrices with $β$-bit integer entries can be computed in $D^{3-δ}\operatorname{poly}(β,\log D)$ time for a fixed rational $0<δ<1$, then min-sum subset convolution on an $n$-element universe can be solved in $(2+2^{-δ})^n 2^{O(\sqrt n\log(n+1))}\operatorname{poly}(n,β)$ time. Instantiating this reduction with the recent breakthrough on subcubic min-plus matrix product by Alman and Vassilevska Williams gives a Las Vegas algorithm with expected running time $O^*(2.9987^n)$ and a deterministic algorithm with running time $O^*(2.9997^n)$, strictly breaking the longstanding $3^n$ computational barrier. Notably, these speedups translate directly to database query optimization, yielding the same expected and deterministic running-time bounds for join ordering under the $C_{\mathrm{out}}$ cost function.

cs.DS↗

Tight Bounds for Equivalence Testing with Non-Adaptive Conditional Samples

We study distribution testing with access to non-adaptive conditional samples. Specifically, we give tight bounds for equivalence testing, determining whether two unknown distributions are equal to or $\varepsilon$-far from each other in total variation distance. Our algorithm and lower bound show that $\tilde Θ\left(\frac{\log n}{\varepsilon^2}\right)$ queries are necessary and sufficient for this problem. These results demonstrate that the complexity of uniformity, identity, and equivalence testing with non-adaptive conditional samples are all $\tilde Θ(\log n)$.

cs.DS↗