arXiv · 2609.39830
MerKurio: sequence extraction and annotation based on matching k-mers
Abstract
K-mers, short subsequences of length k, play an important role in computational analyses using biological sequence data. A critical downstream processing step involves getting back to the source sequences of selected k-mers for further analysis or validation. Here, we present MerKurio, a high-performance command line tool written in Rust with two main functionalities: 1) extracting sequence records from FASTA/FASTQ files using k-mers and 2) annotating or filtering aligned sequences in SAM/BAM files with k-mer tags. The tool automatically selects the pattern matching algorithm via empirically derived rules based on query characteristics and supports paired-end reads, compressed input files, and reverse complement searching. Benchmark comparisons demonstrate that MerKurio outperforms existing tools in terms of speed while providing additional utility including detailed matching statistics and comprehensive file format support. MerKurio is a fast and user-friendly tool for sequence record extraction and tagging of aligned sequences based on k-mers. MerKurio is freely available at https://github.com/lschoenm/MerKurio.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lukas Schönmann, Heinz Himmelbauer, Juliane C. Dohm. 2026-09-30. MerKurio: sequence extraction and annotation based on matching k-mers. https://arxiv.org/abs/2609.39830
Cite the original work for its findings. Save a collection to share your selection of sources.