arXiv · 2608.20576
mLS-GKM: Efficient Multi-class Regulatory Sequence Classification with Gapped k-mer SVMs
Abstract
Gapped k-mer support vector machines (gkm-SVMs) are widely used for classifying regulatory DNA sequences and identifying the sequence features underlying those predictions. Although LS-GKM provides an efficient implementation of gkm-based kernels, it is restricted to binary classification and does not provide calibrated probability outputs. Here, we present mLS-GKM, an extension of LS-GKM that adds multiclass classification, probability-calibrated predictions, parallelised inference, memory efficient sequence interpretation and checkpointing during training. Across 322 ENCODE ChIP-seq datasets, classifiers trained using mLS-GKM are identical to those produced by LS-GKM, while gkmpredict and gkmexplain run 22x and 80x faster respectively at 64 threads, and gkmexplain peak memory is reduced by more than 65%. As a practical demonstration, mLS-GKM was able to accurately distinguish between enhancers, promoters and CTCF binding sites directly from sequence and identified biologically relevant regulatory motifs. Together, these improvements extend gkm-SVMs to multiclass problems and substantially improve their scalability, enabling efficient interpretation of regulatory DNA sequences in large and complex datasets
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kieran Howard, Nathan Harmston. 2026-08-20. mLS-GKM: Efficient Multi-class Regulatory Sequence Classification with Gapped k-mer SVMs. https://arxiv.org/abs/2608.20576
Cite the original work for its findings. Save a collection to share your selection of sources.