arXiv · 1307.3515
QuorUM: an error corrector for Illumina reads
Abstract
Motivation: Illumina Sequencing data can provide high coverage of a genome by relatively short (100 bp150 bp) reads at a low cost. Our goal is to produce trimmed and error-corrected reads to improve genome assemblies. Our error correction procedure aims at producing a set of error-corrected reads (1) minimizing the number of distinct false k-mers, i.e. that are not present in the genome, in the set of reads and (2) maximizing the number that are true, i.e. that are present in the genome. Because coverage of a genome by Illumina reads varies greatly from point to point, we cannot simply eliminate k-mers that occur rarely. Results: Our software, called QuorUM, provides reasonably accurate correction and is suitable for large data sets (1 billion bases checked and corrected per day per core). Availability: QuorUM is distributed as an independent software package and as a module of the MaSuRCA assembly software. Both are available under the GPL open source license at http://www.genome.umd.edu. Contact: gmarcais@umd.edu
Explore related subjects
Keep this discovery
Guillaume Marçais, James A. Yorke, Aleksey Zimin. 2013-07-12. QuorUM: an error corrector for Illumina reads. https://arxiv.org/abs/1307.3515
Cite the original work for its findings. Save a collection to share your selection of sources.