Back to Search View Original Cite This Article

Abstract

<jats:p> <jats:bold>Motivation</jats:bold> : Novel sequencing technologies (for example, different single-cell chemistries and protocols) produce complex data. Often, the sequenced reads themselves encode critical technical information, such as the cell or molecule of origin. For effective pre-processing of this data and subsequent downstream analysis, it is required to efficiently and accurately identify, extract, and potentially normalize this information. <jats:bold>Results</jats:bold> : We introduce seqproc, a general-purpose sequence pre-processing tool based on a concise descriptive grammar to specify sequence matching and transformations. seqproc compiles a user-provided sequence geometry and transformation description into an execution graph, executed by the ANTISEQUENCE library. We demonstrate that \seqproc is faster on most chemistries, substantially more memory efficient, and at least as accurate as alternative tools that provide similar functionality, while having a more concise description syntax. <jats:bold>Availability</jats:bold> : seqproc is written in Rust and can be executed as a binary program or used as a Rust crate. It is licensed under the BSD 3-clause license and the source code is available at &lt;a href="https://github.com/COMBINE-lab/seqproc"&gt;https://github.com/COMBINE-lab/seqproc&lt;/a&gt;. </jats:p>

Show More

Keywords

seqproc sequence chemistries data information

Related Articles

PORE

About

Connect