Skip to content

Feature request: increase speed of C++ API for streaming reads from alignment-based SRA records #794

Description

@jgans

The fasterq-dump utility can extract all reads (i.e., both aligned, unaligned and partially aligned) from an alignment-based SRA record much faster than the C++ API can (by iterating over a ncbi::NGS::openReadCollection). The speed-up appears to be due to the strategy used to parse the SRA record (with only a small fraction of the speed up provided by multi-threading). This strategy appears to be unique to the fasterq-dump utility and not be used by either the C++ API or other SRA utilities (like sam-dump).

For example, using a prefetched SRA record on a four core laptop:

$ time fasterq-dump ERR1138826 --fasta-unsorted
spots read      : 4,227,697
reads read      : 8,455,394
reads written   : 8,455,394

real 0m30.136s
user 1m31.743s
sys 0m3.474s

while the C++ API, sam-dump and using the VDB C API to access reads via the SEQUENCE table all have comparable, and much slower, performance:

$ time sam-dump --fasta --unaligned ERR1138826 > ERR1138826.fasta
real 4m12.311s
user 5m12.168s
sys 0m11.230s

For C++ applications that need to stream all of the reads in an alignment-based SRA record (in any order), it would be very useful to have the fasterq-dump access strategy incorporated in the C++ API!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions