The fasterq-dump utility can extract all reads (i.e., both aligned, unaligned and partially aligned) from an alignment-based SRA record much faster than the C++ API can (by iterating over a ncbi::NGS::openReadCollection). The speed-up appears to be due to the strategy used to parse the SRA record (with only a small fraction of the speed up provided by multi-threading). This strategy appears to be unique to the fasterq-dump utility and not be used by either the C++ API or other SRA utilities (like sam-dump).
For example, using a prefetched SRA record on a four core laptop:
$ time fasterq-dump ERR1138826 --fasta-unsorted
spots read : 4,227,697
reads read : 8,455,394
reads written : 8,455,394
real 0m30.136s
user 1m31.743s
sys 0m3.474s
while the C++ API, sam-dump and using the VDB C API to access reads via the SEQUENCE table all have comparable, and much slower, performance:
$ time sam-dump --fasta --unaligned ERR1138826 > ERR1138826.fasta
real 4m12.311s
user 5m12.168s
sys 0m11.230s
For C++ applications that need to stream all of the reads in an alignment-based SRA record (in any order), it would be very useful to have the fasterq-dump access strategy incorporated in the C++ API!
The
fasterq-dumputility can extract all reads (i.e., both aligned, unaligned and partially aligned) from an alignment-based SRA record much faster than the C++ API can (by iterating over ancbi::NGS::openReadCollection). The speed-up appears to be due to the strategy used to parse the SRA record (with only a small fraction of the speed up provided by multi-threading). This strategy appears to be unique to thefasterq-dumputility and not be used by either the C++ API or other SRA utilities (likesam-dump).For example, using a prefetched SRA record on a four core laptop:
while the C++ API,
sam-dumpand using the VDB C API to access reads via theSEQUENCEtable all have comparable, and much slower, performance:For C++ applications that need to stream all of the reads in an alignment-based SRA record (in any order), it would be very useful to have the
fasterq-dumpaccess strategy incorporated in the C++ API!