Hi EGAPx team,
We are trying to run EGAPx on a huge genome (approx. 35 Gb). We tried running it on one of the shorter chromosomes (about 0.97 Gb in length) with one pair of PE RNAseq fastq files, it finished without issues.
After we started EGAPx on the whole genome with all of the RNAseq data, it crashed after a few hours. EGAPx is run locally on a workstation, with a Singulariy container, the version is 0.5.2.
Is the problem chromosome length or is it another issue? If the chromosomes are too long, what would be the best way to cut them for the annotation?
Here is the Nextflow report:
Workflow execution completed unsuccessfully!
The exit status of the task that caused the workflow execution to fail was: 3.
The full error message was:
Error executing process > 'egapx:target_proteins_plane:best_aligned_prot:run_best_aligned_prot'
Caused by:
Process `egapx:target_proteins_plane:best_aligned_prot:run_best_aligned_prot` terminated with an error exit status (3)
Command executed:
mkdir -p output
mkdir -p tmp
lds2_indexer -source indexed -db tmp/lds_index
echo "aligns.9.asn
aligns.2.asn
aligns.6.asn
aligns.10.asn
aligns.1.asn
aligns.5.asn
aligns.3.asn
aligns.7.asn
aligns.8.asn
aligns.4.asn" > align.mft
sort align.mft > rm_me.tmp
mv rm_me.tmp align.mft
best_placement -asm_alns_filter 'reciprocity = 3' -lds2 tmp/lds_index -nogenbank -gc_path Prang.HiC-gencoll.asn -in_alns align.mft -out_alns output/align.asn -out_rpt output/report.txt
rm -rf tmp
Command exit status:
3
Command output:
(empty)
Command error:
+ mkdir -p output
+ mkdir -p tmp
+ lds2_indexer -source indexed -db tmp/lds_index
+ echo 'aligns.9.asn
aligns.2.asn
aligns.6.asn
aligns.10.asn
aligns.1.asn
aligns.5.asn
aligns.3.asn
aligns.7.asn
aligns.8.asn
aligns.4.asn'
+ sort align.mft
+ mv rm_me.tmp align.mft
+ best_placement -asm_alns_filter 'reciprocity = 3' -lds2 tmp/lds_index -nogenbank -gc_path Prang.HiC-gencoll.asn -in_alns align.mft -out_alns output/align.asn -out_rpt output/report.txt
Loading input alignments
Loading gencoll annotation set from file
161347 query sequences to process on 323 subjects
Using default scoring model
Starting ranking
Error: (CSeqalignException::eOutOfRange) Can not convert row to seq-interval - invalid from/to value
Error: (106.16) Application's execution failed (CSeqalignException::eOutOfRange) Can not convert row to seq-interval - invalid from/to value
Work dir:
/path/to/egapxWork/8e/568200874f43162c1092d3508244db
Container:
/path/to/software/egapx/egapx_0.5.2.sif
Tip: view the complete command output by changing to the process work dir and entering the command `cat .command.out`
Please let us know if you require any more information.
Thanks
Hi EGAPx team,
We are trying to run EGAPx on a huge genome (approx. 35 Gb). We tried running it on one of the shorter chromosomes (about 0.97 Gb in length) with one pair of PE RNAseq fastq files, it finished without issues.
After we started EGAPx on the whole genome with all of the RNAseq data, it crashed after a few hours. EGAPx is run locally on a workstation, with a Singulariy container, the version is 0.5.2.
Is the problem chromosome length or is it another issue? If the chromosomes are too long, what would be the best way to cut them for the annotation?
Here is the Nextflow report:
Please let us know if you require any more information.
Thanks