Hello thanks a lot for the great tool. We are using your tool as part of a pipeline which retrieves proteins from uniprot. We compile the retrieved proteins in a master fasta and perform foldseek easy-search against a database created from fasta files with prostt5. Recently, I realized the output of foldseek easy-search can be different even when the query fasta files have the same proteins but are in a different order because of the upstream protein retrieval scripts. However, the fasta files are the same byte-by-byte (as decided with md5sum after seqkit sort). The difference is up to ~300 very significant (~1e-40) alignments missing from one of the fastas.
We use the latest foldseek (Version: 10.941cd33) with MMseqs Version: 10.941cd33 in an HPC environment with the commands:
# Download the weights
foldseek databases ProstT5 "$FS_dir/weights" tmp_rep2
# Create the database
foldseek createdb mapping_73strains_genomicfp_ec_wslen.fasta "$FS_dir/DB" --prostt5-model "$FS_dir/weights" --gpu 1 --threads 1
# Prepare the database for GPU search
foldseek makepaddedseqdb "$FS_dir/DB" "$FS_dir/DB_pad" --threads 1
# Create the tmp path under scratch and refer to it.
export TMPDIR=/scratch/$USER/foldseek_tmp_rep1_2
mkdir -p $TMPDIR
# Search
foldseek easy-search reproducibility.fasta "$FS_dir/DB_pad" "$FS_dir/rep_FS.m8" $TMPDIR --prostt5-model "$FS_dir/weights" --gpu 1 --threads 1 --format-mode 4 --format-output "qheader,theader,fident,alnlen,evalue,bits" --exhaustive-search 1 -e 0.001
Reason for capping the threading at 1 is described soedinglab/MMseqs2#996.
Can you explain if this is an expected behaviour? If yes, what do you suggest to ensure we get as many hits as possible in a reproducible way? If not, do you have suggestions to avoid it? Happy to share queries if needed. Thanks a lot! Have a nice day.
Hello thanks a lot for the great tool. We are using your tool as part of a pipeline which retrieves proteins from uniprot. We compile the retrieved proteins in a master fasta and perform foldseek easy-search against a database created from fasta files with prostt5. Recently, I realized the output of foldseek easy-search can be different even when the query fasta files have the same proteins but are in a different order because of the upstream protein retrieval scripts. However, the fasta files are the same byte-by-byte (as decided with md5sum after seqkit sort). The difference is up to ~300 very significant (~1e-40) alignments missing from one of the fastas.
We use the latest foldseek (Version: 10.941cd33) with MMseqs Version: 10.941cd33 in an HPC environment with the commands:
Reason for capping the threading at 1 is described soedinglab/MMseqs2#996.
Can you explain if this is an expected behaviour? If yes, what do you suggest to ensure we get as many hits as possible in a reproducible way? If not, do you have suggestions to avoid it? Happy to share queries if needed. Thanks a lot! Have a nice day.