Clustering tool optimized for sparse distance matrices: ~16M elements with ~190M connections clustered in minutes.
-
Updated
Sep 12, 2026 - C++
Clustering tool optimized for sparse distance matrices: ~16M elements with ~190M connections clustered in minutes.
CD HIT cluster file parser
Representative sequence selection for large bioinformatics datasets
Community builds for the main bioinformatics software - Samtools, Tabix, Bcftools, CD-HIT
Creating group files from Cd-Hit output clusters and automatic importing into CLANS savefile.
RnD project Collaboration
Development of a structure-driven HMM for the Kunitz domain (PF00014), combining curated 3D alignments and robust statistical evaluation. Project created during the MSc in Bioinformatics at the University of Bologna for the Laboratory of Bioinformatics 1 course.
Pipeline for discovering secreted multidomain fungal CAZymes using SignalP 6, DeepTMHMM, dbCAN, CD-HIT and homology-based candidate prioritization.
Build a well-annotated protein phylogeny (Pfam domain architectures, taxonomy, support) from a FASTA file in one command; phyloXML for Archaeopteryx.
To associate your repository with the cd-hit topic, visit your repo's landing page and select "manage topics."