CLAlign aligns protein sequences or protein structures with residue-level embeddings from protein language models.
CLAlign learns residue-level representations for protein alignment through contrastive learning. During training, sequence or structure inputs are encoded by a protein language model, and aligned residue pairs are pulled together while non-matching residues are separated in the residue-level similarity matrix. At inference time, CLAlign embeds the query and reference proteins, builds a substitution matrix from embedding similarities, and applies dynamic programming for global or local alignment.
CLAlign currently targets Python 3.12.7+ and CUDA 12.8 builds of PyTorch. CLAlign was tested on Ubuntu 24.04 with an NVIDIA RTX 4090 GPU. On first use, the ProtT5 base model is downloaded automatically from Hugging Face if it is not already cached.
Clone the repository and install from the repository root with:
git clone https://github.com/yourh/CLAlign.git
cd CLAlign
conda create -n CLAlign python=3.12 -y
conda activate CLAlign
pip install . \
--extra-index-url https://download.pytorch.org/whl/cu128 \
--extra-index-url https://urob.github.io/numpy-mkl \
-f https://data.pyg.org/whl/torch-2.7.1+cu128.htmlThe package includes the default LoRA adapters for:
CLAlign-ProtT5CLAlign-ProstT5CLAlign-ESMIF
For a quick pairwise alignment, create a text file containing two protein
sequences, one per line. This example uses the MALISAM pair
d1cs1a_d1q8ia1:
AGAVLTNTGMSAIHLVTTVFLDLVLHSCTKYLNGHSDVVAGVVIAKDPDVVTELAWWANNIGVTG
QESVAFIPADQVPRAQHILQGEQGFRLTPLALKDFHRQPVYGLYCRAHRQLMNYEKRLREGGVTVYEADVR
Then run:
clalign align pair pair.txtThe pairwise demo can also be run on CPU:
clalign --device cpu align pair pair.txtGPU acceleration is recommended for larger workloads.
Example output:
Alignment Score: 0.282163
AGA-V-L-TNT-GMSAIHLVTTVFLDLVLHSCTKYLNGHSDVVAGVVIAKDPDVVTELAWWANNIGVT------G
|| | | ||| |||||||||||||||||| |||| | |||||||||||||||||||||||||||||| |
-QESVAFIPADQVPRAQHILQGEQGFRLTP-LALK-D-FHRQPVYGLYCRAHRQLMNYEKRLREGGVTVYEADVR
Align query proteins against reference proteins:
MKL_NUM_THREADS=1 clalign align query query.fasta reference.fasta results.csvThe output is a CSV file with query ID, reference ID, aligned strings, start
positions, and alignment score. By default, CLAlign reports hits with at least
20% query coverage. Change this cutoff with --threshold.
For local alignment:
MKL_NUM_THREADS=1 clalign align query query.fasta reference.fasta results.csv --localFor large searches, generate embeddings once and reuse them:
clalign generate-embs query.fasta query.clalign
clalign generate-embs reference.fasta reference.clalign
MKL_NUM_THREADS=1 clalign --use-precomputed align query \
query.fasta reference.fasta results.csv \
--qe query.clalign \
--re reference.clalignIf an embedding path without .npy is provided, CLAlign will also try the same
path with .npy appended.
CLAlign can be used as an embedding-based remote homology search tool by aligning a query FASTA against a reference database FASTA:
MKL_NUM_THREADS=1 clalign align query query.fasta reference.fasta homology_hits.csv \
--only-score \
--threshold 0.2 \
--keep 100--keep controls how many top reference hits are retained per query. Values
between 0 and 1 are interpreted as a fraction of the reference set; values
greater than or equal to 1 are interpreted as an absolute hit count.
By default, CLAlign loads the bundled CLAlign-ProtT5 LoRA adapter on top of
Rostlab/prot_t5_xl_uniref50.
To use the ProstT5 adapter for sequence alignment:
MKL_NUM_THREADS=1 clalign \
--base-model-name Rostlab/ProstT5 \
--model-path CLAlign-ProstT5 \
align query query.fasta reference.fasta prostt5_results.csvTo use the ESM-IF adapter for structure alignment, provide text files listing
PDB IDs and point --pdb-dir to the directory containing the corresponding
PDB files. Each line in the input list should match a file name without the
suffix, for example 1abc_A for pdbs/1abc_A.pdb.
MKL_NUM_THREADS=1 clalign \
--base-model-name esm_if1_gvp4_t16_142M_UR50 \
--model-path CLAlign-ESMIF \
--pdb-dir pdbs \
align query query_pdbs.txt reference_pdbs.txt esmif_results.csvCommon options:
clalign --help
clalign align query --helpUseful flags include:
--model-path: path to a LoRA adapter or model directory--lora/--no-lora: enable or disable LoRA loading--base-model-name: base protein language model--max-length: maximum model input length--batch-size: embedding batch size--device: device such ascudaorcpu--amp: enable autocast mixed precision--pdb-dir: read structure inputs from PDB files listed by name--local: use local alignment--only-score: skip storing alignment strings--threshold: minimum query coverage for reported hits--use-precomputed: require--qe/--reembeddings instead of loading a model
CLAlign is released under the MIT License. This repository includes vendored
ESM 2.0.1 source code under src/esm/, distributed under its original MIT
License. See THIRD_PARTY_NOTICES.md for details.
