Reproduction and evaluation of DeepCDA (Deep Cross-Domain Compound-Protein Affinity prediction). The repository contains a from-scratch PyTorch re-implementation of the DeepCDA encoder and its two-sided attention mechanism, an Adversarial Discriminative Domain Adaptation (ADDA) study for cross-dataset transfer (Davis to KIBA and back), and the benchmark baselines used for comparison.
- Original paper: Abbasi et al., Bioinformatics 2020
- Original code: LBBSoft/DeepCDA
- Full write-up:
Report.pdf, covering methodology, experiments, and results.
.
├── main.py # Entry point: encoder training / validation / testing
├── ada_train.py # Entry point: adversarial domain adaptation (ADDA)
├── models.py # DeepCDA encoder (CNN + LSTM + two-sided attention) and Discriminator
├── dataset.py # Dataset loading, folding, target scaling, DataLoader construction
├── train_test_utils.py # Training / validation / test loops for the encoder
├── ada_train_utils.py # Discriminator training and adversarial encoder-adaptation utilities
├── utils.py # Seeding, logging, metric computation, pickle helpers
├── configs.yaml # Config consumed by main.py
├── ada_config.yaml # Config consumed by ada_train.py
├── requirements.txt # Python dependencies
├── Report.pdf # Final project report
└── benchmarks/ # Baseline models used for comparison in the report
├── SimBoost/ # Feature-engineering + gradient-boosting baseline
└── DeepDTA/ # CNN-only deep-learning baseline (legacy Keras/TF 1.x)
KronRLS is discussed in the report as a third baseline but its evaluation code is not included here.
Download the folded Davis / KIBA datasets and place them under ../Data_folded/ (a sibling of
this folder), so the paths in dataset.py resolve
(../Data_folded/Davis/..., ../Data_folded/KIBA/...):
https://drive.google.com/open?id=15KotSJWknMOAnHM68RpOh_rqMISsMwsE
The benchmark scripts under benchmarks/ expect their own data/ subfolder. See each
benchmark's notes below.
Install Miniconda, then create the environment and install dependencies (run this from inside the project folder):
conda create -n deepcda python=3.11
conda activate deepcda
pip install -r requirements.txtNote: The version pins in
requirements.txtwere captured on macOS. On another OS a given pinned version may not be available; remove the version numbers ifpipfails to resolve them.
Edit configs.yaml to set the dataset, protein / SMILES filter sizes, embedding dimension,
channel dimension, number of epochs, learning rate, and other hyperparameters. Hyperparameters
must be given as lists where the example config uses lists. Set train: True, then run:
python3 main.py --config_path configs.yamlSet test: True in configs.yaml, keeping the same hyperparameters used at training time
(otherwise the saved weights will not match), then run the same command:
python3 main.py --config_path configs.yamlFirst train and save the encoder weights (above). Then edit ada_config.yaml, in particular
saved_model_folder (path to the trained encoder weights) and save_folder (where adapted
weights are written), and run:
python3 ada_train.py --config_path ada_config.yamlTo evaluate the adapted model, reuse the encoder test script: in configs.yaml, point
save_folder at the domain-adaptation output folder, set test: True, and run:
python3 main.py --config_path configs.yamlMetrics are written to metrics.csv in the current working directory.
Each baseline is self-contained in its own folder and is run from inside that folder. Both expect
a data/<dataset>/ subfolder holding the standard DeepDTA-format benchmark files
(ligands_can.txt, proteins.txt, Y, folds/test_fold_setting1.txt, and, for SimBoost, the
*_similarities_*.txt matrices).
Feature-engineering baseline (similarity averages, PageRank / neighbour counts, NMF latent
factors) fed to an XGBoost regressor. Trained model artifacts are checked in as
davis_simboost.pkl / kiba_simboost.pkl; the evaluation scripts re-fit from scratch.
cd benchmarks/SimBoost
python evaluate_davis.py # writes simboost_results_davis.txt
python evaluate_kiba.py # writes simboost_results_kiba.txtCNN-only deep-learning baseline. The pretrained Keras weights (combined_davis.h5,
combined_kiba.h5, about 23 MB each) are not committed; download them separately and drop
them in benchmarks/DeepDTA/ (see that folder's README.md). Because these are legacy
TensorFlow 1.x / Keras 2.3 models, benchmarks/DeepDTA/README.md documents running them via a
tensorflow/tensorflow:1.15.2-py3 Docker container (AMD64 emulation on Apple Silicon). Select
the dataset at the top of evaluate.py (DATASET_NAME / MODEL_PATH), then:
cd benchmarks/DeepDTA
python evaluate.py # writes results_<dataset>.txtDeepCDA (our re-implementation, using the GitHub-variant attention, full batch ratio) vs.
baselines on the held-out test split. Arrows show the better direction; r2m is the QSAR
regression-toward-the-mean metric.
Davis
| Model | MSE (down) | Pearson (up) | r2m (up) | CI (up) | AUPR (up) |
|---|---|---|---|---|---|
| SimBoost | 0.2528 | 0.8315 | 0.6387 | 0.8878 | n/a |
| DeepDTA | 0.2766 | 0.8141 | 0.6368 | 0.8677 | 0.7143 |
| DeepCDA (original) | 0.248 | 0.857 | 0.649 | 0.891 | 0.739 |
| DeepCDA (ours) | 0.207 | 0.741 | 0.664 | 0.883 | 0.735 |
KIBA
| Model | MSE (down) | Pearson (up) | r2m (up) | CI (up) | AUPR (up) |
|---|---|---|---|---|---|
| SimBoost | 0.2420 | 0.8026 | 0.6369 | 0.8330 | n/a |
| DeepDTA | 0.2106 | 0.8442 | 0.6856 | 0.8557 | 0.7816 |
| DeepCDA (original) | 0.176 | 0.855 | 0.682 | 0.889 | 0.812 |
| DeepCDA (ours) | 0.167 | 0.753 | 0.655 | 0.841 | 0.789 |
See Report.pdf for the full tables (including KronRLS, the Eq. 5 attention variant, and the
domain-adaptation results) and discussion.