Skip to content

Repository files navigation

DeepCDA Reproduction

Reproduction and evaluation of DeepCDA (Deep Cross-Domain Compound-Protein Affinity prediction). The repository contains a from-scratch PyTorch re-implementation of the DeepCDA encoder and its two-sided attention mechanism, an Adversarial Discriminative Domain Adaptation (ADDA) study for cross-dataset transfer (Davis to KIBA and back), and the benchmark baselines used for comparison.

Repository structure

.
├── main.py                 # Entry point: encoder training / validation / testing
├── ada_train.py            # Entry point: adversarial domain adaptation (ADDA)
├── models.py               # DeepCDA encoder (CNN + LSTM + two-sided attention) and Discriminator
├── dataset.py              # Dataset loading, folding, target scaling, DataLoader construction
├── train_test_utils.py     # Training / validation / test loops for the encoder
├── ada_train_utils.py      # Discriminator training and adversarial encoder-adaptation utilities
├── utils.py                # Seeding, logging, metric computation, pickle helpers
├── configs.yaml            # Config consumed by main.py
├── ada_config.yaml         # Config consumed by ada_train.py
├── requirements.txt        # Python dependencies
├── Report.pdf              # Final project report
└── benchmarks/             # Baseline models used for comparison in the report
    ├── SimBoost/           # Feature-engineering + gradient-boosting baseline
    └── DeepDTA/            # CNN-only deep-learning baseline (legacy Keras/TF 1.x)

KronRLS is discussed in the report as a third baseline but its evaluation code is not included here.

Data

Download the folded Davis / KIBA datasets and place them under ../Data_folded/ (a sibling of this folder), so the paths in dataset.py resolve (../Data_folded/Davis/..., ../Data_folded/KIBA/...):

https://drive.google.com/open?id=15KotSJWknMOAnHM68RpOh_rqMISsMwsE

The benchmark scripts under benchmarks/ expect their own data/ subfolder. See each benchmark's notes below.

Environment setup

Install Miniconda, then create the environment and install dependencies (run this from inside the project folder):

conda create -n deepcda python=3.11
conda activate deepcda
pip install -r requirements.txt

Note: The version pins in requirements.txt were captured on macOS. On another OS a given pinned version may not be available; remove the version numbers if pip fails to resolve them.

DeepCDA encoder

Training

Edit configs.yaml to set the dataset, protein / SMILES filter sizes, embedding dimension, channel dimension, number of epochs, learning rate, and other hyperparameters. Hyperparameters must be given as lists where the example config uses lists. Set train: True, then run:

python3 main.py --config_path configs.yaml

Testing

Set test: True in configs.yaml, keeping the same hyperparameters used at training time (otherwise the saved weights will not match), then run the same command:

python3 main.py --config_path configs.yaml

Domain adaptation (ADDA)

First train and save the encoder weights (above). Then edit ada_config.yaml, in particular saved_model_folder (path to the trained encoder weights) and save_folder (where adapted weights are written), and run:

python3 ada_train.py --config_path ada_config.yaml

To evaluate the adapted model, reuse the encoder test script: in configs.yaml, point save_folder at the domain-adaptation output folder, set test: True, and run:

python3 main.py --config_path configs.yaml

Metrics are written to metrics.csv in the current working directory.

Benchmarks

Each baseline is self-contained in its own folder and is run from inside that folder. Both expect a data/<dataset>/ subfolder holding the standard DeepDTA-format benchmark files (ligands_can.txt, proteins.txt, Y, folds/test_fold_setting1.txt, and, for SimBoost, the *_similarities_*.txt matrices).

SimBoost (benchmarks/SimBoost/)

Feature-engineering baseline (similarity averages, PageRank / neighbour counts, NMF latent factors) fed to an XGBoost regressor. Trained model artifacts are checked in as davis_simboost.pkl / kiba_simboost.pkl; the evaluation scripts re-fit from scratch.

cd benchmarks/SimBoost
python evaluate_davis.py     # writes simboost_results_davis.txt
python evaluate_kiba.py      # writes simboost_results_kiba.txt

DeepDTA (benchmarks/DeepDTA/)

CNN-only deep-learning baseline. The pretrained Keras weights (combined_davis.h5, combined_kiba.h5, about 23 MB each) are not committed; download them separately and drop them in benchmarks/DeepDTA/ (see that folder's README.md). Because these are legacy TensorFlow 1.x / Keras 2.3 models, benchmarks/DeepDTA/README.md documents running them via a tensorflow/tensorflow:1.15.2-py3 Docker container (AMD64 emulation on Apple Silicon). Select the dataset at the top of evaluate.py (DATASET_NAME / MODEL_PATH), then:

cd benchmarks/DeepDTA
python evaluate.py           # writes results_<dataset>.txt

Results (from the report)

DeepCDA (our re-implementation, using the GitHub-variant attention, full batch ratio) vs. baselines on the held-out test split. Arrows show the better direction; r2m is the QSAR regression-toward-the-mean metric.

Davis

Model MSE (down) Pearson (up) r2m (up) CI (up) AUPR (up)
SimBoost 0.2528 0.8315 0.6387 0.8878 n/a
DeepDTA 0.2766 0.8141 0.6368 0.8677 0.7143
DeepCDA (original) 0.248 0.857 0.649 0.891 0.739
DeepCDA (ours) 0.207 0.741 0.664 0.883 0.735

KIBA

Model MSE (down) Pearson (up) r2m (up) CI (up) AUPR (up)
SimBoost 0.2420 0.8026 0.6369 0.8330 n/a
DeepDTA 0.2106 0.8442 0.6856 0.8557 0.7816
DeepCDA (original) 0.176 0.855 0.682 0.889 0.812
DeepCDA (ours) 0.167 0.753 0.655 0.841 0.789

See Report.pdf for the full tables (including KronRLS, the Eq. 5 attention variant, and the domain-adaptation results) and discussion.

About

Deep Learning & Domain Adaptation: reproduction and evaluation of DeepCDA for compound-protein affinity prediction (encoder, ADDA, SimBoost/DeepDTA benchmarks)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages