The official source code of our CVPR 2026 paper, "SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval".
We used Anaconda to setup a deep learning workspace that supports PyTorch. Run the following script to install all the required packages.
conda create -n SAVE python==3.9 -y
conda activate SAVE
git clone https://github.com/ruc-aimc-lab/SAVE.git
cd SAVE
pip install -r requirements.txtThree pre-trained checkpoints are required. Place them under ./pretrained/ at the repository root, with the following layout:
| Model | Source | Path |
|---|---|---|
| CLIP ViT-B/32 | OpenAI | pretrained/CLIP/ViT-B-32.pt |
| AST (AudioSet 10/10) | AST repository | pretrained/AST/audioset_10_10.pth |
| ImageBind huge | ImageBind | pretrained/ImageBind/imagebind_huge.pth |
For convenience, we provide a script that downloads all three checkpoints and places them at the paths above:
bash preprocess/download_pretrained.shIf you prefer a different location, override the path via environment variable, e.g. export AST_PATH=/somewhere/else/audioset_10_10.pth.
Use MSR-VTT as the running example below; the same recipe applies to VATEX, Charades, and LSMDC by replacing msrvtt with the corresponding name.
-
Annotations & ASR. We provide caption annotations, data splits, and extracted ASR transcriptions at Google Drive. Download the folder and place its contents under
./data/. The layout on Google Drive already matches what our code expects, so no renaming is needed. -
Raw videos. Follow the guide from CLIP4Clip: Data Preparing to obtain the raw
.mp4clips, and put (or symlink) them intodata/msrvtt/VideoData/. -
Audio. Extract 16 kHz mono
.wavfrom each clip:python preprocess/extract_audio.py
This populates
data/msrvtt/AudioData/. Clips with no audio track are skipped; the dataloader substitutessilent_file.wavat training time. -
Teacher features (ImageBind). We have already extracted and uploaded the teacher features to the Google Drive. You simply need to download and place them under
data/msrvtt/FeatureData/.(Optional) If you prefer to extract them from scratch, clone ImageBind locally and run our script:
IMAGEBIND_DIR=/path/to/ImageBind bash fe.sh
This populates
data/msrvtt/FeatureData/ImageBind/{Audio,Video}Feature/.
Before starting to run the code, please organize the data and weights in the following format (taking MSR-VTT as the example):
SAVE
├── pretrained
│ ├── CLIP
│ │ └── ViT-B-32.pt
│ ├── AST
│ │ └── audioset_10_10.pth
│ └── ImageBind
│ └── imagebind_huge.pth
└── data
└── msrvtt
├── Annotations
│ ├── MSRVTT_data.json
│ ├── MSRVTT_train.9k.csv
│ ├── MSRVTT_train.7k.csv
│ ├── MSRVTT_JSFUSION_test.csv
│ └── ...
├── VideoData
│ ├── video0.mp4
│ └── ...
├── AudioData
│ ├── video0.wav
│ └── ...
├── FeatureData
│ └── ImageBind
│ ├── AudioFeature
│ │ ├── video0.pt
│ │ └── ...
│ └── VideoFeature
│ ├── video0.pt
│ └── ...
└── msrvtt10k_asr_text.jsonAfter everything is in place, run a one-shot sanity check:
python preprocess/verify_data.pyYou can train SAVE on specific dataset splits using the following commands:
# MSR-VTT 9k
bash scripts/run_msrvtt-9k.sh
# MSR-VTT 7k
bash scripts/run_msrvtt-7k.shBy default the script trains on 2 GPUs with global batch size 128. Override via env vars when needed, e.g. NPROC=4 BATCH_SIZE=128 bash scripts/run_msrvtt-9k.sh.
After training, evaluate the models using:
INIT_MODEL=./outputs/save_msrvtt9k/pytorch_model.bin.4 bash scripts/run_eval.shIf you find SAVE useful in your work, please cite:
@inproceedings{save,
title={SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval},
author={Zhao, Ruixiang and Xu, Zhihao and Lan, Bangxiang and Xin, Zijie and Liu, Jingyu and Li, Xirong},
booktitle={CVPR},
year={2026}
}Our codebase builds on AVIGATE, CLIP4Clip, AST, and ImageBind. We thank the original authors for their open-sourcing.
If you encounter any issue when running the code, please feel free to reach us either by creating a new issue in the GitHub or by emailing
- Zhihao Xu (xuzhihao@ruc.edu.cn)
- Ruixiang Zhao (ruixiangzhao@ruc.edu.cn)
