This is the source code implementation of Hypergraph-based Adapter (HGAdapter) in paper HGAdapter: Hypergraph-based Adapters in Language Models for Code Summarization and Clone Detection. HGAdapter is a novel parameter-efficient fine-tuning method for code understanding with pre-trained language models. It introduces three high-order structural correlations within code — AST family correlation, lexical correlation, and line correlation — into language models through an improved hypergraph neural network integrated with adapter tuning. HGAdapter can be plugged into pre-trained language models for code-related tasks including code summarization and code clone detection.
Our paper is accepted by the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025) as a findings long paper.
Official paper link https://aclanthology.org/2025.findings-emnlp.800/
DOI reference number 10.18653/v1/2025.findings-emnlp.800
Arxiv preprint paper link https://arxiv.org/abs/2510.17591
Code implementation link https://github.com/qiankunmu/HGAdapter
We conduct experiment in Ubuntu 18.04.6 LTS and in python 3.12.
We mainly implement our method by pytorch 2.4, transformers 4.48 and adapters 1.1.
We train our model on a RTX 3090.
The required environments are listed in requirements.txt.
CodeSearchNet is the dataset for code summarization task. We use the edition from CodeXGLUE. The following process is adapted from CodeXGLUE, you can also refer to this.
First, you need to download datasets.zip, then put it into Code-summarization-CodeSearchNet/data.
Then
unzip dataset.zip
cd dataset
Second, you can download ruby.zip, javascript.zip, java.zip, python.zip, php.zip and go.zip from Hugging Face.
Then, put them into Code-summarization-CodeSearchNet/data/dataset and unzip them.
unzip python.zip
unzip java.zip
unzip ruby.zip
unzip javascript.zip
unzip go.zip
unzip php.zip
rm *.zip
rm *.pkl
Third, run preprocess.py to get the data.
python preprocess.py
rm -r */final
rm -r */*.txt
BigCloneBench is the dataset for code clone detection task.
We use the edition from CodeXGLUE.
You need to download the datasets.
You can download data.jsonl, train.txt, valid.txt and test.txt, then put them into Clone-detection-BigCloneBench/data/dataset.
1. For code summarization task, you can run trainLlama.py, it can train, valid and test the Llama series models with HGAdapter inserted.
You can modify pretrain_model_name_or_path in trainLlama.py to choose the pre-trained model, such as
pretrain_model_name_or_path = "CodeLlama-7b"
You can also choose "TinyLlama_v1.1_math_code" or other Llama series models.
You can run trainHyper.py, it can train, valid and test the RoBERTa series models with HGAdapter inserted.
You can modify pretrain_model_name_or_path in trainHyper.py to choose the pre-trained model, such as
pretrain_model_name_or_path = "codebert-base"
You can also choose "roberta-base", "graphcodebert-base", "unixcoder-base" or other RoBERTa series models.
You can run trainQwen2.py, it can train, valid and test the Qwen2.5 series models with HGAdapter inserted.
You can modify pretrain_model_name_or_path in trainQwen2.py to choose the pre-trained model, such as
pretrain_model_name_or_path = "Qwen2.5-Coder-0.5B"
You can also choose other Qwen2.5 series models.
2. For code clone detection, you can run trainHyper.py, it can train, valid and test the RoBERTa series models with HGAdapter inserted.
You can modify pretrain_model_name_or_path in trainHyper.py to choose the pre-trained model, such as
pretrain_model_name_or_path = "codebert-base"
You can also choose "graphcodebert-base" and "unixcoder-base".
3. You can also locally download pre-trained models into pre_train_models/, modify the pretrain_model_name_or_path to the corresponding path.
4. After the run is complete, the model will be saved to the work_dir, and the results will be outputted and saved as a CSV in work_dir.
The implementation of the HGAdapter is located in the models folder.
You can manually modify hyper-parameters such as train_batch_size, num_epochs in trainXXX.py.
If you use our code, please cite us.
@inproceedings{Yang2025HGAdapter,
title={HGAdapter: Hypergraph-based Adapters in Language Models for Code Summarization and Clone Detection},
author={Guang Yang and Yujie Zhu},
booktitle={Findings of the Association for Computational Linguistics: {EMNLP} 2025},
pages={14823–14833},
year={2025},
doi={10.18653/v1/2025.findings-emnlp.800}
}