Skip to content
qiankunmuPublic

About

[EMNLP 2025] This is the source code implementation of Hypergraph-based Adapter (HGAdapter) in paper HGAdapter: Hypergraph-based Adapters in Language Models for Code Summarization and Clone Detection.

Topics

Resources

Stars

2 stars

Watchers

1 watching

Forks

Repository files navigation

Hypergraph-based Adapter (HGAdapter)

Introduction

This is the source code implementation of Hypergraph-based Adapter (HGAdapter) in paper HGAdapter: Hypergraph-based Adapters in Language Models for Code Summarization and Clone Detection. HGAdapter is a novel parameter-efficient fine-tuning method for code understanding with pre-trained language models. It introduces three high-order structural correlations within code — AST family correlation, lexical correlation, and line correlation — into language models through an improved hypergraph neural network integrated with adapter tuning. HGAdapter can be plugged into pre-trained language models for code-related tasks including code summarization and code clone detection.

Our paper is accepted by the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025) as a findings long paper.

Related Links

paper    Official paper link https://aclanthology.org/2025.findings-emnlp.800/
doi         DOI reference number 10.18653/v1/2025.findings-emnlp.800
arxiv    Arxiv preprint paper link https://arxiv.org/abs/2510.17591
GitHub   Code implementation link https://github.com/qiankunmu/HGAdapter

Requirments

We conduct experiment in Ubuntu 18.04.6 LTS and in python 3.12. We mainly implement our method by pytorch 2.4, transformers 4.48 and adapters 1.1. We train our model on a RTX 3090. The required environments are listed in requirements.txt.

Datasets

CodeSearchNet

CodeSearchNet is the dataset for code summarization task. We use the edition from CodeXGLUE. The following process is adapted from CodeXGLUE, you can also refer to this.

First, you need to download datasets.zip, then put it into Code-summarization-CodeSearchNet/data.
Then

unzip dataset.zip
cd dataset

Second, you can download ruby.zip, javascript.zip, java.zip, python.zip, php.zip and go.zip from Hugging Face.
Then, put them into Code-summarization-CodeSearchNet/data/dataset and unzip them.

unzip python.zip
unzip java.zip
unzip ruby.zip
unzip javascript.zip
unzip go.zip
unzip php.zip
rm *.zip
rm *.pkl

Third, run preprocess.py to get the data.

python preprocess.py
rm -r */final
rm -r */*.txt

BigCloneBench

BigCloneBench is the dataset for code clone detection task. We use the edition from CodeXGLUE. You need to download the datasets. You can download data.jsonl, train.txt, valid.txt and test.txt, then put them into Clone-detection-BigCloneBench/data/dataset.

Usages

1. For code summarization task, you can run trainLlama.py, it can train, valid and test the Llama series models with HGAdapter inserted. You can modify pretrain_model_name_or_path in trainLlama.py to choose the pre-trained model, such as

pretrain_model_name_or_path = "CodeLlama-7b"

You can also choose "TinyLlama_v1.1_math_code" or other Llama series models.

You can run trainHyper.py, it can train, valid and test the RoBERTa series models with HGAdapter inserted. You can modify pretrain_model_name_or_path in trainHyper.py to choose the pre-trained model, such as

pretrain_model_name_or_path = "codebert-base"

You can also choose "roberta-base", "graphcodebert-base", "unixcoder-base" or other RoBERTa series models.

You can run trainQwen2.py, it can train, valid and test the Qwen2.5 series models with HGAdapter inserted. You can modify pretrain_model_name_or_path in trainQwen2.py to choose the pre-trained model, such as

pretrain_model_name_or_path = "Qwen2.5-Coder-0.5B"

You can also choose other Qwen2.5 series models.

2. For code clone detection, you can run trainHyper.py, it can train, valid and test the RoBERTa series models with HGAdapter inserted. You can modify pretrain_model_name_or_path in trainHyper.py to choose the pre-trained model, such as

pretrain_model_name_or_path = "codebert-base"

You can also choose "graphcodebert-base" and "unixcoder-base".

3. You can also locally download pre-trained models into pre_train_models/, modify the pretrain_model_name_or_path to the corresponding path.

4. After the run is complete, the model will be saved to the work_dir, and the results will be outputted and saved as a CSV in work_dir.
The implementation of the HGAdapter is located in the models folder.
You can manually modify hyper-parameters such as train_batch_size, num_epochs in trainXXX.py.

Citation

If you use our code, please cite us.

@inproceedings{Yang2025HGAdapter,
  title={HGAdapter: Hypergraph-based Adapters in Language Models for Code Summarization and Clone Detection},
  author={Guang Yang and Yujie Zhu},
  booktitle={Findings of the Association for Computational Linguistics: {EMNLP} 2025},
  pages={14823–14833},
  year={2025},
  doi={10.18653/v1/2025.findings-emnlp.800}
}

About

[EMNLP 2025] This is the source code implementation of Hypergraph-based Adapter (HGAdapter) in paper HGAdapter: Hypergraph-based Adapters in Language Models for Code Summarization and Clone Detection.

Topics

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages