Skip to content

Repository files navigation

Analyzing the Energy Consumption of Generative Text-to-Music Models

Francesca Ronchini*1, Riccardo Passoni*2, Luca Comanducci1, Romain Serizel3, Fabio Antonacci1

1 Dipartimento di Elettronica, Informazione e Bioingegneria - Politecnico di Milano, Milan, Italy
2 Institute of Sound Recording (IoSR), University of Surrey, Guildford, UK
3 Université de Lorraine, CNRS, Inria, Loria, Nancy, France
*These authors contributed equally

Abstract

Music generation via multimodal machine learning generative models has recently gained significant traction in both the research community and the general public. Among such models, an important category is represented by Text-To-Music, which take only a textual description or caption as input to specify the type of composition the user wishes to generate. While generative artificial intelligence models achieve remarkable performance, they also require substantial computational power during both training and inference. The latter, in particular, becomes increasingly critical as these models are deployed more widely in real-world music-making applications. In this paper, we address this issue by analyzing the energy consumption of seven diffusion-based and two autoregressive text-to-music models. Through a series of experiments, we evaluate how model-dependent generation parameters influence energy usage during inference. Since audio generation quality, prompt adherence, and energy efficiency are all important, we employ Pareto-optimal analysis to explore the trade-offs among these competing objectives. Our results highlight the balance between model performance and energy usage, offering guidance for designing more environmentally conscious generative audio systems.

Install & Usage

In order to run the Jupyter notebooks, you need to clone the repo, create a virtual environment, and install the needed packages.

You can create the virtual environment and install the needed packages using conda with the following command:

conda env create -f requirements.yml

Once everything is installed, you can run the Jupyter Notebook following the instruction reported on it and reproduce the results.

The scripts contained in the 'inferences' folder can be run by creating environments specific to the desired model; further information is provided in the folder's README.

Additional information

For more details: "Analyzing the Energy Consumption of Generative Text-to-Music Models" (Francesca Ronchini, Riccardo Passoni, Luca Comanducci, Romain Serizel, Fabio Antonacci)

If you use code or comments from this work, please cite:

@ARTICLE{ronchini2026,
  author={Ronchini, Francesca and Passoni, Riccardo and Comanducci, Luca and Serizel, Romain and Antonacci, Fabio},
  journal={},
  title={Analyzing the Energy Consumption of Generative Text-to-Music Models},
  year={2026},
  volume={},
  number={},
  pages={},
  keywords={},
  doi={}

About

Analyzing the Energy Consumption of Generative Text-to-Music Models

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages