Ron Aluf, Alon Canfi, Eliya Nachmani · Ben-Gurion University of the Negev · NeurIPS 2026
GS-Codec replaces the vector quantizer of a neural audio codec with a parametric decomposition.
Each 3-second segment of the encoder latent is represented as a weighted sum of 1D Gaussian
primitives, fitted by an inner optimization loop during training and regressed in a single forward
pass by the GS Predictor at inference. Scalar quantization is applied only after training, so a
single checkpoint covers a continuous range of bitrates by changing the number of primitives
Training (top): the encoder latent is approximated by
pip install git+https://github.com/ronaluf/gs-codecFor training and evaluation, clone the repository and install the extras:
git clone https://github.com/ronaluf/gs-codec && cd gs-codec
pip install -e ".[train,eval]"from gscodec import GSCodec
from gscodec.data.audio import audio_read, audio_write
codec = GSCodec.from_pretrained("ronaluf/gs-codec-24khz")
wav, sr = audio_read("speech.wav")
code = codec.encode(wav, sample_rate=sr, n_gaussians=102, n_bits=5) # 5.95 kbps
open("speech.gsc", "wb").write(code.to_bytes())
audio_write("speech_decoded", codec.decode(code), codec.sample_rate)From the command line:
gscodec encode speech.wav speech.gsc --n_gaussians 102 --n_bits 5
gscodec decode speech.gsc speech_decoded.wav
gscodec reconstruct speech.wav speech_decoded.wav --mode iterativeencode accepts:
| Argument | Values | Description |
|---|---|---|
mode |
predictor (default), iterative
|
GS Predictor (single pass) or inner Adam fitting (300 steps) |
n_gaussians |
any |
primitives per 3 s segment |
n_bits |
bits per scale and amplitude (k-means codebooks) | |
freeze_positions |
False (default), True
|
centers predicted and sent with 10 bits, or fixed to a uniform grid |
Per 3-second segment with
with
gscodec bitrate --n_gaussians 102 --n_bits 5 # 17850 bits / 3 s = 5.950 kbps
gscodec bitrate --n_gaussians 109 --n_bits 5 --freeze_positions # 17985 bits / 3 s = 5.995 kbpsPretrained checkpoints will be released soon.
Training uses Dora and Hydra. Build a manifest for each
split (see egs/README.md), then train the codec with the iterative
Gaussian-splatting bottleneck:
dora -P gscodec run solver=compression/gaussian_24khz dset=audio/librilightTrain the GS Predictor on the frozen codec:
dora -P gscodec run solver=compression/gs_predictor_24khz dset=audio/librilight \
gaussian_splat.amortized_pretrained_codec_path=outputs/xps/<codec_sig>/checkpoint.thExperiments are written to outputs/ (set GSCODEC_DORA_DIR to change it). Any config value can
be overridden on the command line, e.g. gaussian_splat.n_gaussians=100 dataset.batch_size=8.
Export a checkpoint to a model directory:
python scripts/export_checkpoint.py outputs/xps/<sig>/checkpoint.th pretrained/gs-codec-24khz# 1. Codebooks on a held-out set
python scripts/calibrate_codebooks.py pretrained/gs-codec-24khz --files LibriTTS/dev-clean \
--mode predictor --n_gaussians 102 --n_bits 5
# 2. Encode to bitstreams, decode from disk
python scripts/reconstruct.py pretrained/gs-codec-24khz --files LibriSpeech/test-clean \
--out_dir runs/pred_102_b5 --n_gaussians 102 --n_bits 5 --min_duration 4 --max_duration 10
# 3. PESQ, STOI, UTMOS, WER (and SIM with --sim_checkpoint)
python scripts/evaluate.py runs/pred_102_b5 --transcripts LibriSpeech/test-cleangscodec/
codec.py GSCodec API and bitstream format
ptq.py k-means codebooks, bit packing, bit accounting
quantization/
gaussian_splat.py Gaussian-splatting bottleneck (inner Adam fit, STE, commitment loss)
amortized_predictor.py GS Predictor
models/ modules/ SEANet encoder/decoder with Snake activations
solvers/ losses/ adversarial/ data/ optim/ training
config/ Hydra configs (codec, predictor, datasets)
scripts/ export, calibration, reconstruction, evaluation
tests/
Code is released under the MIT License. Parts adapted from
AudioCraft and
DAC keep their original MIT notices in
LICENSE.audiocraft, LICENSE.dac and file headers.
The training framework and SEANet autoencoder build on AudioCraft and EnCodec. The Snake activation follows DAC, and UTMOS uses SpeechMOS.
@inproceedings{aluf2026gscodec,
title = {{GS-Codec}: A Gaussian-Splatting Bottleneck for Neural Audio Coding},
author = {Aluf, Ron and Canfi, Alon and Nachmani, Eliya},
booktitle = {Advances in Neural Information Processing Systems},
year = {2026}
}
