The repo contains:
-
Official Implementation The official implementation of OASIS and AoS, which couples the per-token null posterior of
Softmax_1attention with depth-wise Attention Residuals. -
Supported Backbones OASIS decoder layers for
Llama,Qwen3andPhi-4. -
Training Code Causal LM continued pre-training with OASIS (
run_clm_oasis.py) and vanilla/outlier-free baselines (run_clm.py). -
Quantization & Outlier Testing Code Post-training WnAn quantization and outlier statistics (activation ∞-norm and kurtosis) via
validate_*.py.
Large Transformers develop massive activation outliers because attention heads that "want to do nothing" dump their probability mass on sink tokens. Outlier-efficient attention (Softmax_1, from OutEffHop) fixes this by adding a null slot to the softmax denominator:
The mass routed to the null slot,
-
Null posterior. Each attention layer computes the per-head null mass and averages it over heads into a branch-level statistic
$\psi_{l,t}$ . -
Attention Residuals. The fixed residual
$h_l = h_{l-1} + f_l(h_{l-1})$ is replaced by a learned attention over all previous layer outputs (including the embedding), applied independently per token. -
OASIS coupling. The depth-routing logits are shifted by the centered null statistic,
$$g^{\text{new}}{l,t} = g^{\text{old}}{l,t} - \beta,(\psi_{l,t} - \bar\psi_t), \qquad \beta = \mathrm{softplus}(\beta_{\text{raw}}),$$
so layers that abstained on a token receive less weight when the hidden state is aggregated.
$\beta$ is initialized near zero, which makes OASIS a near no-op at the start of fine-tuning.
# create and activate virtual python environment
conda create -n oasis python=3.10
conda activate oasis
# install required packages
cd source_code
pip install -r requirement.txt
Note: the OASIS modules (
*_oasis_attention.py) importtransformers.masking_utilsandGradientCheckpointingLayer, so they need a recenttransformersrelease, newer than the version pinned inrequirement.txt.vutils/softmax_1.pyimportstriton, so a CUDA machine is required.
OASIS/
├── source_code/
│ ├── run_clm_oasis.py # OASIS training (Llama / Qwen3 / Phi-4)
│ ├── run_clm.py, run_clm_ddp.py # causal LM baselines
│ ├── validate_clm.py # evaluation + quantization
│ ├── transformers_language/
│ │ ├── models/
│ │ │ ├── llama_oasis_attention.py
│ │ │ ├── qwen_oasis_attention.py
│ │ │ ├── phi4_oasis_attention.py
│ │ │ ├── softmax.py # SOFTMAX_MAPPING (vanilla, softmax1, clipped, ...)
│ │ │ └── quantized_*.py
│ │ ├── args.py
│ │ └── dataset_setups.py # wikitext_2 | wikitext_103 | bookcorpus_and_wiki
│ ├── quantization/ # fake-quant utilities
│ ├── vutils/softmax_1.py # Softmax_1 (+ Triton flash-attention kernel)
│ ├── accelerate_configs/
│ └── model_configs/
└── scripts/ # Slurm submission scripts
All commands below are run from source_code/.
OASIS requires --attn_softmax softmax1. With vanilla softmax the null posterior is identically zero, so the model reduces to plain Attention Residuals.
accelerate launch --config_file accelerate_configs/1gpu_fp16.yaml run_clm_oasis.py \
--pad_to_max_length \
--wd_LN_gamma \
--with_tracking \
--report_to wandb \
--run_name oasis_llama3_1b \
--seed 1000 \
--dataset_setup bookcorpus_and_wiki \
--preprocessing_num_workers 10 \
--data_cache_dir your/path/to/.hf_data \
--model_cache_dir your/path/to/.hf_cache \
--model_name_or_path meta-llama/Llama-3.2-1B \
--tokenizer_name meta-llama/Llama-3.2-1B \
--max_seq_length 2048 \
--block_size 512 \
--learning_rate 0.0004 \
--lr_scheduler_type linear \
--max_train_steps 2000 \
--num_warmup_steps 2000 \
--per_device_train_batch_size 6 \
--per_device_eval_batch_size 6 \
--gradient_accumulation_steps 32 \
--max_grad_norm 1.0 \
--weight_decay 0.1 \
--checkpointing_steps 500 \
--attn_softmax softmax1 \
--attn_res_softmax_fn vanilla \
--max_checkpointing_number 2 \
--output_dir your/path/to/saveReplace the model with Qwen/Qwen3-0.6B or microsoft/Phi-4-mini-instruct to train the other backbones. If no model is given, the script defaults to microsoft/Phi-4-mini-instruct.
On a Slurm cluster, adjust the account, partition and paths in the scripts and submit from the repo root:
sbatch scripts/submit_phi4_oasis_train.sh # Phi-4 + OASIS
sbatch scripts/submit_phi4_vanilla_train.sh # Phi-4 vanilla baseline# Causal LM (Llama / Qwen3) without OASIS coupling: same arguments as 4.1, using run_clm_ddp.py
sh run_qwen.sh
# OPT outlier experiments
sh ../scripts/submit_outlier_opt.shChoose the attention variant with --attn_softmax (vanilla, softmax1, clipped(...), clippedsoftmax1(...); see transformers_language/models/softmax.py).
To evaluate perplexity and outlier statistics (max activation ∞-norm and kurtosis), run:
accelerate launch --config_file accelerate_configs/1gpu_no_mp.yaml validate_clm.py \
--seed 1000 \
--dataset_setup wikitext_2 \
--block_size 512 \
--per_device_eval_batch_size 4 \
--attn_softmax softmax1 \
--attn_res_softmax_fn vanilla \
--data_cache_dir your/path/to/.hf_data \
--model_cache_dir your/path/to/.hf_cache \
--model_name_or_path your/path/to/checkpoint \
--output_dir your/path/to/metricsTo run WnAn (Weights-nbit, Activations-nbit) post-training quantization, add:
--quantize \
--n_bits n \
--n_bits_act nOn Slurm: sbatch scripts/submit_phi4_validate.sh (Phi-4 quantization is not yet implemented, so this script runs without --quantize). For OPT, use scripts/submit_outlier_valid_opt.sh.
| Argument | Description |
|---|---|
--attn_softmax |
Softmax used inside token attention. Must be softmax1 for OASIS to have a non-zero null posterior. |
--attn_res_softmax_fn |
Softmax used for the depth-wise Attention Residual aggregation. |
--dataset_setup |
wikitext_2, wikitext_103 or bookcorpus_and_wiki. |
--block_size |
Training / evaluation sequence length after grouping. |
--quantize, --n_bits, --n_bits_act |
Enable post-training quantization and set its bit-widths. |
If you have any question regarding our paper or codes, please feel free to start an issue.
If you use OASIS in your work, please kindly cite our paper:
OASIS
@inproceedings{luo2026attention,
title = {Attention Sinks and Outliers in Attention Residuals},
author = {Haozheng Luo and Haoran Dai and Shaoyang Zhang and
Xi Chen and Eric Hanchen Jiang and Yijiang Li and
Jingyuan Huang and Chenghao Qiu and Chenwei Xu and
Zhenyu Pan and Haotian Zhang and Binghui Wang and Yan Chen},
booktitle = {The Fortieth Annual Conference on Neural Information Processing Systems},
year = {2026},
url = {https://openreview.net/forum?id=yjVSLVS0Dh}
}
OutEffHop
@inproceedings{hu2024outlier,
title={Outlier-Efficient Hopfield Layers for Large Transformer-Based Models},
author={Jerry Yao-Chieh Hu and Pei-Hsuan Chang and Haozheng Luo and Hong-Yu Chen and Weijian Li and Wei-Po Wang and Han Liu},
booktitle={Forty-first International Conference on Machine Learning},
year={2024}
}
We appreciate the following GitHub repos a lot for their valuable code and efforts.
- OutEffHop (https://github.com/MAGICS-LAB/OutEffHop)
- Outlier-free Transformers (https://github.com/Qualcomm-AI-research/outlier-free-transformers)
- GERM (https://github.com/MAGICS-LAB/GERM)
- Hugging Face Transformers (https://github.com/huggingface/transformers)
- FROST (https://github.com/robinzixuan/FROST)
- hf-attention-normalizers (https://github.com/robinzixuan/hf-attention-normalizers)
- Attention Residaul [https://github.com/kyegomez/attn_res]