Task- and Type-Aware KV Cache Compression for Unified Multimodal Models
Paper · Project Page · Overview · Results · Installation · Quick Start · Implementations · Citation
Image generation with UniCache + BAGEL at approximately 60% logical KV compression.
UniCache is a training-free framework for task- and type-aware KV cache compression. It identifies the cache segments activated by each task and assigns suitable compression policies through offline calibration. During inference, it coordinates their storage budgets through attention-guided allocation and task-aware temporal scheduling, then applies the assigned policies independently and in parallel.
This inference-only release targets BAGEL-7B-MoT and supports image understanding, text-to-image generation, and image editing. It includes both a PyTorch reference implementation and a physical CUDA inference engine.
| BAGEL method | MME Total ↑ | GenEval Overall ↑ | PIE Structure ↓ | PIE PSNR ↑ |
|---|---|---|---|---|
| Full KV | 2373.21 | 0.781 | 0.101 | 18.805 |
| Global H2O | 2372.60 | 0.777 | 0.125 | 14.663 |
| Global KIVI | 2362.31 | 0.777 | 0.098 | 18.576 |
| UniCache | 2377.36 | 0.780 | 0.100 | 18.969 |
UniCache reaches approximately 80% logical KV compression for understanding and editing, and approximately 60% for generation. Its physical engine improves long-context throughput by up to 1.78x on a single A100.
Image-editing results at high logical KV compression, compared with Full KV, global H2O, and global KIVI.
| Backend | Implementation | Included presets | Purpose |
|---|---|---|---|
torch |
PyTorch attention masks and KIVI-style fake quantization | Attention-guided allocation; constant total budget for understanding, conditional-attention decay for generation/editing | Algorithm inspection and quality evaluation |
engine |
Physically compacted GQA KV; packed KIVI with CUDA/Triton and FlashAttention | Fixed per-type budgets and frozen H2O selection after initial scoring | Physical storage and runtime evaluation |
full |
Full-KV BAGEL | No compression | Baseline |
| Cache type | Understanding | Generation | Editing |
|---|---|---|---|
| Instruction | H2O-style eviction | H2O-style eviction | H2O-style eviction |
| Source ViT | H2O-style eviction | Absent | H2O-style eviction |
| Source VAE | Absent | Absent | KIVI-style quantization |
| Boundary / decoded text / current latent | Protected where present | Protected where present | Protected where present |
Model inference requires Linux, Python 3.10/3.11, and an NVIDIA CUDA GPU. BAGEL weights use roughly 30 GB in BF16.
Run the following commands from the repository root:
python3.10 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install torch==2.5.1 torchvision==0.20.1 \
--index-url https://download.pytorch.org/whl/cu121
python -m pip install -r requirements/torch.txtInstall a CUDA toolkit and C++ compiler compatible with your PyTorch build;
nvcc must be available when building the extensions.
python -m pip install ninja packaging setuptools wheel
python -m pip install --no-build-isolation -r requirements/engine.txt
# Set the architecture for your GPU; 8.0 is A100.
export TORCH_CUDA_ARCH_LIST=8.0
bash scripts/install_kivi.sh
export UNICACHE_KIVI_ROOT="$PWD/third_party/KIVI"
python scripts/check_environment.py --backend engine --task editinginstall_kivi.sh fetches pinned KIVI, applies the included multi-query
addressing patch, and builds its extension. The engine uses KIVI for source-VAE
KV in image editing and H2O-style eviction for sparse cache types.
Obtain the official BAGEL weights separately and comply with their license:
python scripts/download_model.py --output-dir checkpoints/BAGEL-7B-MoTSupply your own source image for understanding/editing. Use a different output directory for every run.
# Understanding
python infer.py --backend torch --task understanding \
--model-path checkpoints/BAGEL-7B-MoT --image /path/to/source.jpg \
--prompt "Describe the image." --max-new-tokens 128 \
--output-dir outputs/understanding
# Text-to-image
python infer.py --backend torch --task text_to_image \
--model-path checkpoints/BAGEL-7B-MoT \
--prompt "A lighthouse above a turquoise sea, watercolor painting." \
--image-size 512 --timesteps 50 --seed 42 \
--output-dir outputs/generation
# Editing
python infer.py --backend torch --task editing \
--model-path checkpoints/BAGEL-7B-MoT --image /path/to/source.jpg \
--prompt "Replace the background with a glacier lake, preserving the subject." \
--image-size 512 --min-image-size 400 --timesteps 50 --seed 42 \
--output-dir outputs/editingReplace --backend torch with --backend engine or --backend full to run the
physical implementation or uncompressed baseline. Each backend loads its
task-specific configuration automatically.
Convenience wrappers are also available:
bash scripts/infer_torch.sh --task text_to_image --prompt "A lighthouse." \
--image-size 512 --output-dir outputs/torch_demo
bash scripts/infer_engine.sh --task text_to_image --prompt "A lighthouse." \
--image-size 512 --output-dir outputs/engine_demoUse --config /path/to/config.json to override a backend preset. To inspect
the compression plan without loading weights:
python infer.py --backend torch --task text_to_image --prompt "A lighthouse." \
--plan-only --output-dir outputs/plangenerated.pngoranswer.txt: inference result.resolved_config.json/unicache_plan.json: effective settings and plan.hook_policy_summary.json: runtime operators and cache accounting.budget_allocations.jsonl,budget_by_cache_type.json,budget_average.json: dynamic budget diagnostics.run_metadata.json: inputs, seed, backend, software versions and runtime flags.
Add --efficiency-events outputs/run/events.jsonl --warmup-runs 1 --measure-runs 3
to record inference timings, excluding model loading.
UniCache/
infer.py # Shared three-task CLI; torch / engine / full
model_loader.py # BAGEL checkpoint loading
inferencer.py # BAGEL interleaved inference
configs/{torch,engine}/ # Three inference presets per implementation
unicache/ # Segments, policies, schedules, adapters, storage
modeling/ # BAGEL model and inference dependencies
data/ # Image transforms and token helpers only
patches/ # Pinned KIVI multi-query patch
scripts/ # Dependency setup, download, inference wrappers
requirements/ # Separate dependency recipes
tests/ # Small contract and numerical checks, no datasets
python -m unittest discover -s tests -vTests cover selection, cache accounting, protection, configuration, and attention equivalence. CUDA tests require FlashAttention and KIVI.
If you find UniCache useful, please cite our paper:
@misc{yang2026unicachetasktypeawarekv,
title={UniCache: Task- and Type-Aware KV Cache Compression for Unified Multimodal Models},
author={Wanqi Yang and Yuexiao Ma and Mei Xie and Xiawu Zheng and Shiwei Liu},
year={2026},
eprint={2609.32831},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2609.32831},
}Code is released under the Apache 2.0 license. Adapted BAGEL code and external method attributions are listed in NOTICE. BAGEL weights and the KIVI dependency must be obtained separately under their own terms.




