Skip to content

Repository files navigation

UniCache

Task- and Type-Aware KV Cache Compression for Unified Multimodal Models

Paper · Project Page · Overview · Results · Installation · Quick Start · Implementations · Citation

arXiv

Sandstone arch and moon Flower shop after rain Garden inside a space station

Image generation with UniCache + BAGEL at approximately 60% logical KV compression.

Overview

UniCache is a training-free framework for task- and type-aware KV cache compression. It identifies the cache segments activated by each task and assigns suitable compression policies through offline calibration. During inference, it coordinates their storage budgets through attention-guided allocation and task-aware temporal scheduling, then applies the assigned policies independently and in parallel.

This inference-only release targets BAGEL-7B-MoT and supports image understanding, text-to-image generation, and image editing. It includes both a PyTorch reference implementation and a physical CUDA inference engine.

UniCache framework

Results

BAGEL method MME Total ↑ GenEval Overall ↑ PIE Structure ↓ PIE PSNR ↑
Full KV 2373.21 0.781 0.101 18.805
Global H2O 2372.60 0.777 0.125 14.663
Global KIVI 2362.31 0.777 0.098 18.576
UniCache 2377.36 0.780 0.100 18.969

UniCache reaches approximately 80% logical KV compression for understanding and editing, and approximately 60% for generation. Its physical engine improves long-context throughput by up to 1.78x on a single A100.

UniCache image-editing comparison

Image-editing results at high logical KV compression, compared with Full KV, global H2O, and global KIVI.

Two Implementations

Backend Implementation Included presets Purpose
torch PyTorch attention masks and KIVI-style fake quantization Attention-guided allocation; constant total budget for understanding, conditional-attention decay for generation/editing Algorithm inspection and quality evaluation
engine Physically compacted GQA KV; packed KIVI with CUDA/Triton and FlashAttention Fixed per-type budgets and frozen H2O selection after initial scoring Physical storage and runtime evaluation
full Full-KV BAGEL No compression Baseline

Cache Policy

Cache type Understanding Generation Editing
Instruction H2O-style eviction H2O-style eviction H2O-style eviction
Source ViT H2O-style eviction Absent H2O-style eviction
Source VAE Absent Absent KIVI-style quantization
Boundary / decoded text / current latent Protected where present Protected where present Protected where present

Installation

Model inference requires Linux, Python 3.10/3.11, and an NVIDIA CUDA GPU. BAGEL weights use roughly 30 GB in BF16.

Run the following commands from the repository root:

python3.10 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install torch==2.5.1 torchvision==0.20.1 \
  --index-url https://download.pytorch.org/whl/cu121
python -m pip install -r requirements/torch.txt

Physical Engine

Install a CUDA toolkit and C++ compiler compatible with your PyTorch build; nvcc must be available when building the extensions.

python -m pip install ninja packaging setuptools wheel
python -m pip install --no-build-isolation -r requirements/engine.txt
# Set the architecture for your GPU; 8.0 is A100.
export TORCH_CUDA_ARCH_LIST=8.0
bash scripts/install_kivi.sh
export UNICACHE_KIVI_ROOT="$PWD/third_party/KIVI"
python scripts/check_environment.py --backend engine --task editing

install_kivi.sh fetches pinned KIVI, applies the included multi-query addressing patch, and builds its extension. The engine uses KIVI for source-VAE KV in image editing and H2O-style eviction for sparse cache types.

Model Weights

Obtain the official BAGEL weights separately and comply with their license:

python scripts/download_model.py --output-dir checkpoints/BAGEL-7B-MoT

Quick Start

Supply your own source image for understanding/editing. Use a different output directory for every run.

# Understanding
python infer.py --backend torch --task understanding \
  --model-path checkpoints/BAGEL-7B-MoT --image /path/to/source.jpg \
  --prompt "Describe the image." --max-new-tokens 128 \
  --output-dir outputs/understanding

# Text-to-image
python infer.py --backend torch --task text_to_image \
  --model-path checkpoints/BAGEL-7B-MoT \
  --prompt "A lighthouse above a turquoise sea, watercolor painting." \
  --image-size 512 --timesteps 50 --seed 42 \
  --output-dir outputs/generation

# Editing
python infer.py --backend torch --task editing \
  --model-path checkpoints/BAGEL-7B-MoT --image /path/to/source.jpg \
  --prompt "Replace the background with a glacier lake, preserving the subject." \
  --image-size 512 --min-image-size 400 --timesteps 50 --seed 42 \
  --output-dir outputs/editing

Replace --backend torch with --backend engine or --backend full to run the physical implementation or uncompressed baseline. Each backend loads its task-specific configuration automatically.

Convenience wrappers are also available:

bash scripts/infer_torch.sh --task text_to_image --prompt "A lighthouse." \
  --image-size 512 --output-dir outputs/torch_demo
bash scripts/infer_engine.sh --task text_to_image --prompt "A lighthouse." \
  --image-size 512 --output-dir outputs/engine_demo

Use --config /path/to/config.json to override a backend preset. To inspect the compression plan without loading weights:

python infer.py --backend torch --task text_to_image --prompt "A lighthouse." \
  --plan-only --output-dir outputs/plan

Outputs

  • generated.png or answer.txt: inference result.
  • resolved_config.json / unicache_plan.json: effective settings and plan.
  • hook_policy_summary.json: runtime operators and cache accounting.
  • budget_allocations.jsonl, budget_by_cache_type.json, budget_average.json: dynamic budget diagnostics.
  • run_metadata.json: inputs, seed, backend, software versions and runtime flags.

Add --efficiency-events outputs/run/events.jsonl --warmup-runs 1 --measure-runs 3 to record inference timings, excluding model loading.

Layout

UniCache/
  infer.py                  # Shared three-task CLI; torch / engine / full
  model_loader.py           # BAGEL checkpoint loading
  inferencer.py             # BAGEL interleaved inference
  configs/{torch,engine}/   # Three inference presets per implementation
  unicache/                 # Segments, policies, schedules, adapters, storage
  modeling/                 # BAGEL model and inference dependencies
  data/                     # Image transforms and token helpers only
  patches/                  # Pinned KIVI multi-query patch
  scripts/                  # Dependency setup, download, inference wrappers
  requirements/             # Separate dependency recipes
  tests/                    # Small contract and numerical checks, no datasets

Tests

python -m unittest discover -s tests -v

Tests cover selection, cache accounting, protection, configuration, and attention equivalence. CUDA tests require FlashAttention and KIVI.

Citation

If you find UniCache useful, please cite our paper:

@misc{yang2026unicachetasktypeawarekv,
      title={UniCache: Task- and Type-Aware KV Cache Compression for Unified Multimodal Models}, 
      author={Wanqi Yang and Yuexiao Ma and Mei Xie and Xiawu Zheng and Shiwei Liu},
      year={2026},
      eprint={2609.32831},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2609.32831}, 
}

License and Attribution

Code is released under the Apache 2.0 license. Adapted BAGEL code and external method attributions are listed in NOTICE. BAGEL weights and the KIVI dependency must be obtained separately under their own terms.

About

Task- and type-aware KV cache compression for unified multimodal models

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages