Skip to content

[CUDA backend ONLY] Use just K-cache for MLA + FA: 47% saving on KV-cache size - #13529

Closed
jukofyork wants to merge 1 commit into
ggml-org:masterfrom
jukofyork:mla-fa-disable-v-cache
Closed

[CUDA backend ONLY] Use just K-cache for MLA + FA: 47% saving on KV-cache size#13529
jukofyork wants to merge 1 commit into
ggml-org:masterfrom
jukofyork:mla-fa-disable-v-cache

Revived PR

2f2fd15
Select commit
Loading
Failed to load commit list.

Workflow runs completed with no jobs