-
Notifications
You must be signed in to change notification settings - Fork 451
Pull requests: ikawrakow/ik_llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
docs: --cache-ram-similarity is a minimum, not a maximum
#2395
opened Aug 31, 2026 by
mattafaak
Loading…
Support SWA compression with DFlash/DSpark
#2384
opened Aug 30, 2026 by
SamuelOliveirads
Collaborator
Loading…
server : add --sleep-idle-seconds to auto-unload model on idle
#2383
opened Aug 29, 2026 by
Ahmed-Yaseen99
Loading…
model: add LFM2 / LFM2.5 (Liquid AI) runtime + conversion support
#2382
opened Aug 29, 2026 by
yoirtbryggffdcfgcbcbcffdbfxx
Loading…
2 of 4 tasks
gemma4: compacted sliding-window KV cache (--swa-compress)
#2378
opened Aug 29, 2026 by
azilber
Loading…
server: disk KV block cache for cross-process prompt prefix reuse
#2377
opened Aug 29, 2026 by
AYwlilwYA
Loading…
model: Add GLM-5.3-Flash (glm5next) runtime support
#2376
opened Aug 28, 2026 by
Skelectric
Contributor
Loading…
2 of 4 tasks
Map dense Qwen DFlash back to LLM_ARCH_DFLASH
#2370
opened Aug 28, 2026 by
SamuelOliveirads
Collaborator
Loading…
qwen4exp: MTP (NextN) self-speculative decoding support
#2369
opened Aug 28, 2026 by
jcr211
Loading…
2 of 4 tasks
Check if enough shared memory available for MLA on CUDA
#2354
opened Aug 25, 2026 by
ikawrakow
Owner
Loading…
Dspark confidence method in spec-autotune
#2326
opened Aug 16, 2026 by
SamuelOliveirads
Collaborator
Loading…
PR: Transfer ATSInfer Tensor Placement Solver into
ik_llama.cpp
#2259
opened Aug 5, 2026 by
giveen
Loading…
2 of 4 tasks
deepseek4: add --dsv4-cache-cpu to keep compressed-attention K caches in host memory
#2239
opened Aug 2, 2026 by
nmccrory
Loading…
deepseek4: fix multi-stream (parallel sequence) decoding
#2229
opened Aug 2, 2026 by
nmccrory
Loading…
window-sized SWA ring KV cache (--swa-compress to opt in)
#2171
opened Jul 23, 2026 by
joaomdsg
Loading…
CUDA sm_60: bound the fp16 flash-attention error on P100 prefill instead of disabling the fast path
#2151
opened Jul 18, 2026 by
mb8565
Contributor
Loading…
ggml : fix build and MoE inference on CPUs without AVX2 (AVX1-only)
#2138
opened Jul 15, 2026 by
neomindryan
Loading…
2 of 4 tasks
llama: disable fused up-gate when an offload backend lacks GGML_OP_FUSED_UP_GATE (fixes silent 3.5x Metal slowdown)
#2133
opened Jul 15, 2026 by
hchengit
Contributor
Loading…
CUDA: update captured graph when a source tensor's shape changes
#2096
opened Jul 8, 2026 by
fibrahimov
Loading…
Previous Next
ProTip!
Type g p on any issue or pull request to go back to the pull request listing page.