forked from ggml-org/llama.cpp
-
Notifications
You must be signed in to change notification settings - Fork 1
Pull requests: GenerelSchwerz/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
docs : link CLAUDE.md to AGENTS.md
documentation
Improvements or additions to documentation
#74
opened Sep 6, 2026 by
GenerelSchwerz
Owner
Loading…
cuda: compile only the reachable native FlashAttention routes (follow-up to #55)
CUDA
devops
documentation
Improvements or additions to documentation
ggml
testing
#72
opened Sep 5, 2026 by
Piggidragon
•
Draft
llama-bench: add kv-offload/workspace sweep flags
documentation
Improvements or additions to documentation
examples
#70
opened Sep 4, 2026 by
Piggidragon
•
Draft
4 tasks done
llama : add an attention split separate from the tensor split
devops
documentation
Improvements or additions to documentation
ggml
testing
#69
opened Sep 3, 2026 by
Piggidragon
Loading…
kv-cache : spend the partial residency budget on the slowest link first
testing
#68
opened Sep 3, 2026 by
Piggidragon
Loading…
kv-cache : resolve the partial KV residency set once for the model
testing
#67
opened Sep 3, 2026 by
Piggidragon
Loading…
ggml-meta : split a host-resident KV cache by head
devops
documentation
Improvements or additions to documentation
ggml
testing
#66
opened Sep 3, 2026 by
Piggidragon
Loading…
llama : split a tied output projection under split mode tensor
testing
#65
opened Sep 3, 2026 by
Piggidragon
Loading…
ggml : report allocation failure from the meta buffer type
ggml
testing
#64
opened Sep 3, 2026 by
Piggidragon
Loading…
llama-bench placement flags, two host-KV fixes, and H1/H2/H3/H5/H13 measured
documentation
Improvements or additions to documentation
examples
ggml
#62
opened Sep 2, 2026 by
Piggidragon
•
Draft
server : allow cache reuse for text-only prompts with mmproj loaded
server
#60
opened Sep 1, 2026 by
Piggidragon
Loading…
cuda: optimize native quant FlashAttention integration (follow-up to #50)
#55
opened Aug 28, 2026 by
GenerelSchwerz
Owner
Loading…
ggml-cuda: coalesce contiguous staging copies
#46
opened Aug 26, 2026 by
GenerelSchwerz
Owner
•
Draft
ggml-cuda: prefetch overflow expert siblings
#45
opened Aug 26, 2026 by
GenerelSchwerz
Owner
•
Draft
ggml-cuda: reuse cached experts during overflow staging
#44
opened Aug 26, 2026 by
GenerelSchwerz
Owner
•
Draft
ggml-cuda: skip duplicate decode sibling prefetch
#43
opened Aug 26, 2026 by
GenerelSchwerz
Owner
•
Draft
ggml-cuda : reuse routing IDs across MoE siblings
#42
opened Aug 26, 2026 by
GenerelSchwerz
Owner
•
Draft
1 task done
sched: pipeline the delivery of a host-resident KV cache
documentation
Improvements or additions to documentation
examples
ggml
testing
#39
opened Aug 26, 2026 by
Piggidragon
Loading…
2 tasks
kv: replace eligible dense causal masks with compact prefixes
CUDA
documentation
Improvements or additions to documentation
ggml
server
testing
#7
opened Aug 21, 2026 by
GenerelSchwerz
Owner
Loading…
ProTip!
Type g p on any issue or pull request to go back to the pull request listing page.