forked from ggml-org/llama.cpp
-
Notifications
You must be signed in to change notification settings - Fork 11
Pull requests: mxxm-t/mx-llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
cuda: test RCCL with a small all-reduce after startup and fall back cleanly (#7)
CUDA
ggml
#9
opened Sep 4, 2026 by
JCraigWasTaken
Loading…
speculative: sample the MTP draft and verify it by rejection sampling
documentation
Improvements or additions to documentation
server
testing
#8
opened Sep 3, 2026 by
JCraigWasTaken
Loading…
llama: enable tensor split for bailingmoe3 + nemotron_h_moe
#5
opened Sep 1, 2026 by
assistmeister
Loading…
ProTip!
Mix and match filters to narrow down what you’re looking for.