Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
633 commits
Select commit Hold shift + click to select a range
60b06ab
metal : fix FA support checks (#29122)
ggerganov Sep 19, 2026
5b59b83
metal : add MoE and SSM_CONV fusion optimizations (#28948)
ggerganov Sep 19, 2026
eb1e1f4
json-schema : accept escaped hyphen in regex patterns (#29127)
ChihebBENCHEIKH1 Sep 19, 2026
1af554f
server : improve startup log messages (#29125)
ggerganov Sep 19, 2026
7d4b92b
hexagon: enable support for TOP_K op (#29113)
aparmp-quic Sep 19, 2026
851cb34
hexagon: add support for GEGLU_QUICK (#29114)
aparmp-quic Sep 19, 2026
e613ef2
hexagon: enable I32 GET_ROWS (#29116)
aparmp-quic Sep 19, 2026
59657a6
chat : add dedicated Ling 3.0 (Bailing V3) parser (#28682)
aetherbird Sep 19, 2026
f072b10
chat : fix gemma4 required tool grammar (#29115)
aldehir Sep 19, 2026
9a9f939
metal: add F16 input to the FWHT (#29094)
bri-prism Sep 20, 2026
4260903
fix(mamba) : make time-step projection input contiguous (#28832)
abetlen Sep 20, 2026
b23efaa
ui: Fix mobile breakpoint + content overflow issues (#29108)
allozaur Sep 20, 2026
3cf0325
CUDA: enable sparse fa for qwen4 (#28770)
am17an Sep 20, 2026
3d82ef6
common/peg : handle invalid utf-8 sequences in the AST (#29161)
aldehir Sep 20, 2026
a894dae
metal : support arbitrary hc in dsv4_hc_pre (#29169)
ggerganov Sep 20, 2026
ce8caa6
CUDA: tune FA for Gemma 4 on Ampere or newer (#29152)
JohannesGaessler Sep 20, 2026
62668d6
convert: enable --fuse-qkv for muse-glimmer (#29203)
dfriehs Sep 21, 2026
932a68e
webgpu : add fused gdn + cpy (#28976)
yomaytk Sep 21, 2026
8aa161b
metal : fix deprecation warnings from macOS 27 SDK (#29136)
nikwen Sep 21, 2026
68d9053
cuda : tune MMVQ to MMQ crossover for SM70 (Volta) (#28912)
cyyself Sep 21, 2026
0c3626e
hexagon: overhaul of buffer and DMA handling to support 64bit mapping…
max-krasnyansky Sep 21, 2026
6ad1af5
ci : Upgrade CUDA to 13.4 for Ubuntu CUDA Release Builds (#29202)
sam-india-007 Sep 21, 2026
8034c1d
ggml-cpu: ARM Repack kernels for Q1_0 (#23492)
pl752 Sep 21, 2026
1aa2954
sycl : coalesce MKL-FA softmax loads instead of one work-item per row…
anantshri Sep 21, 2026
26394b4
json: Fixed json enum handling (#28518)
Silverside Sep 21, 2026
335b21f
ggml-metal : simplify fusion pattern op list declaration (#29206)
ggerganov Sep 21, 2026
711f60b
tests : remove stale comment (#29140)
mostafafaheem Sep 21, 2026
982a332
server : do not forward --api-key-file to router-spawned child instan…
nandan2003 Sep 21, 2026
e0dff58
args: add env vars for temperature, top-p, min-p and penalties (#27380)
kucharskim Sep 21, 2026
542e920
ci : refactor build-self-hosted into backend-specific workflows (#28991)
CISC Sep 21, 2026
1d72b05
tests/test-backend-ops : allow regex entries in the -o filter (#29204)
ggerganov Sep 21, 2026
161755f
test-llama-archs : make tensor data stdev configurable and improve he…
ggerganov Sep 21, 2026
1884824
CUDA: Follow up of #25635, refactoring FA shared smem swizzle (#28536)
ynankani Sep 21, 2026
af91114
sycl : pinned memory use right device context instead of 0 (#28895)
lslusarczyk Sep 21, 2026
bb3c853
sycl : support gated DSV4_HC_PRE and optional HC_POST comb matrix (#2…
cwriter Sep 21, 2026
ec91ab5
docker : bump cuda to 13.4.1 (#29207)
CISC Sep 21, 2026
6f41ac5
vendor : update cpp-httplib to 0.57.0 (#29214)
angt Sep 21, 2026
c21284c
ggml : fix dimension and stride truncation in ggml_permute (#29227)
leejet Sep 21, 2026
e6cef81
cuda : accelerate conv2d with implicit GEMM (#29135)
leejet Sep 21, 2026
f4e276a
ggml-cuda : convert contiguous tensors four elements at a time (#29155)
pwilkin Sep 21, 2026
b1c2863
cuda: fix sm_70 tile compilation error (#29224)
lingyezhixing Sep 21, 2026
9655061
llama-context : report graph inputs and input tensors during sched re…
ggerganov Sep 21, 2026
c641dfa
test-save-load-state : compare logits with NMSE and feed expected tok…
ggerganov Sep 21, 2026
fb34fc2
metal : fix mask bounds in flash attention block pre-pass (#29220)
masterFoad Sep 21, 2026
ff0dbb9
vendor : update cpp-httplib to 0.57.1 (#29239)
angt Sep 21, 2026
5836771
hexagon: new HMX-optimized GATED_DELTA_NET (#29199)
max-krasnyansky Sep 21, 2026
c550d2f
ci : update Level Zero SDK to v1.33.1 and enable the L0/oneDNN CMake …
Asahi-Prv Sep 22, 2026
ec5a12b
opencl: add A8 Q4_0 non-MoE dp4a binary kernel (#29055)
shaofeiqi Sep 22, 2026
8cfc315
Add close button to UI toasts (#28246)
agustinmista Sep 22, 2026
0ee9435
ci : publish snapdragon builds in release workflow (#29007)
ykhrustalev Sep 22, 2026
7ab4ee7
chat : Fix Muse Glimmer tool-call first parser error (#29242)
NickM-27 Sep 22, 2026
a60f9ae
cmake : allow repeated find_package calls for llama (#29228)
miyanyan Sep 22, 2026
bfd73a8
convert: add MiMo-V2.6 support (#29257)
AesSedai Sep 22, 2026
828fdf2
spec : support DFlash for HunyuanOCR (#28890)
wendadawen Sep 22, 2026
217f81c
server: Add support for binding to multiple addresses (#28690)
erusev Sep 22, 2026
348f853
jinja: use const for statement::execute and ::visit (#29271)
ngxson Sep 22, 2026
9b421fa
ui : Accept WEBM video files (#28622)
EpicEric Sep 22, 2026
c350a40
Performance tune for gemma4-26b-a4b flash attention shape. (#28450)
frobnitzem Sep 22, 2026
f95b0d9
ggml : IQ1_M build prefix sums once per block (#28706)
bartowski1182 Sep 22, 2026
0f8a414
metal : gate mul_mm_id src1 rescale behind ggml_prec (#29029)
mdegans Sep 22, 2026
73c941b
mtmd: add various sanity checks (#29276)
ngxson Sep 22, 2026
4ceb171
vulkan: add Intel Xe flash attention optimization kernels (2/3, Xe-LP…
fish-jiang Sep 22, 2026
4098fdc
server: support input_image in function_call_output (#20663) (#22575)
Empressia Sep 22, 2026
bbf99b1
server: do not pass log file to children (#29212)
dfriehs Sep 22, 2026
9919911
server: fix router eviction races with the existing queue (#29217)
ServeurpersoCom Sep 22, 2026
d5f6649
opencl: add bin kernel `kernel_gemm_noshuffle_q4_k_q8_1_dp4a_ila_a8_b…
shaofeiqi Sep 22, 2026
709fe75
jinja : fix dangling reference warning in for_statement (#29279)
ggerganov Sep 22, 2026
f46bc30
HIP : optimize IQ2/IQ3 (`__vsub4` `__vcmpne4`) using SWAR (#27962)
yanjs Sep 22, 2026
e6ab7c1
hex-dma: introduce direct-mapped DMA cache that is better suited for …
max-krasnyansky Sep 22, 2026
441df11
sampler: reduce the size of the probe (#29285)
max-krasnyansky Sep 23, 2026
08b1d2a
vulkan: hide internal symbols to prevent duplicate-dlopen state destr…
ewintr Sep 23, 2026
4d7d770
sycl : support op get_rows_back, only support fp32/fp16 (#25266)
arthw Sep 23, 2026
5e48b31
sycl: extend MMVQ GLU fusion, add rms_norm+scale and ssm_conv+silu fu…
anantshri Sep 23, 2026
384a534
sycl : support new UT case for mul_mat_hadamard fp16 (#29218)
arthw Sep 23, 2026
1a67982
cuda: top-k MoE should always fire (#28432)
am17an Sep 23, 2026
9425611
ggml-meta: resolve multi buffer views (#29266)
0cc4m Sep 23, 2026
b1ff4ca
vulkan: add IQ4_XS MMQ/MMV matmul kernels (#28415)
pwilkin Sep 23, 2026
e97545d
sycl : fix compile warnings
ggerganov Sep 23, 2026
503549c
ggml : bump version to 0.25.0 (ggml/1635)
ggerganov Sep 23, 2026
45062d4
sync : ggml
ggerganov Sep 23, 2026
183d2a0
make-release : update summary prompt
ggerganov Sep 23, 2026
86b2daa
ci : run python (jinja) test (#29302)
CISC Sep 23, 2026
633733d
model : support Gemma4 DSpark draft backbone (#29226)
hthadicherla Sep 23, 2026
18f9f7b
model-conversion : add causal-compare-logits recipe (#29305)
danbev Sep 23, 2026
26758d3
ci : fix build-cmake runner target (#29299)
CISC Sep 23, 2026
bcbc936
server: Dedup the draft HF model via dedup-cache-models (#27934)
DreamingWater Sep 23, 2026
057494f
server: accept OpenAI video_url content type and data: video URIs (#2…
calebrio02 Sep 23, 2026
ee3ecce
metal : key the fa-vec tuned table by family instead of SKU (#29075)
forforever73 Sep 23, 2026
4e416ee
jinja : parse unary +/- before variables (#29244)
cs-fisha Sep 23, 2026
42916d8
server: fix token counting API crash on sleep (#29309)
willweimike Sep 23, 2026
dc9879c
CUDA: enable sparse-fa for dsv4 prefill (again) (#29298)
am17an Sep 23, 2026
9575389
metal: add the missing f32 x bf16 mul_mv variants (#28741)
ServeurpersoCom Sep 23, 2026
bddf826
common : keep HF cache dir as path, expose UTF-8 only for logs (#29320)
angt Sep 23, 2026
66fba63
CUDA: add a reserve to avoid spurious warning on older GCC builds (#2…
am17an Sep 23, 2026
e4e2f62
ggml : bump version to 0.25.1 (ggml/1637)
ggerganov Sep 23, 2026
177cd8c
sync : ggml
ggerganov Sep 23, 2026
7fe450e
llama.cpp : bump version to 0.5.0 (#29333)
ggerganov Sep 23, 2026
fee39dd
opencl: add A8 Q6_K non-MoE dp4a binary kernel (#29057)
shaofeiqi Sep 23, 2026
6e60f35
ci : use hf-jobs-cpu-xl runner in server sanitize workflow (#29297)
ggerganov Sep 23, 2026
d2e5458
tests: add `-b/--backend` option to test-llama-archs for testing a sp…
yomaytk Sep 23, 2026
b9ae43a
server: allow preset to set log file (#29334)
ngxson Sep 23, 2026
bd4f514
convert : allow vision target for DFlash/Dspark (#29339)
tdakhran Sep 23, 2026
013b31c
scripts : make-release-desc - link previous release in changelog titl…
ggerganov Sep 24, 2026
9710a32
hexagon: reject MUL_MAT_ID when src1 precision is F32 (#29348)
jhen0409 Sep 24, 2026
4c5957c
test-save-load-state : print a per-model results table in --models mo…
ggerganov Sep 24, 2026
2b70583
server,common : fix the GCC 12 stringop-overread false positive (agai…
angt Sep 24, 2026
f830688
model : add Ling 3.0 VL support (#29151)
aetherbird Sep 24, 2026
53ed051
cuda : add conv3d with implicit GEMM (#29137)
leejet Sep 24, 2026
3423f94
vulkan: tune KHR cooperative matrix support for Adreno GPUs (#29328)
Raman-Raje Sep 24, 2026
6b790a9
vulkan: handle misalignment in conv_2d and conv_3d (#29365)
0cc4m Sep 24, 2026
70c4e15
vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (…
0cc4m Sep 24, 2026
70596c4
ci : use hf-jobs-cpu-performance, disable pytest workers (#29369)
ggerganov Sep 24, 2026
308883b
server : change default pytest workers to 4 (#29376)
danbev Sep 24, 2026
fc343a8
llama: add llama_batch_ext (#24669)
ngxson Sep 24, 2026
945064f
ui : fix missing svg use and animation elements in preview and downlo…
nandan2003 Sep 24, 2026
8212c78
test: flush status (#28352)
kurquhar Sep 24, 2026
a72e04a
cuda : add F16 kernel support for CONV_2D_DW (#29064)
crion99 Sep 24, 2026
97a418b
hexagon: support I32 CPY and CONT (#29379)
kurquhar Sep 24, 2026
07fc586
hexagon: dynamic quantizer improvements (#29395)
kurquhar Sep 24, 2026
5cf3a35
llama-grammar: fix numeric truncation for token_id parsing (#29382)
apach301 Sep 24, 2026
a02c7f5
hexagon: handle multi-sequence in concat_2d (#29344)
jhen0409 Sep 24, 2026
bced459
sync : ggml (#29396)
ggerganov Sep 24, 2026
cdc0642
metal : optimize sparse FA + clean-up (#29377)
ggerganov Sep 24, 2026
84e76d8
metal : fix graph capture and handle empty graphs (#29390)
ggerganov Sep 24, 2026
ed319fe
hexagon: use DMA for contiguous dim1 CONCAT (#29404)
kurquhar Sep 25, 2026
4de0926
hexagon: add q5_k quant type support (#29123)
jhen0409 Sep 25, 2026
f805c57
llama : fix tensor split for fused qkv with uneven K/V head sizes (#2…
am17an Sep 25, 2026
1ab7e5a
CUDA: fuse RMS_NORM + SCALE into one kernel (#29393)
InflexCZE Sep 25, 2026
f9af9be
musa: fix PH1 (MTT S5000) operator failures and build issues (#29193)
yeahdongcn Sep 25, 2026
cd74ef6
[SYCL] support sparse FA (#28796)
arthw Sep 25, 2026
66963a8
rpc: include nb in the get_alloc_size cache key and floor the result …
Jesssullivan Sep 25, 2026
d028c69
HIP: bump HIP_VERSION requried for fp8 to avoid missing __hip_fp8_e4m…
IMbackK Sep 25, 2026
e9f824d
llama : add `llama_prec_policy` + model-driven W4A4 path (#24364)
ynankani Sep 25, 2026
5a75f14
metal : split fa kernels into per-dtype libraries (#29329)
ggerganov Sep 25, 2026
e351231
metal: FWHT kernels for block widths above 512 (#29095)
bri-prism Sep 25, 2026
27b20ba
common : extract shared unicode path/string helpers (#29415)
angt Sep 25, 2026
d81aef1
gguf-py : TemplateProcessing has final word on add_special_token (#29…
CISC Sep 25, 2026
b248f4a
gguf-py : ByteLevel processing defaults bos/eos to False (#29422)
CISC Sep 25, 2026
e85e15c
Fixing the vulkan build issue of legacy GLSLC version that has no coo…
sliu39 Sep 25, 2026
a25c986
opencl: add bin kernel `kernel_gemm_noshuffle_q5_k_f32_32b_trans_ila_…
shaofeiqi Sep 25, 2026
fcc8915
mtmd: fix mel preprocessor in LFM2 audio (#29403)
ykhrustalev Sep 25, 2026
4b1a27f
common,rpc : simplify fs_create_directory_with_parents() (#29432)
angt Sep 25, 2026
171e884
vendor : update cpp-httplib to 0.58.0 (#29407)
cabelo Sep 25, 2026
4e74811
hexagon: find software divide calls using binary inspection tool (#29…
trivikram-reddy1 Sep 26, 2026
9f70b2c
opencl: add A8 Q8_0 non-MoE dp4a binary kernel (#29439)
shaofeiqi Sep 26, 2026
d834d44
ggml-cpu: tiled mul_mat for k-quants (#27851)
jbooth Sep 26, 2026
965f897
polished Readme and llama-bench (#28968)
truecoder34 Sep 26, 2026
a1de614
jinja : support noncall test statements with arg (#29443)
CISC Sep 26, 2026
08618ff
llama : fix K/V and recurrent state cleanup after failed restores (#2…
CHIPMUNK-T0T Sep 26, 2026
86a24a1
jinja : fix compile error (#29468)
CISC Sep 26, 2026
81bc6b8
jinja : implement sameas test (#29448)
CISC Sep 26, 2026
2145525
Revert "Change max context length for auto-fitting with unified KV (#…
gaugarg-nv Sep 26, 2026
fcb3074
server : fix wake_fd warning on Windows (#29479)
angt Sep 26, 2026
6f856c7
cuda: add F16 input to the FWHT (#29096)
bri-prism Sep 26, 2026
694ec23
musa: build the docker images from the PH1 MUSA SDK image (#29481)
yeahdongcn Sep 26, 2026
9588757
cuda: support Nemotron 3 Puzzle state size 96 for ssm scan (#28717)
anavp-nvidia Sep 26, 2026
2b129cc
hexagon: support for backend sampler (#29502)
max-krasnyansky Sep 27, 2026
7ac59a6
hexagon: support tiled Q4_0 and Q8_0 GET_ROWS (#29511)
kurquhar Sep 27, 2026
85ca3b5
hrm : fix layer placement of `z_l_init` weight (#29512)
ggerganov Sep 27, 2026
187664b
llama-bench : fix OOB access of hf_file (#29515)
angt Sep 27, 2026
7fb2b08
ci : enable GGML_SCHED_DEBUG_REALLOC=1 for ctest workflows (#29514)
ggerganov Sep 27, 2026
d7fb90e
RPC: use RDMA completion channel to not spin (#29440)
am17an Sep 27, 2026
da6c28e
common : throw instead of abort on grammar without llguidance (#29516)
angt Sep 27, 2026
cea7462
vulkan: fix argsort kernel selection for Adreno (#29469)
0cc4m Sep 27, 2026
2ebd9ae
HIP: Enable fattn-mma kernel on cdna for dkq > 256 for large batch si…
IMbackK Sep 27, 2026
36d7b08
CUDA: tune fp16 tile FlashAttention configs for head sizes 40-112 (#2…
animeshsri14 Sep 27, 2026
c829670
sycl: FWHT kernels for block widths above 512 (#29243)
bri-prism Sep 27, 2026
c9064dd
opencl: refine bin kernel loading condition (#29503)
lhez Sep 27, 2026
33c923d
jinja : add support for dict builtin (#29477)
CISC Sep 27, 2026
6fd50a4
ci : bump ty to 0.0.84 (#29529)
CISC Sep 27, 2026
9adc7f4
convert : export YaRN scaling parameters for PLaMo-3 (#29528)
tokinasin Sep 27, 2026
ecf9741
jinja : support sequence indices in selectattr and rejectattr
bitalov Sep 27, 2026
63dded7
model : add K2 Horizon dense and MoVA support
bitalov Sep 27, 2026
93819d9
chat : support K2 Horizon reasoning and tool calls
bitalov Sep 27, 2026
d70e235
Merge the original K2 Horizon branch into the validated integration
bitalov Sep 27, 2026
136887b
common : make string_split<T> throw on invalid values (#29518)
angt Sep 27, 2026
f13798d
conversion: remove obsolete K2 Aurora alias
bitalov Sep 27, 2026
a97cce8
common : avoid side effects around params parsing (#29537)
ggerganov Sep 27, 2026
4da6337
server : allow RANK pooling batch splitting for causal LLM rerankers …
timothywang21 Sep 27, 2026
5262471
vulkan: fix wrong results when a mul_mat reads a slice of a larger ca…
ServeurpersoCom Sep 28, 2026
81ef10e
tests : fix ggml init (#29554)
ggerganov Sep 28, 2026
0c6a6a7
Enables Windows ARM64 build with MSVC cl.exe (#28362)
sarahwu185 Sep 28, 2026
ed7ac35
context : do not re-reserve the scheduler when toggling causal_attn (…
sihanyu03 Sep 28, 2026
4364bf7
metal: support left and circular padding in GGML_OP_PAD (#29561)
ServeurpersoCom Sep 28, 2026
c2a9e16
HIP: fix template skip for DKQ > 256 mfma kernels (#29559)
IMbackK Sep 28, 2026
03a667a
vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain (#…
fxgsell Sep 28, 2026
f916130
ci : ignore more vgpr spills in > 256 DQK fattn kernels (#29571)
IMbackK Sep 28, 2026
6f767fe
ggml-cpu: enable tiled flash attention for non-vector-multiple head d…
SongXiaoXi Sep 28, 2026
d77dd08
tests : refactor test-recurrent-state-rollback (#29426)
ggerganov Sep 28, 2026
f00a64c
webgpu: Handle unaligned writes in ggml_backend_webgpu_buffer_set_te…
jbooth Sep 28, 2026
6c7a87f
common : fix HF cache paths on Windows (#29475)
angt Sep 28, 2026
540e938
k2-horizon: enforce response schemas and load YaRN betas
bitalov Sep 28, 2026
f1ea206
batch: migrate speculative, mtmd and server to batch_ext (#29385)
ngxson Sep 28, 2026
14ebbd5
ggml-openvino: mark unaligned batch-stride views unsupported (#29603)
ravi9 Sep 28, 2026
57b557c
models: pad on the left with ggml_pad_ext (#29567)
ServeurpersoCom Sep 28, 2026
66e665c
vulkan: include functional header (#29597)
LinuxUserGD Sep 28, 2026
680a036
server : support typed content (vision/audio/video) input for /v1/emb…
timothywang21 Sep 28, 2026
1c47294
hex-scripts: show trace events smaller than 100nsec in perfetto (#29614)
trivikram-reddy1 Sep 28, 2026
526c43b
mtmd: fix GCC 15 stringop-overflow in decode_embd_batch (#29607)
ServeurpersoCom Sep 28, 2026
fc07d78
ci : update the oneAPI toolkit to 2026.1 (#29273)
Asahi-Prv Sep 29, 2026
46e17a6
tests : skip pytest workers when PYTEST_WORKERS=1 (#29610)
ggerganov Sep 29, 2026
76a5bc8
common : use fs::path for cache dirs (#29595)
angt Sep 29, 2026
139997d
chat : fix Muse Glimmer ignoring response_format json_schema with --j…
vaibhavdedhia Sep 29, 2026
0bc845d
vulkan : reuse descriptor sets when bindings are constant (#29280)
realrasengan Sep 29, 2026
18bbc46
metal: FWHT perf optimizations (#29602)
bri-prism Sep 29, 2026
6d78fb0
llama : fix init in several tools/examples (#29632)
ggerganov Sep 29, 2026
c8cda8b
ci: remove gpu-rocm keyed directory logs (#28940)
taronaeo Sep 29, 2026
c13e04e
ggml : speed up model loading (#29598)
angt Sep 29, 2026
18b74ff
musa: build the docker image and CI container from the MUSA SDK image…
yeahdongcn Sep 29, 2026
86ea01d
ggml-zdnn: fix 0-row tensor crash (#29636)
taronaeo Sep 29, 2026
8019dc5
ggml : collect all input tensors into graph_inputs (#29634)
ggerganov Sep 29, 2026
31385c9
common : add fs_write_atomic() (#29642)
angt Sep 29, 2026
c85b92c
tests : adjust server string regex to also match m2 utlra results (#2…
ggerganov Sep 29, 2026
00af635
common : use fs::path for config dir (#29649)
angt Sep 29, 2026
ba0ba54
server : remove the built-in UI's service worker when the UI is not s…
erusev Sep 29, 2026
d280808
common : stop accepting draft tokens at EOG (#29638)
ethanhq Sep 29, 2026
e904318
hexagon: add FP32 GELU_ERF and GEGLU_ERF support (#29631)
aparmp-quic Sep 29, 2026
b5cf8ce
ggml : require input tensors to be GGML_OP_NONE (#29647)
ggerganov Sep 29, 2026
284153e
ggml : accumulate f16 dot products in f32 on AVX512-FP16 (#29545)
angt Sep 29, 2026
a3f84fa
vocab : keep </s> NORMAL in PLaMo-2 and PLaMo-3 (#29580)
tokinasin Sep 29, 2026
da89bb3
ggml : fix c++ odr by properly using GGML_COMMON_DECL_CPP (#29504)
XZiar Sep 29, 2026
94a0ae3
vulkan: MOE aware mat_mul_id tile selection (#29182)
Ankk98 Sep 29, 2026
83dd71f
vulkan : Load F32 A matrix 2 at a time when its 2-aligned (#29254)
TheBlueMatt Sep 29, 2026
5c200e0
vulkan: Tune GDN kernel, fix Intel performance (#29476)
0cc4m Sep 29, 2026
6dbbac4
opencl: fix get_tensor for q5_K adreno gemm_nonshuffle kernel (#29555)
lhez Sep 29, 2026
cee37ff
ci: add zdnn backend build but not test (#29541)
taronaeo Sep 29, 2026
748d422
ggml-cuda: HIP: optimize packed byte subtraction (`__vsubss4` -> `__v…
thelittlefireman Sep 29, 2026
6a2743f
CUDA: bitonic argsort handles rows wider than one block (#28957)
ServeurpersoCom Sep 29, 2026
7fee178
hexagon: optimize concat op (#29673)
trivikram-reddy1 Sep 29, 2026
48de2a1
model : support classifier_pooling for rerankers (#29627)
boshjerns Sep 29, 2026
d3954b9
ggml : check row bounds in get_rows_back (#29575)
angt Sep 29, 2026
a6ea155
gguf : reject tensor size that wraps after padding (#26979)
x14ngch3n Sep 29, 2026
19e28a2
Hexagon f16 activation ops (#29209)
cqderek Sep 29, 2026
eae11d2
ggml-zdnn: impl buffer reset, fix memory leaks (#29637)
taronaeo Sep 30, 2026
931351e
vendor: update BoringSSL to 0.20260929.0 (#29669)
cabelo Sep 30, 2026
649dcb1
add GLM-5.3-Flash (GLM5-Next) support (#27773)
timkhronos Sep 30, 2026
2a53ace
SYCL: reduce tensor allreduce sync with pinned host buffers (#29604)
Captain-Tripps Sep 30, 2026
72db1e0
ci : add models backend check (#29651)
CISC Sep 30, 2026
272aad8
musa : define __CUDA_ARCH__ for device passes (#29508)
yeahdongcn Sep 30, 2026
db00347
ci : fix Fusion / metal by adding glm5-next to MTL.csv (#29712)
ServeurpersoCom Sep 30, 2026
25747b0
openvino: serve GET_ROWS on a weight view from the base Constant (#28…
ServeurpersoCom Sep 30, 2026
fa2bde5
ui : type-safe API types, fetch helpers and download-ready models sto…
allozaur Sep 30, 2026
f653250
ui : model id grammar for sidecars, quants and capability parsing (#2…
allozaur Sep 30, 2026
9b43336
ui : Hugging Face Hub data layer (#27947)
allozaur Sep 30, 2026
4cfb6d1
ui : model memory-fit estimation (#27957)
allozaur Sep 30, 2026
8664eae
ui : model download pipeline (#27959)
allozaur Sep 30, 2026
4a096b8
ui : shared model display primitives (#29644)
allozaur Sep 30, 2026
8df332d
model-conversion : add --add-bos to run org model script (#29558)
danbev Sep 30, 2026
cbee4a2
renaming template fixture
bitalov Sep 30, 2026
757e62e
constants conflict fix
bitalov Sep 30, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
6 changes: 3 additions & 3 deletions .devops/intel.Dockerfile
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
ARG ONEAPI_VERSION=2025.3.3-0-devel-ubuntu24.04
ARG ONEAPI_VERSION=2026.1.1-devel-ubuntu24.04
ARG BUILD_DATE=N/A
ARG APP_VERSION=N/A
ARG APP_REVISION=N/A
Expand All @@ -19,7 +19,7 @@ RUN npm ci
COPY tools/ui/ ./
RUN LLAMA_BUILD_NUMBER="$APP_VERSION" npm run build

FROM docker.io/intel/deep-learning-essentials:$ONEAPI_VERSION AS build
FROM docker.io/intel/oneapi-toolkit:$ONEAPI_VERSION AS build

ARG GGML_SYCL_F16=ON
ARG LEVEL_ZERO_VERSION=1.28.2
Expand Down Expand Up @@ -59,7 +59,7 @@ RUN mkdir -p /app/full \
&& cp requirements.txt /app/full \
&& cp .devops/tools.sh /app/full/tools.sh

FROM docker.io/intel/deep-learning-essentials:$ONEAPI_VERSION AS base
FROM docker.io/intel/oneapi-toolkit:$ONEAPI_VERSION AS base

ARG BUILD_DATE=N/A
ARG APP_VERSION=N/A
Expand Down
15 changes: 10 additions & 5 deletions .devops/musa.Dockerfile
Original file line number Diff line number Diff line change
@@ -1,10 +1,9 @@
ARG UBUNTU_VERSION=22.04
# This needs to generally match the container host's environment.
ARG MUSA_VERSION=rc4.3.0
# Target the MUSA build image
ARG BASE_MUSA_DEV_CONTAINER=docker.io/mthreads/musa:${MUSA_VERSION}-devel-ubuntu${UBUNTU_VERSION}-amd64
ARG BASE_MUSA_DEV_CONTAINER=registry.mthreads.com/mcconline/musa_sdk:5.2.0-devel-ubuntu${UBUNTU_VERSION}-s5000

ARG BASE_MUSA_RUN_CONTAINER=docker.io/mthreads/musa:${MUSA_VERSION}-runtime-ubuntu${UBUNTU_VERSION}-amd64
ARG BASE_MUSA_RUN_CONTAINER=registry.mthreads.com/mcconline/musa_sdk:5.2.0-runtime-ubuntu${UBUNTU_VERSION}-s5000

ARG BUILD_DATE=N/A
ARG APP_VERSION=N/A
Expand Down Expand Up @@ -37,7 +36,10 @@ RUN apt-get update && \
python3-pip \
git \
libssl-dev \
libgomp1
libgomp1 \
musa-mualg-5-2 \
musa-muthrust-5-2 \
libmthreads-compute

WORKDIR /app

Expand Down Expand Up @@ -80,13 +82,16 @@ LABEL org.opencontainers.image.created=$BUILD_DATE \
org.opencontainers.image.source=$IMAGE_SOURCE

RUN apt-get update \
&& apt-get install -y libgomp1 curl ffmpeg \
&& apt-get install -y libgomp1 curl ffmpeg libmthreads-compute \
&& apt autoremove -y \
&& apt clean -y \
&& rm -rf /tmp/* /var/tmp/* \
&& find /var/cache/apt/archives /var/lib/apt/lists -not -name lock -type f -delete \
&& find /var/cache -type f -delete

# The MUSA runtime image does not register its library directory
RUN echo "/usr/local/musa/lib" > /etc/ld.so.conf.d/musa-runtime.conf && ldconfig

COPY --from=build /app/lib/ /app

### Full
Expand Down
12 changes: 6 additions & 6 deletions .devops/nix/package.nix
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,7 @@
]
&& blas.meta.available,
useCuda ? config.cudaSupport,
useMetalKit ? stdenv.isAarch64 && stdenv.isDarwin,
useMetalKit ? stdenv.hostPlatform.isAarch64 && stdenv.hostPlatform.isDarwin,
# Increases the runtime closure size by ~700M
useMpi ? false,
useRocm ? config.rocmSupport,
Expand Down Expand Up @@ -92,7 +92,7 @@ let

cudaBuildInputs = with cudaPackages; [
cuda_cudart
cuda_cccl # <nv/target>
cccl # <nv/target>
libcublas
];

Expand Down Expand Up @@ -166,7 +166,7 @@ effectiveStdenv.mkDerivation (finalAttrs: {
# `xcrun` is used find the path of the Metal compiler, which is varible
# and not on $PATH
# see https://github.com/ggml-org/llama.cpp/pull/6118 for discussion
__noChroot = effectiveStdenv.isDarwin && useMetalKit && precompileMetalShaders;
__noChroot = effectiveStdenv.hostPlatform.isDarwin && useMetalKit && precompileMetalShaders;

nativeBuildInputs =
[
Expand All @@ -181,10 +181,10 @@ effectiveStdenv.mkDerivation (finalAttrs: {
autoAddDriverRunpath
]
++ optionals (effectiveStdenv.hostPlatform.isGnu && enableStatic) [ glibc.static ]
++ optionals (effectiveStdenv.isDarwin && useMetalKit && precompileMetalShaders) [ xcrunHost ];
++ optionals (effectiveStdenv.hostPlatform.isDarwin && useMetalKit && precompileMetalShaders) [ xcrunHost ];

buildInputs =
optionals effectiveStdenv.isDarwin darwinBuildInputs
optionals effectiveStdenv.hostPlatform.isDarwin darwinBuildInputs
++ optionals useCuda cudaBuildInputs
++ optionals useMpi [ mpi ]
++ optionals useRocm rocmBuildInputs
Expand Down Expand Up @@ -245,7 +245,7 @@ effectiveStdenv.mkDerivation (finalAttrs: {

# Configurations that are known to result in build failures. Can be
# overridden by importing Nixpkgs with `allowBroken = true`.
broken = (useMetalKit && !effectiveStdenv.isDarwin);
broken = (useMetalKit && !effectiveStdenv.hostPlatform.isDarwin);

description = "Inference of LLaMA model in pure C/C++${descriptionSuffix}";
homepage = "https://github.com/ggml-org/llama.cpp/";
Expand Down
20 changes: 10 additions & 10 deletions .devops/openvino.Dockerfile
Original file line number Diff line number Diff line change
@@ -1,18 +1,18 @@
ARG OPENVINO_VERSION_MAJOR=2026.3
ARG OPENVINO_VERSION_FULL=2026.3.0.22451.bd8d6542e3c
ARG OPENVINO_VERSION_MAJOR=2026.4
ARG OPENVINO_VERSION_FULL=2026.4.0.22959.99c81491cc3
ARG UBUNTU_VERSION=24.04

# Intel GPU driver versions. https://github.com/intel/compute-runtime/releases
ARG IGC_VERSION=v2.38.2
ARG IGC_VERSION_FULL=2_2.38.2+22051
ARG COMPUTE_RUNTIME_VERSION=26.27.39122.11
ARG COMPUTE_RUNTIME_VERSION_FULL=26.27.39122.11-0
ARG IGC_VERSION=v2.40.13
ARG IGC_VERSION_FULL=2_2.40.13+22418
ARG COMPUTE_RUNTIME_VERSION=26.31.39395.13
ARG COMPUTE_RUNTIME_VERSION_FULL=26.31.39395.13-0
ARG IGDGMM_VERSION=22.10.0

# Intel NPU driver versions. https://github.com/intel/linux-npu-driver/releases
ARG NPU_DRIVER_VERSION=v1.35.0
ARG NPU_DRIVER_FULL=v1.35.0.20260722-29947505341
ARG LIBZE1_VERSION=1.28.2-1~24.04~ppa1
ARG NPU_DRIVER_VERSION=v1.38.0
ARG NPU_DRIVER_FULL=v1.38.0.20260910-34487311128
ARG LIBZE1_VERSION=1.32.0-1~24.04~ppa1

# Optional proxy build arguments
ARG http_proxy=
Expand Down Expand Up @@ -173,7 +173,7 @@ RUN --mount=type=cache,target=/var/cache/intel-npu,sharing=locked \
fi; \
DEB=/var/cache/intel-npu/libze1_${LIBZE1_VERSION}_amd64.deb; \
if [ ! -f "$DEB" ]; then \
wget -q -O "$DEB" https://snapshot.ppa.launchpadcontent.net/kobuk-team/intel-graphics/ubuntu/20260606T100000Z/pool/main/l/level-zero-loader/libze1_${LIBZE1_VERSION}_amd64.deb; \
wget -q -O "$DEB" https://snapshot.ppa.launchpadcontent.net/kobuk-team/intel-graphics/ubuntu/20260830T100000Z/pool/main/l/level-zero-loader/libze1_${LIBZE1_VERSION}_amd64.deb; \
fi; \
mkdir /tmp/npu/ && cd /tmp/npu/ && tar -xf "$TGZ" && cp "$DEB" .; \
apt-get update; \
Expand Down
2 changes: 1 addition & 1 deletion .ecrc
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
{
"Exclude": ["^\\.gitmodules$", "stb_image\\.h"],
"Exclude": ["^\\.gitmodules$", "stb_image\\.h", "examples/test-cmake/build/", "examples/test-cmake/build-subdir/"],
"Disable": {
"IndentSize": true
}
Expand Down
2 changes: 1 addition & 1 deletion .github/ISSUE_TEMPLATE/config.yml
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
blank_issues_enabled: true
blank_issues_enabled: false
contact_links:
- name: Got an idea?
url: https://github.com/ggml-org/llama.cpp/discussions/categories/ideas
Expand Down
2 changes: 1 addition & 1 deletion .github/actions/get-tag-name/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ runs:
run: |
BUILD_NUMBER="$(git rev-list --count HEAD)"
SHORT_HASH="$(git rev-parse --short=7 HEAD)"
if [[ "${{ env.BRANCH_NAME }}" == "master" ]]; then
if [[ "${{ env.BRANCH_NAME }}" == "master" || "${{ env.BRANCH_NAME }}" == "b${BUILD_NUMBER}" ]]; then
echo "name=b${BUILD_NUMBER}" >> $GITHUB_OUTPUT
else
SAFE_NAME=$(echo "${{ env.BRANCH_NAME }}" | tr '/' '-')
Expand Down
Loading
Loading