-
-
Notifications
You must be signed in to change notification settings - Fork 278
Pull requests: Neroued/ninfer
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
kv: thread compute stream through activate() so membership publish is stream-ordered
#211
opened Sep 7, 2026 by
ranxianglei
Loading…
perf(ops): warm L2 for the next projection from the MoE down tail, issuing the hint before the block barrier
#206
opened Sep 7, 2026 by
MichaelDementii
Contributor
Loading…
perf(text): skip NFC normalisation when the text is already pure ASCII
#205
opened Sep 7, 2026 by
MichaelDementii
Contributor
Loading…
perf(nvfp4): rasterise the W4A4 TMA CTAs token-fastest, and fetch each activation-scale box once
#204
opened Sep 7, 2026 by
MichaelDementii
Contributor
Loading…
perf(ops): prefetch the shared-expert down weights from the D1 tail
#203
opened Sep 7, 2026 by
MichaelDementii
Contributor
Loading…
perf(runtime): keep the Linear Attention state in L2 across the decode graph
#202
opened Sep 7, 2026 by
MichaelDementii
Contributor
Loading…
perf(ops): per-schedule activation cache policy in w8_rowsplit_gemm_mma, with ca on the BM=16 tiles
#201
opened Sep 7, 2026 by
MichaelDementii
Contributor
Loading…
perf(sparse_moe): one measured Rows2 window of nineteen tokens for both routed-down codecs
#200
opened Sep 7, 2026 by
MichaelDementii
Contributor
Loading…
perf(sparse_moe): keep two Q4 group quads in flight in the routed gate/up dot product
#199
opened Sep 7, 2026 by
MichaelDementii
Contributor
Loading…
perf(moe): choose the routed gate/up pipeline depth by route, size the grid by work
#198
opened Sep 7, 2026 by
MichaelDementii
Contributor
Loading…
Fall back to a preset of the same weights format when no (model, weights) row matches: prefill-cost prediction goes from 3.1x to 1.15x on a registered artifact
#195
opened Sep 6, 2026 by
MichaelDementii
Contributor
Loading…
perf(nvfp4): drop the guarded expf slow path from the fused SwiGLU epilogue
#194
opened Sep 6, 2026 by
MichaelDementii
Contributor
Loading…
perf(frontend): store the BPE merge rules in a flat open-addressed table
#193
opened Sep 6, 2026 by
MichaelDementii
Contributor
Loading…
feat(serve): override the frontend chat template via --chat-template FILE
#183
opened Sep 5, 2026 by
wojciak
Loading…
feat(kv): rk2v4-e8 compressed-KV (E8-root, 208 B/head-token) on the paged-KV engine
#173
opened Sep 4, 2026 by
danielfparkernz
Loading…
The fp8 A8 GEMM stages its operands through TMA: prefill +2.7% at chunk 1024 and +4.7% at 4096
#167
opened Sep 3, 2026 by
MichaelDementii
Contributor
Loading…
feat(serve): expose llama.cpp-compatible model metadata on /v1/models
#162
opened Sep 2, 2026 by
hecrj
Loading…
The NVFP4 TMA route reads activation scales as one tile: operator 0.81x to 0.87x, prefill +4.1%, bitwise identical
#160
opened Sep 2, 2026 by
MichaelDementii
Contributor
Loading…
feat(serve): automatic shared-prefix write at the system/developer frontier
#152
opened Sep 1, 2026 by
Astrangemaninhere
Loading…
OpenAI Responses API: compatible with reasoning summary and encryption
#148
opened Sep 1, 2026 by
Sha1rholder
Loading…
fix(qwen3.8): wire-format detect nvfp4 artifact profile
#107
opened Aug 28, 2026 by
koloved
Loading…
build: cache C++ and CUDA compilation in container builds
#97
opened Aug 26, 2026 by
DuncanBetts
Loading…
feat(platform): native Windows (MSVC + CUDA) build for ninfer-serve
#84
opened Aug 22, 2026 by
devan-carlin
Loading…
Previous Next
ProTip!
Find all pull requests that aren't related to any open issues with -linked:issue.