-
Notifications
You must be signed in to change notification settings - Fork 106
Pull requests: InfiniTensor/InfiniLM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
feat(nvidia): add Mamba2 and RWKV5 adapters
#591
opened Sep 20, 2026 by
lgeln10
Loading…
7 of 8 tasks
feat(nvidia): 为 v0.2.9c 增加可复用的 GGUF Route B 支持
#589
opened Sep 20, 2026 by
xindongliu594
Loading…
40 of 46 tasks
feat: add mtp draft models for speculative decoding
#588
opened Sep 20, 2026 by
accelerator-llc
Loading…
34 of 41 tasks
feat: add FP8 blockwise weight quantization and FP8(E4M3) KV cache support
#586
opened Sep 20, 2026 by
shsaihdsaiudh
Loading…
38 of 42 tasks
perf: Qwen3 offline inference — hybrid attention backend, fused rms_norm_rope, chunked prefill + prompt-lookup spec decoding
#585
opened Sep 20, 2026 by
shsaihdsaiudh
Loading…
feat(qwen): add paged greedy MTP with FP8 block weights
#584
opened Sep 19, 2026 by
big-hip
Loading…
40 of 41 tasks
feat(minimax): support MiniMax-Text-01 and MiniMax-M2
#583
opened Sep 19, 2026 by
y258dd
Loading…
25 of 42 tasks
feat(quantization): support compressed-tensors W8A8 and mixed modules
#582
opened Sep 19, 2026 by
sam0336
Loading…
35 of 44 tasks
feat(nvidia): support model type: granitemoehybrid
#580
opened Sep 18, 2026 by
WHoutstanding
Loading…
41 of 49 tasks
perf: reuse running reservations within each scheduler step
#579
opened Sep 17, 2026 by
T4t4KAU
Loading…
9 of 11 tasks
perf: bound KV page reclamation candidate scans
#578
opened Sep 17, 2026 by
T4t4KAU
Loading…
8 of 10 tasks
perf: build KV slot mappings by page ranges
#577
opened Sep 17, 2026 by
T4t4KAU
Loading…
9 of 11 tasks
perf: avoid scanning KV cache pages for capacity checks
#576
opened Sep 17, 2026 by
T4t4KAU
Loading…
15 of 22 tasks
feat(mamba2): integrate model loading and recurrent inference
#575
opened Sep 17, 2026 by
big-hip
Loading…
38 of 39 tasks
feat(engine): integrate prefix eviction and chunked parallel execution
#573
opened Sep 16, 2026 by
big-hip
Loading…
32 of 41 tasks
feat: add priority-aware request scheduling
#571
opened Sep 13, 2026 by
xiaoba17
Loading…
26 of 36 tasks
fix(moore): enable Mate flash-attn inference
#567
opened Sep 10, 2026 by
spike-zhu
Collaborator
Loading…
1 of 49 tasks
feat(nvidia): add reusable GGUF Route B support for Qwen3.5
#559
opened Sep 3, 2026 by
xindongliu594
Loading…
32 of 48 tasks
perf: fuse unquantized Qwen3-Next GDN projections
#555
opened Sep 2, 2026 by
T4t4KAU
Loading…
26 of 39 tasks
feat(server): add agent support with tool-call and reasoning parsing
#554
opened Sep 1, 2026 by
rubik-hua
Contributor
Loading…
Previous Next
ProTip!
no:milestone will show everything without a milestone.