Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
3200 commits
Select commit Hold shift + click to select a range
e1ed5e0
[None][test] Waive 9 failed cases for main in QA CI (#14792)
tensorrt-cicd Jun 3, 2026
aa4276d
[None][test] Update DSV32 32k4k config to avoid timeout issue (#14856)
chenfeiz0326 Jun 3, 2026
bcdf418
[None][chore] Bump version to 1.3.0rc18 (#14872)
yuanjingx87 Jun 3, 2026
66577ac
[None][infra] Waive 5 failed cases for main in post-merge 2755 (#14883)
ZhanruiSunCh Jun 3, 2026
06388ec
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 3, 2026
328ef0b
[None][fix] LTX-2 audio PE pad: use token-axis seq_dim=1 for token-ma…
luyiyun1021 Jun 3, 2026
c798fd9
[#13082][fix] Fix-multimodal embedding mismatch (#13240)
aashirvad08 Jun 3, 2026
e94830c
[None][fix] Pipe stderr separately in subprocess calls to improve err…
yufeiwu-nv Jun 3, 2026
f7f92f4
[None][fix] Use renamed get_param_count_and_checkpoint_size in hybrid…
yufeiwu-nv Jun 3, 2026
8edd72e
[None][test] Remove duplicate test cases in llm_perf_core file (#14749)
yufeiwu-nv Jun 3, 2026
d0cfcde
[None][test] Remove 28 closed-bug waive entries for main (#14545)
tensorrt-cicd Jun 3, 2026
514afc8
[TRTLLM-13022][test] remove deprecated models from tests (#14660)
xinhe-nv Jun 3, 2026
17ccf33
[None][feat] Reserve one more slots for attention_dp in mixed mamba c…
Wanli-Jiang Jun 3, 2026
c938efa
[https://nvbugs/6195110][fix] Restore DeepSeek shared-weights vanilla…
zhaoyangwang-nvidia Jun 3, 2026
a24c3d4
[#12359][feat] AutoDeploy: Support SSM replay kernel for MTP with Fla…
galagam Jun 3, 2026
abc6ba2
[None][test] Waive 1 failed cases for main in QA CI (#14857)
tensorrt-cicd Jun 3, 2026
af9568a
[None][chore] add attention module owner for VisualGen (#14814)
zhenhuaw-me Jun 3, 2026
a336495
[None][fix] release v1 KV blocks on MAX_UTILIZATION pause (#14723)
eopXD Jun 3, 2026
6ab5005
[None][perf] Reduce OpenAI stream postprocess overhead (#14708)
2ez4bz Jun 3, 2026
3959914
[None][fix] propagate chat prompt token ids (#14420) (#14859)
reasonsolo Jun 3, 2026
3630e16
[https://nvbugs/6211193][fix] etcd listen all interfaces (#14863)
reasonsolo Jun 3, 2026
7e8082e
[#5247][fix] auto-detect local cnn_dailymail dataset by directory lay…
guan404ming Jun 3, 2026
e5b8094
[https://nvbugs/6248987][fix] Made the slow-tokenizer swap lazy and i…
tensorrt-cicd Jun 3, 2026
a163d74
[None][chore] redact internal NVIDIA URLs from exec-slurm-compile ski…
ssam18 Jun 3, 2026
b2bb0ad
[None][feat] VisualGen: Attention2D + Ulysses & Multi-GPU LPIPS Evals…
juney-nvidia Jun 3, 2026
fbf66e9
[TRTLLM-13077][feat] Decompose post_load_weights() (#14770)
chienchunhung Jun 3, 2026
7374d1f
[None][fix] Fix config sharing issue for Qwen3-VL (#14766)
2ez4bz Jun 3, 2026
5f64e7d
[https://nvbugs/6104831][fix] Enforce request and buffer index lifecy…
chienchunhung Jun 3, 2026
c16ce24
[None][feat] Add encoder CUDA graph support to llm.encode() (#14326)
tingyangk Jun 4, 2026
6dc60cb
[None][feat] Support Step-3.7-Flash model (#14711)
kaiyux Jun 4, 2026
d64b217
[https://nvbugs/6050489][chore] unwaive tests (#14866)
bo-nv Jun 4, 2026
e1212ad
[None][test] Waive 1 failed cases for main in QA CI (#14896)
tensorrt-cicd Jun 4, 2026
86d08f4
[None][infra] Source code and container vulnerability fix (#14025)
yuanjingx87 Jun 4, 2026
33efef2
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 4, 2026
846d0f4
[None][infra] Waive 11 failed cases for main in post-merge 2757 (#14925)
ZhanruiSunCh Jun 4, 2026
b7ca2a6
[None][test] update rtx6k test list (#14929)
xinhe-nv Jun 4, 2026
00187c0
[None][fix] Update dataset identifier for cnn_dailymail to use namesp…
yufeiwu-nv Jun 4, 2026
864240b
[https://nvbugs/5979673][fix] Unwaive test_agent_multi_backends.py::t…
Shixiaowei02 Jun 4, 2026
4437cb9
[None][fix] Add nemotron-v3 as the proper nemotron-h reasoning parser…
Wanli-Jiang Jun 4, 2026
27af2e5
[https://nvbugs/6193836][test] Use EP=8 + attention DP for minimax_m2…
ruodil Jun 4, 2026
023ac82
[TRTLLM-8236][infra] fix platform tag for public wheel (#14616)
niukuo Jun 4, 2026
0f5dc5a
[None][test] update bug ids in waives (#14946)
xinhe-nv Jun 4, 2026
8b0eba9
[https://nvbugs/6244474][fix] AutoDeploy: skip explicit shape-prop af…
tensorrt-cicd Jun 4, 2026
8c39de8
[None][infra] fix cbts json decode (#14928)
crazydemo Jun 4, 2026
2a934fc
[https://nvbugs/6222480][fix] Fix stress (#14949)
xinhe-nv Jun 4, 2026
941c778
[None][test] Decrease P1 models number and merge sanity test list int…
yufeiwu-nv Jun 4, 2026
8361d42
[None][perf] Use a Triton kernel for Cpp mamba hybrid state update (#…
VALLIS-NERIA Jun 4, 2026
c17611c
[None][chore] Autodeploy unwaive 5888827, 6200112 (#14894)
galagam Jun 4, 2026
222d9e8
[NVBUG-6248780][fix] Add --decoupled flag to benchmark_core_model in …
karljang Jun 4, 2026
0718049
[TRTLLM-12870][feat] Support num_images_per_prompt for FLUX pipelines…
karljang Jun 4, 2026
33b0a32
[TRTLLM-12214][perf] DeepGemmFusedMoE: fuse masked gather + finalize-…
xwang233 Jun 4, 2026
a8c4007
[None][fix] Fix AutoDeploy accuracy tests (#13925)
bmarimuthu-nv Jun 4, 2026
910826b
[TRTLLMINF-69][infra] Migrate A100X-FMHA-Post-Merge-1 and A100X-Trito…
mlefeb01 Jun 4, 2026
8e5d9e2
[TRTLLM-11508][refactor] Merge Eagle3 and MTP-eagle one-model workers…
zhaoyangwang-nvidia Jun 4, 2026
a50b5e2
[https://nvbugs/6143787][fix] Add `kv_cache_config = KvCacheConfig(fr…
tensorrt-cicd Jun 4, 2026
df2d5b9
[https://nvbugs/6248764][fix] Normalize non-sliding KV windows to ful…
eopXD Jun 5, 2026
bd17d1b
[https://nvbugs/6240420][fix] Clamp KV pool window sizes to max_seq_l…
eopXD Jun 5, 2026
81e86a5
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 5, 2026
45d4d54
[None][test] Fix the ci disagg perf local submit test scope too large…
fredricz-20070104 Jun 5, 2026
f0ca418
[None][test] remove outdated model in perf test (#14992)
ruodil Jun 5, 2026
4574851
[TRTLLM-12648][test] implement disagg cancellation injector thread (#…
chienchunhung Jun 5, 2026
21ffdc7
[None][feat] Add AutoDeploy support for StepFun Step-3.7-Flash (#14759)
bmarimuthu-nv Jun 5, 2026
316430f
[None] [waive] Waive the failed step3p7 test case due to ckpt update …
kaiyux Jun 5, 2026
d5de55e
[https://nvbugs/6210714][fix] Fix mamba block calculation (#14524)
VALLIS-NERIA Jun 5, 2026
6818233
[https://nvbugs/5546507][https://nvbugs/5612313][test] Remove obsolet…
xinhe-nv Jun 5, 2026
fdcdcb3
[None][fix] AutoDeploy: Move hf_id_to_local_model_dir to function for…
bmarimuthu-nv Jun 5, 2026
6387eac
[TRTLLM-12893][infra] Parallelize post stages: Rerun Report, Test Cov…
ZhanruiSunCh Jun 5, 2026
2336e47
[None][infra] Waive 11 failed cases for main in post-merge 2760 (#15003)
ZhanruiSunCh Jun 5, 2026
58da60a
[None][fix] Uncomment Qwen3.5 and DSR1 from model registry so that th…
taylor-yb-lee Jun 5, 2026
73c824c
[TRTLLM-11410][feat] Cosmos3 Support (#14824)
NVShreyas Jun 5, 2026
fb5bd44
[https://nvbugs/5859886][fix] Remove the waiver (#14948)
ziyixiong-nv Jun 5, 2026
501b5c2
[https://nvbugs/6248744][fix] Added `trust_remote_code=True` to the `…
tensorrt-cicd Jun 5, 2026
37ece3f
[https://nvbugs/6160629][fix] AutoDeploy: Fix manual seed setting for…
galagam Jun 5, 2026
86f9602
[TRTLLM-12714][feat] KVCacheManagerV2: wire PyExecutor rebalance hook…
thorjohnsen Jun 5, 2026
3e17560
[None][feat] add Wan I2V generation example (#14981)
o-stoner Jun 5, 2026
52ba2bb
[TRTLLM-12527][feat] Parallelize multi-shard visual-gen checkpoint lo…
yibinl-nvidia Jun 5, 2026
d639c57
[https://nvbugs/6250866][fix] Fix deep ep partial warp sync for gptos…
dongfengy Jun 6, 2026
3b21093
[https://nvbugs/6272668][infra] Unwaive DSR1 and Qwen3.5 again (#15010)
taylor-yb-lee Jun 6, 2026
d7a5872
[None][feat] Afmoe trinity support (#13148)
alyosha-swamy Jun 6, 2026
4279e5b
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 6, 2026
e47f26e
[TRTLLM-13027][ci] Relocate under-using tests to right-sized stages (…
QiJune Jun 6, 2026
520262d
[None][feat] Add LTX-2 visual generation example (#14976)
yibinl-nvidia Jun 6, 2026
ec6b284
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 7, 2026
bedad85
[None][feat] AutoDeploy: Fix hardcoded configs (#14943)
taylor-yb-lee Jun 7, 2026
47666de
[#13718][feat] AutoDeploy MoE all-to-all: cache + runtime max-tokens …
greg-kwasniewski1 Jun 7, 2026
428cc3e
[TRTLLM-13177][doc] Add Nemotron 3 Ultra doc (#14964)
nv-guomingz Jun 7, 2026
dcd4e90
[#10710][feat] Make explicit CLI flags take precedence over --config …
marinayanov Jun 7, 2026
8be182d
[https://nvbugs/6260907][fix] unwaive test (#15058)
bo-nv Jun 8, 2026
b8d17d7
[None][chore] Increase GB200-4_GPUs-PyTorch shards (#14836)
tburt-nv Jun 8, 2026
71debd5
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 8, 2026
0e0ee27
[TRTLLM-12648][test] implement disagg cancellation canary thread (#15…
chienchunhung Jun 8, 2026
98a88f7
[TRTLLM-12507][feat] Cudagraph support for routed-expert MoE LoRA wit…
brb-nv Jun 8, 2026
86f33e6
[https://nvbugs/6245317][test] set Harmony tiktoken env for GPT-OSS d…
dongfengy Jun 8, 2026
b4d44d3
[https://nvbugs/6153955][test] unwaive GPT-OSS w4 DP4 CUTLASS (#14884)
dongfengy Jun 8, 2026
ca2bc5e
[None][perf] kv_cache_manager_v2: batch block-key SHA-256 hashing (#1…
lancelly Jun 8, 2026
2cad6db
[TRTLLM-13259][ci] Merge DGX_H100 DeepSeek and GptOss stages (#15035)
QiJune Jun 8, 2026
5fa68a4
[None][infra] Waive 11 failed cases for main in post-merge 2765 (#15080)
ZhanruiSunCh Jun 8, 2026
2632530
[None][infra] Waive 3 failed cases for main in post-merge 2765 (#15082)
ZhanruiSunCh Jun 8, 2026
7e49baa
[None][test] waive weekly qa ci failure cases (#15077)
crazydemo Jun 8, 2026
02f6b2f
[None][feat] AutoDeploy: propagate layer_type hint across pattern-mat…
greg-kwasniewski1 Jun 8, 2026
28dc25e
[None][test] Waive 15 failed cases for main in QA CI (#15056)
tensorrt-cicd Jun 8, 2026
2febb37
[None][infra] Waive 1 failed cases for main in pre-merge 41894 (#15089)
ZhanruiSunCh Jun 8, 2026
09c21b6
[TRTLLM-13262][ci] Move non-default-feature tests to post merge (#15038)
QiJune Jun 8, 2026
c93c63d
[None][feat] Enable disk cache config for KV cache v2 (#14845)
reasonsolo Jun 8, 2026
6dee167
[https://nvbugs/6185446][fix] Add warmup for trtllm-gen fmha JIT kern…
pengbowang-nv Jun 8, 2026
b14794c
[https://nvbugs/6162940][chore] Unwaive fixed test (#15078)
longlee0622 Jun 8, 2026
2bf4d3d
[None][perf] Support Gemma RMSNorm + interleaved mRoPE in fused_qk_no…
nv-guomingz Jun 8, 2026
9eaa468
[None][test] Half K25 Agg Multi Round to Solve Timeout Issue (#15083)
chenfeiz0326 Jun 8, 2026
9af8a16
[None][infra] Reduce Docker image layer count in release stage (#14972)
tburt-nv Jun 8, 2026
cb01607
[#14828][feat] AutoDeploy: support multi KV cache memory pool in trtl…
MrGeva Jun 8, 2026
15d06c0
[None][doc] Refine Nemotron Ultra doc (#15113)
nv-guomingz Jun 8, 2026
8036cde
[None][infra] Waive TestQwen3NextInstruct nvfp4 cases (#15086)
mzweilz Jun 8, 2026
1998324
[https://nvbugs/6248757][fix] Avoid running all reduce in aux stream …
tensorrt-cicd Jun 8, 2026
900d069
[https://nvbugs/6221483][fix] AutoDeploy: Fix Eagle metadata host syn…
govind-ramnarayan Jun 8, 2026
9827c21
[None][feat] add FLUX visual generation examples (#14987)
karljang Jun 8, 2026
b222246
[https://nvbugs/6261164][fix] In the kvcache insert transform (`_Inse…
tensorrt-cicd Jun 8, 2026
c1e9b00
[https://nvbugs/6211189][fix] Lower the reference to 46.5 (matching c…
tensorrt-cicd Jun 9, 2026
bfb4537
[None][refactor] split VisualGen pipeline and model configs (#14956)
bobboli Jun 9, 2026
5e3af40
[TRTLLM-11457][feat] Async Ulysses pipeline (Enabled for LTX-2 + WAN)…
luyiyun1021 Jun 9, 2026
09ebc59
[TRTLLM-11548][doc] Add Qwen3.5 deployment guide doc (#15111)
nv-guomingz Jun 9, 2026
a33dec7
[https://nvbugs/6181383][fix] Build inner text/vision/audio sub-confi…
tensorrt-cicd Jun 9, 2026
041ed83
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 9, 2026
2490441
[https://nvbugs/6273850][chore] waive TestQwen3_5_4B::test_bf16 for a…
tburt-nv Jun 9, 2026
64497e2
[None][doc] Add docs for AutoDeploy transforms (#15122)
bmarimuthu-nv Jun 9, 2026
9349fcc
[None][infra] Waive 4 failed cases for main in post-merge 2769 (#15140)
ZhanruiSunCh Jun 9, 2026
28845dd
[https://nvbugs/6227203][fix] Remove redundant TikTokenTokenizer shim…
tianyuxbear Jun 9, 2026
a197a5e
[None][fix] tunable_fp4_quantize: rename misnamed kwarg + add real SF…
luyiyun1021 Jun 9, 2026
e9402ab
[None][test] Fix gen_only missing prev_device_step_time race in perf …
tensorrt-cicd Jun 9, 2026
6bf3e49
[None][test] Fix disagg test result dir (#14864)
fredricz-20070104 Jun 9, 2026
a7e4a9b
[TRTLLM-13332][test] Remove TestLlama4ScoutInstruct tests (#15144)
QiJune Jun 9, 2026
6f7aea5
[https://nvbugs/6266705][fix] Gate FlashInfer GDN kernels to supporte…
nv-guomingz Jun 9, 2026
6254f3a
[https://nvbugs/6255037][fix] Count DSA indexer K-cache correctly as …
eopXD Jun 9, 2026
a90fd15
[https://nvbugs/6194812][test] Update llm_perf_core.yml to require a …
yufeiwu-nv Jun 9, 2026
34a94ee
[TRTLLMINF-112][infra] Reduce the waiting time between check node is …
EmmaQiaoCh Jun 9, 2026
b852703
[None][infra] Waive 1 failed cases for main in pre-merge 41821 (#15135)
ZhanruiSunCh Jun 9, 2026
178f4e6
[None][infra] CBTS Layer 3: pass test-db via Artifactory instead of e…
crazydemo Jun 9, 2026
45e2523
[TRTLLM-13264][feat] Add native bias epilogue to NVFP4 GEMM (#15053)
luyiyun1021 Jun 9, 2026
2ee96cf
[https://nvbugs/6278380][unwaive] unwaive ad cases (#15148)
crazydemo Jun 9, 2026
104b9d7
[https://nvbugs/6244474][fix] AutoDeploy: Remove llama perf test from…
MrGeva Jun 9, 2026
ba6ba1f
[https://nvbugs/6212252][fix] Select CUTLASS MoE backend on non-Black…
xxi-nv Jun 9, 2026
d620851
[TRTLLM-13302][feat] Register NVIDIA Wan2.2-T2V quantized checkpoints…
zhenhuaw-me Jun 9, 2026
484e6c9
[None][chore] add VisualGen team as the codeowner of the VisualGen At…
zhenhuaw-me Jun 9, 2026
487330e
[None][feat] Default on FlashInferTrtllmGenAttention (#14618)
yihwang-nv Jun 9, 2026
58fbfb9
[None][infra] Test DFW with BSL branch (#14597)
yuanjingx87 Jun 9, 2026
451dbb8
[TRTLLM-12214][perf] customMoeRoutingKernel: lower BLOCK_SIZE to 128,…
xwang233 Jun 9, 2026
f0ba8c7
[TRTLLM-12214][perf] DeepGemmFusedMoE: skip redundant data expand via…
xwang233 Jun 9, 2026
736dc22
[TRTLLM-12648][test] implement disagg cancellation load thread (#15124)
chienchunhung Jun 9, 2026
f1d39ea
[None][fix] Fix regression from SageAttention kernel: Use static sche…
xrq-phys Jun 9, 2026
680c6c4
[TRTLLM-12467][feat] EPD improvements (#13864)
venkywonka Jun 9, 2026
0edbbfe
[None][feat] Expose stored block-hash chain to KV cache connector (#1…
jthomson04 Jun 9, 2026
358505c
[#12805][fix] Fall back to local cache when loading tokenizer for gat…
1MrazorT1 Jun 9, 2026
3ddef66
[None][feat] Support partial RoPE fusion for Hopper kernels in XQA fo…
DomBrown Jun 9, 2026
b0216c6
[None][infra] Add nv-xtf, rahul-steiger-nv, tedzhouhk, tensorrt-cicd …
ZhanruiSunCh Jun 9, 2026
98393f3
[None][feat] Add Prometheus metrics for prompt cache, speculative dec…
vedularaghu Jun 9, 2026
f39a79c
[None][chore] Unwaive DSV32 helix tests (#14871)
brb-nv Jun 9, 2026
e0a909a
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 9, 2026
ddef2d0
[None][fix] unset UCX_TLS=tcp (#15008)
tburt-nv Jun 9, 2026
884520c
[None][feat] Port 13 AutoDeploy custom models to sharding IR + opt th…
greg-kwasniewski1 Jun 9, 2026
48d2b89
[None][chore] Make image paths absolute in blog22 (#15177)
brb-nv Jun 9, 2026
0f7e1db
Fix PyExecutor FPM iteration timing (#14922)
tedzhouhk Jun 9, 2026
ffcd8e6
[#13816][feat] AutoDeploy: Optimize gpt-oss-120b perf (#14202)
taylor-yb-lee Jun 9, 2026
9a7f76f
[None][fix] Register Multimodal Placeholders for Qwen3.5 MoE VLM Serv…
anurags25 Jun 9, 2026
edfc667
[None][feat] Weight trtllm-bench AR/AL averages by output length (#14…
zhaoyangwang-nvidia Jun 10, 2026
3b945f7
[TRTLLM-13052][feat] Enable TRTLLM moe backend for nemotron-h BF16 ck…
Wanli-Jiang Jun 10, 2026
8e40515
[None][fix] Fix and unwaive nemotron related bugs (#15085)
Wanli-Jiang Jun 10, 2026
9c6cb35
[https://nvbugs/6140226][test] Add DFlash coverage for Qwen3.5 MoE va…
yingguo-trt Jun 10, 2026
2763557
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 10, 2026
9c100bb
[None][test] temporarily waive Cosmos3 B200 failures (#15195)
bobboli Jun 10, 2026
5741389
[NVBUG-6241842][fix] DSA DSL atom-split: guard against MTP draft next…
limin2021 Jun 10, 2026
2148a3e
[#11423][feat] AutoDeploy: Basic Disagg Support (#14057)
govind-ramnarayan Jun 10, 2026
9bc4321
[https://nvbugs/6280060][fix] Scope disagg-ctx cache-transfer quorum …
tensorrt-cicd Jun 10, 2026
27b52b3
[None][test] Add e2e example tests for flux1/2, ltx2, wan_i2v, and co…
chang-l Jun 10, 2026
2878b30
[#12632][feat] Add pipeline cache support for AutoDeploy (#13729)
nvchenghaoz Jun 10, 2026
31e730a
[None][test] Add support for nemotron_3_ultra_550b_nvfp4 model in per…
yufeiwu-nv Jun 10, 2026
b206f68
[None][feat] Indexer TopK: single-block / multi-pass radix (#14268)
dcampora Jun 10, 2026
3b46728
[None][fix] Clear workspace in run_mla_generation to avoid potential …
yihwang-nv Jun 10, 2026
90cb7ff
[None][chore] Unwaive AutoDeploy accuracy tests (#14971)
bmarimuthu-nv Jun 10, 2026
2d196f7
[None][test] Increase kv_transfer_timeout_ms for b200 deepseek-r1 dis…
tensorrt-cicd Jun 10, 2026
62c6521
[None][feat] Enable MTP for Step-3.7 NVFP4 and port Step-3.7VL vision…
kaiyux Jun 10, 2026
9635f7d
[https://nvbugs/6266370][fix] Fix MAX_UTILIZATION reuse token budget …
brb-nv Jun 10, 2026
69f5add
[https://nvbugs/6272573][ci] Unwaive skipped test (#15118)
2ez4bz Jun 10, 2026
74d8a48
[https://nvbugs/6245279][fix] AutoDeploy: Unwaive accuracy tests (#15…
galagam Jun 10, 2026
6db3233
[TRTLLM-12491][feat] Align VisualGen serve request schema with Visual…
zhenhuaw-me Jun 10, 2026
dab3400
[None][test] Add MLA chunked-prefill SM dispatch regression coverage …
DhineshPonnarasan Jun 10, 2026
0be1447
[TRTLLM-12648][test] enable disagg cancellation stress test (#15174)
chienchunhung Jun 10, 2026
03ed843
[None][feat] Preserve cache_salt string in KV cache events (#13051)
jthomson04 Jun 10, 2026
309c764
[https://nvbugs/6104831][fix] Port dataTransceiver shared_ptr<LlmRequ…
chienchunhung Jun 10, 2026
7301075
[None][fix] Fix AutoDeploy transform docs generation (#15228)
bmarimuthu-nv Jun 10, 2026
bb74da1
[None][feat] Targeted warmup-waste cleanup (#14609)
dominicshanshan Jun 11, 2026
0d44f33
[None][fix] Remove TLLM_RUBIN_FEATURES (#15143)
yuxianq Jun 11, 2026
cab198d
[https://nvbugs/6108994][fix] add kv_transfer_timeout_ms to avoid tim…
bo-nv Jun 11, 2026
af5a22e
[TRTLLM-12657][infra] Fix periodic-junit in unittest pytest (#14075)
yiqingy0 Jun 11, 2026
5e3f012
[https://nvbugs/6143883][fix] Preserve ip:port for trtllm-serve visua…
JunyiXu-nv Jun 11, 2026
228829c
[TRTLLM-12958][feat] Enable gen-only spec dec (#14546)
bo-nv Jun 11, 2026
55bbcf4
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 11, 2026
016fb4c
[None][test] Remove 78 closed-bug waive entries for main (#15061)
tensorrt-cicd Jun 11, 2026
205920d
[https://nvbugs/6278399][fix] Add x86_64 path using CU_MEM_HANDLE_TYP…
tensorrt-cicd Jun 11, 2026
a622e30
[TRTLLM-11538][feat] Blackwell custom mask fmha support (#12958)
sunnyqgg Jun 11, 2026
01d8ccb
[None][infra] Waive 6 failed cases for main in post-merge 2773 (#15250)
ZhanruiSunCh Jun 11, 2026
d96c0df
[None][feat] Enhance CuteDSL NVF4 MOE (#15092)
liyuhannnnn Jun 11, 2026
6b7d8cf
[None][infra] Waive 3 failed cases for main in post-merge 2772 (#15253)
ZhanruiSunCh Jun 11, 2026
84b349f
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 11, 2026
835fd61
[None][test] Update K2.5 andGLM-5 into CI Perf Test (#14960)
chenfeiz0326 Jun 11, 2026
54dec4f
[None][feat] enable GQA and cross-attention for attn2d (#14961)
NVShreyas Jun 11, 2026
d3748a3
[#12230][fix] Add bounds checking in autotuner _find_nearest_profile …
mihai-chiorean Jun 11, 2026
80f18fe
[None][refactor] visual_gen Attention: drop redundant enable_ulysses …
luyiyun1021 Jun 11, 2026
9ab3501
[None][fix] Generalize FP8 checkpoint loading for Qwen3.5 (#15067)
amukkara Jun 11, 2026
19b5d0e
[#13858][fix] AutoDeploy fix the piecewise vlm issue (#14006)
nvchenghaoz Jun 11, 2026
be7117c
[TRTLLM-12507][feat] Cudagraph support for per-expert lora in Cutlass…
brb-nv Jun 11, 2026
c859650
[None][test] Remove stale perf sanity waives (#15269)
cascade812 Jun 11, 2026
81f7baf
[None][infra] Waive 8 failed cases for main in pre-merge 42699 (#15273)
ZhanruiSunCh Jun 11, 2026
d19b6a8
[None][fix] Install processor-output validation filter at module impo…
aswinvisva Jun 11, 2026
67e1097
[None][infra] Waive 10 failed cases for main in pre-merge 42753 (#15275)
ZhanruiSunCh Jun 11, 2026
eb5674b
[TRTLLM-12534][fix] Nemotron Nano - properly account for text prompts…
moraxu Jun 11, 2026
b586ebf
[None][doc] Fix stale --disable_xqa reference in legacy docs (#13395)
Erfandarzi Jun 11, 2026
1b360ee
[TRTLLM-11403][doc] Cache-DiT documentation (#15268)
o-stoner Jun 11, 2026
00ed78c
[#15022][fix] Guided decoding (xgrammar) + EAGLE-3 + draft_len_schedu…
chungen04 Jun 11, 2026
ccc0708
[TRTLLM-12154][test] Add Qwen3-32B FP8 disagg stress test (#14278)
brnguyen2 Jun 11, 2026
aef7d47
[TRTLLM-13141][feat] Add backend-agnostic SourceIdentity gate for wei…
chienchunhung Jun 12, 2026
ae9226e
[None][feat] Add PyTorch reset_prefix_cache API (#14970)
milesial Jun 12, 2026
2dd5c67
[None][fix] Stabilize Mamba replay state update (#14841)
sunnyqgg Jun 12, 2026
5a77356
[None][infra] Waive remaining AutoDeploy Disagg tests until fix lands…
govind-ramnarayan Jun 12, 2026
82ddf75
[None][test] Sunset the old disagg test cases for the qa side (#15290)
fredricz-20070104 Jun 12, 2026
02957e9
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 12, 2026
82d3811
[None][infra] Waive 1 failed cases for main in pre-merge 42836 (#15293)
ZhanruiSunCh Jun 12, 2026
2b3af6d
[None][fix] Fix max_context_length value for attention workspace sizi…
pengbowang-nv Jun 12, 2026
fb7a1d0
[TRTLLM-12038][feat] Add accuracy tests for nemotron-v3-ultra (#14808)
Wanli-Jiang Jun 12, 2026
8e2b7b2
[#14672][fix] AutoDeploy: Vendor OpenELMConfig locally to fix OpenELM…
plapagesse Jun 12, 2026
c323881
[https://nvbugs/6035425][fix] Fix KV cache host splitting logic (#14373)
mikeiovine Jun 12, 2026
85d5e6e
[None][refactor] Move KV cache manager V2 to separate file (#14680)
jiaganc Jun 12, 2026
f18d18d
[TRTLLM-12963][refactor] LTX-2 attention: drop dead k_pe parameter; r…
luyiyun1021 Jun 12, 2026
44550bc
[TRTLLM-10184][chore] Remove legacy XQA precompiled code path (#14941)
pengbowang-nv Jun 12, 2026
82ca2c5
[TRTLLM-35882][feat] cute dsl gvr-top multi-cta optimization (#15198)
limin2021 Jun 12, 2026
db7161b
[None][fix] Revert "Add PyTorch reset_prefix_cache API (#14970)" (#15…
xxi-nv Jun 12, 2026
b03b78f
Revert "[None][test] Add support for nemotron_3_ultra_550b_nvfp4 mode…
tburt-nv Jun 12, 2026
646464b
[https://nvbugs/6309375][test] AutoDeploy: Remove stale fallback test…
govind-ramnarayan Jun 12, 2026
19ae053
[None][fix] AutoDeploy: set enable_spec_decode on ADEngine for disagg…
Shixiaowei02 Jun 12, 2026
380d96a
[TRTLLM-12498][feat] Add support for beam search in disaggregated ser…
athena-nv Jun 12, 2026
be7e978
[None][chore] 2 more WAN multi-gpu tests (#15223)
NVShreyas Jun 12, 2026
57bb6ee
[TRTLLM-12721][feat] Add disagg transfer state consensus (#15139)
chienchunhung Jun 12, 2026
c2b7cd9
[None][infra] Waive 1 failed cases for main in pre-merge 43047 (#15326)
ZhanruiSunCh Jun 12, 2026
cd65070
[#12715][fix] disable NCCL_SYMMETRIC tactic on GB10 (DGX Spark) (#12902)
nv-lschneider Jun 13, 2026
706a91f
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 13, 2026
bb32597
[None][feat] AutoDeploy: Qwen3.5: Apply whielist based sharding and a…
taylor-yb-lee Jun 13, 2026
ec47baa
[https://nvbugs/6293015][fix] Add a delegating `@property def vocab_s…
tensorrt-cicd Jun 13, 2026
4e1776a
[TRTLLM-12842][feat] Maximal LLMAPI capture in usage telemetry (#14398)
venkywonka Jun 13, 2026
1283c6b
[TRTLLM-12427][perf] Qwen2.5/3/3.5-VL Performance Optimization (#11943)
yechank-nvidia Jun 13, 2026
a878a51
Add CPU-only fork CI mirror for pre-commit checks
CodersAcademy006 Jul 8, 2026
00e7af5
Test push
CodersAcademy006 Jul 8, 2026
a392a3f
Fix workflow condition logic for PR and comment triggers
CodersAcademy006 Jul 8, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
The diff you're trying to view is too large. We only load the first 3000 changed files.
80 changes: 80 additions & 0 deletions .claude/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,80 @@
# Custom Claude Code Skills & Agents for TensorRT-LLM

## Background: Skills & agents in Claude Code

Claude Code supports two extensibility mechanisms — **skills** and **agents** —
that let teams encode domain expertise into reusable, version-controlled
components.

**Skills** are markdown playbooks that Claude follows step-by-step when
triggered. They are invoked via `/slash-commands` (e.g. `/perf-analysis`) or
matched automatically from natural-language requests. Each skill lives in its
own directory under `.claude/skills/` and can bundle reference materials that
Claude reads during execution. See
[Custom slash commands](https://code.claude.com/docs/en/skills)
for details.

**Agents** (sub-agents) are specialist workers that Claude spawns in a separate
context to handle focused tasks. Each agent has its own system prompt, tool
access, and domain knowledge. Claude delegates to them when it determines a task
fits a specialist's scope, while you can also invoke agents directly. Agent
definitions live under `.claude/agents/`. See
[Custom sub-agents](https://code.claude.com/docs/en/sub-agents)
for details.

## How skills and agents are loaded

For users who are working with Claude Code under TensorRT-LLM project directory,
skills and agents are automatically discovered by Claude Code at startup — no
manual registration needed. Files placed in `.claude/skills/` and
`.claude/agents/` are picked up by convention.

To verify what's loaded, launch Claude Code under TensorRT-LLM project directory
and type `/skills` or `/agents` in the Claude Code prompt to see available
skills and sub-agents.

## How to use skills and agents

There are two ways to trigger skills and agents:

1. **Automatic dispatch** — just describe what you need in plain language
(e.g. "profile this workload", "compile TensorRT-LLM"). Claude Code will
match your request to the appropriate skill or delegate to the right
sub-agent automatically.

2. **Manual invoke** — type `/<skill-name>` (e.g. `/perf-analysis`,
`/trtllm-serve-config-guide`) to explicitly run a skill. For sub-agents, type
`@"<agent-name>" (agent)` (e.g. `@"exec-compile-specialist (agent)"`) to
delegate a task directly. This is useful when you know exactly which workflow you want.

In most cases, automatic dispatch is sufficient — you don't need to memorize
skill or agent names. Manual invoke is there for when you want precise control.

References:
* [Extend Claude with skills](https://code.claude.com/docs/en/skills)
* [Work with subagents](https://code.claude.com/docs/en/sub-agents#work-with-subagents)

## Naming convention

Every skill and agent name uses the format `<prefix>-<descriptive-name>`.
The prefix identifies the primary work area; the descriptive part should be
short and not repeat it.

| Prefix | Domain | Definition |
|---|---|---|
| `ad-` | AutoDeploy | Model onboarding, pipeline debugging, and execution for the AutoDeploy backend |
| `exec-` | Execution infra | Environment setup and job execution (compile, run, container) |
| `kernel-` | Kernel development | Kernel writing, generation, and kernel-specific transforms |
| `perf-` | Performance work | Profiling, analysis, and tuning above the kernel layer (kernel modifications belong under `kernel-`) |
| `trtllm-` | TRT-LLM project workflows | Project-specific workflows: codebase exploration, contribution, dependency upgrades, and serving configuration (static subsystem knowledge belongs in repo docs) |

Guidelines:

* If a skill doesn't fit any prefix, propose a new one and agree on its
boundary before using it.
* Use the prefix of the skill's **primary** domain, even if it orchestrates
across multiple domains.
* Agents follow the same convention.
* Good: `exec-local-compile`, `kernel-cuda-writing`, `perf-host-analysis`
* Bad: `exec-trtllm-compile`, `kernel-cuda-kernel-writing`,
`perf-trtllm-host-analysis`
30 changes: 30 additions & 0 deletions .claude/agent-tests/perf-test-sync/build_prompt.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
"""Promptfoo prompt builder.

Loads perf-test-sync.md, strips Claude Code specific sections, and appends the
user request.

Stripped sections:
- YAML frontmatter between leading `---` markers
- `# Persistent Agent Memory` section and everything after it (Claude Code
memory infrastructure, not relevant to prompt-quality evaluation)
"""

import os
import re

_SCRIPT_DIR = os.path.dirname(os.path.abspath(__file__))
_AGENT_MD = os.path.normpath(os.path.join(_SCRIPT_DIR, "..", "..", "agents", "perf-test-sync.md"))


def _load_agent_body() -> str:
with open(_AGENT_MD, "r", encoding="utf-8") as f:
text = f.read()
text = re.sub(r"\A---\n.*?\n---\n", "", text, count=1, flags=re.DOTALL)
text = text.split("# Persistent Agent Memory", 1)[0].rstrip()
return text


def build(context: dict) -> str:
user_prompt = context["vars"]["prompt"]
agent_body = _load_agent_body()
return f"{agent_body}\n\n## User request\n\n{user_prompt}\n"
Loading