v0.10.1 #197
nvluxiaoz
announced in
Announcements
v0.10.1
#197
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
TensorRT Edge-LLM 0.10.1 Release 2026-09-03
We are excited to announce the release of TensorRT Edge-LLM 0.10.1!
TensorRT Edge-LLM 0.10.1 adds experimental TP=2 inference across two NVIDIA DGX Spark systems, JetSpec speculative decoding for Qwen3-8B, Qwen3.8-27B DSpark support, end-to-end FP8 ViT attention and DART visual-token pruning. This release also redesigns the experimental OpenAI-compatible server to achieve a 60% - 99% reduction of overall disk usage, and a 50% reduction of cold launch time.
Key Features
Other Important Features
Runtime and Performance
Export and Quantization
Server and API
Platform Support and Documentation
Bug Fixes
NVIDIA Contributors
@nvluxiaoz @nvamberl @xiangg-nv @willg-nv @mahu888 @nv-samcheng @duofant @JCalafato @duanyaqi @zhazhang-nv @zhaoyuanh-nvidia @zhijial-nvidia @ever-wong @Jasper-NV @jhalabi-nv @levichen-nvidia @sunghyunp-nvidia @nvyocox @ruocheng-nv @yuanyao-nv @qikail-ctrl @MadFunMaker @kevinfu1
This discussion was created from the release v0.10.1.
All reactions