TensorRT-LLM v0.17 Release - #2725
Conversation
|
This release mentions |
Sorry for the mis-leading. The June |
|
I see. Perhaps the |
|
Are you sure you included the enc dec FP8 in the readme? Because it's not in the 0.17 branch or the main branch |
Thanks for the suggestion. The The reason that this time the release/0.17 branch contains more new features than main is due to the Blackwell release which makes the github update process a little bit special :) Thanks |
Thanks for catching this, @MahmoudAshraf97 . I just communicated with the team and there are something wrong about the doc update process for 0.17 release. We just updated the 0.17 release info now. Pls help review it to see whether there is anything else incorrect. Thanks again for reminding us about this doc issue. June |
Co-authored-by: Kaiyu Xie <26294424+kaiyux@users.noreply.github.com> open source f8c0381a2bc50ee2739c3d8c2be481b31e5f00bd (#2736) Co-authored-by: Kaiyu Xie <26294424+kaiyux@users.noreply.github.com> Add note for blackwell (#2742) Update the docs to workaround the extra-index-url issue (#2744) update README.md (#2751) Fix github io pages (#2761) Update
|
Hi @schetlur-nv the file https://github.com/NVIDIA/TensorRT-LLM/blob/v0.16.0/cpp/tensorrt_llm/kernels/mixtureOfExperts/moe_kernels.cu has been removed from the repo in release v0.17.0. I wonder how deepseek fused gate is supported by our engineer team. Thank you ! |
TensorRT-LLM Release 0.17.0 [UPDATED 1/31]
Key Features and Enhancements
LLMAPI andtrtllm-benchcommand.tensorrt_llm._torch. The following is a list of supported infrastructure, models, and features that can be used with the PyTorch workflow.LLMAPI.examples/multimodal/README.md.userbufferbased AllReduce-Norm fusion kernel.executorAPI.API Changes
paged_context_fmhais enabled.--concurrencysupport for thethroughputsubcommand oftrtllm-bench.Fixed Issues
cluster_keyfor auto parallelism feature. ([feature request] Can we add H200 in infer_cluster_key() method? #2552)__post_init__function ofLLmArgsClass. Thanks for the contribution from @topenkoff in Fix kwarg name #2691.Infrastructure Changes
nvcr.io/nvidia/pytorch:25.01-py3.nvcr.io/nvidia/tritonserver:25.01-py3.