[https://nvbugs/6315845][fix] Dual-location revert of #14851's protective-code removal: (1) pin… - #15472
Open
chenfeiz0326 wants to merge 1 commit into
Open
[https://nvbugs/6315845][fix] Dual-location revert of #14851's protective-code removal: (1) pin…#15472chenfeiz0326 wants to merge 1 commit into
chenfeiz0326 wants to merge 1 commit into
Conversation
…oval: (1) pin `mMaxSeqLenKv Signed-off-by: Chenfei Zhang <chenfeiz@nvidia.com>
Contributor
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (2)
📝 WalkthroughWalkthroughIn two TRTLLM-GEN FMHA runner configuration sites, ChangesmMaxSeqLenKv upper-bound pinning for TRTLLM-GEN FMHA
Estimated code review effort🎯 2 (Simple) | ⏱️ ~10 minutes Possibly related PRs
Suggested reviewers
🚥 Pre-merge checks | ✅ 3 | ❌ 2❌ Failed checks (1 warning, 1 inconclusive)
✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
mMaxSeqLenKvpinning in BOTHmlaGeneration(cpp/tensorrt_llm/common/attentionOp.cpp:1133) andXqaDispatcher::runImpl(cpp/tensorrt_llm/kernels/xqaDispatcher.cpp:~509), replacing them with per-iterationmax_past_kv_length. The dynamic value cycles through seqLenKv buckets the warmup grid never covered, triggering 1-15 NVRTC re-compiles of 7-9 s each and blowing TTFT P99 from 882 ms to 25 s while decode (TPOT/ITL) stays within 1%. Prior attempts addressed onlymlaGeneration(or only the warmup grid).mMaxSeqLenKv = max_attention_window_sizeinmlaGeneration(safe for PagedKv — strides independent, KV CTAs early-exit via seqLensKvPtr); (2) pinmMaxSeqLenKv = (mQkvLayout == PagedKv) ? max_attention_window_size : max_past_kv_lengthinXqaDispatcher::runImplnon-spec-dec-tree branch — QkvLayout-aware so ContiguousKv keeps its true past-kv length. C++-only, two files, 18 insertions / 2 deletions, no Python or warmup-grid changes.Test plan
Links
Summary by CodeRabbit