[https://nvbugs/6029882][fix] Fix attentionOp fp8 mla kvreuse workspace calculation - #14852
Conversation
|
/bot help |
GitHub Bot Help
Provide a user friendly way for developers to interact with a Jenkins server. Run See details below for each supported subcommand. Details
Launch build/test pipelines. All previously running jobs will be killed.
kill
Kill all running builds associated with pull request. skip
Skip testing for latest commit on pull request. reuse-pipeline
Reuse a previous pipeline to validate current commit. This action will also kill all currently running builds associated with the pull request. IMPORTANT NOTE: This is dangerous since lack of user care and validation can cause top of tree to break. |
📝 WalkthroughWalkthroughA test waiver entry for a DeepSeek NVFP4 multi-GPU throughput test variant is removed from the skip list, allowing this previously skipped test case to run. ChangesTest waiver removal
Estimated code review effort🎯 1 (Trivial) | ⏱️ ~2 minutes Possibly related PRs
Suggested reviewers
🚥 Pre-merge checks | ✅ 3 | ❌ 2❌ Failed checks (2 warnings)
✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
|
/bot run --stage-list "GB200-8_GPUs-2_Nodes-PyTorch-1,GB200-8_GPUs-2_Nodes-PyTorch-2" |
|
PR_Github #51523 [ run ] triggered by Bot. Commit: |
|
/bot run --stage-list "GB200-8_GPUs-2_Nodes-PyTorch-2" |
|
PR_Github #51548 [ run ] triggered by Bot. Commit: |
|
PR_Github #51523 [ run ] completed with state
|
|
PR_Github #51548 [ run ] completed with state |
|
/bot run --stage-list "GB200-8_GPUs-2_Nodes-PyTorch-2" --disable-reuse-test |
|
PR_Github #51580 [ run ] triggered by Bot. Commit: |
|
PR_Github #51580 [ run ] completed with state
|
|
/bot run --stage-list "GB200-8_GPUs-2_Nodes-PyTorch-2" --disable-reuse-test |
|
PR_Github #51611 [ run ] triggered by Bot. Commit: |
|
PR_Github #51611 [ run ] completed with state
|
|
/bot run --stage-list "GB200-8_GPUs-2_Nodes-PyTorch-2" --disable-reuse-test |
|
PR_Github #51733 [ run ] triggered by Bot. Commit: |
|
PR_Github #51733 [ run ] completed with state
|
|
/bot run --stage-list "GB200-8_GPUs-2_Nodes-PyTorch-2" --disable-reuse-test |
|
PR_Github #51751 [ run ] triggered by Bot. Commit: |
|
/bot run --stage-list "GB200-8_GPUs-2_Nodes-PyTorch-2" --disable-reuse-test |
|
PR_Github #51751 [ run ] completed with state
|
|
/bot run --stage-list "GB200-8_GPUs-2_Nodes-PyTorch-2" |
|
PR_Github #51776 [ run ] triggered by Bot. Commit: |
|
PR_Github #51776 [ run ] completed with state
|
8b9465d to
7779960
Compare
|
/bot run --stage-list "GB200-8_GPUs-2_Nodes-PyTorch-1,GB200-8_GPUs-2_Nodes-PyTorch-2,GB200-8_GPUs-2_Nodes-PyTorch-3,GB200-8_GPUs-2_Nodes-PyTorch-4,DGX_B300-4_GPUs-PyTorch-Post-Merge-2,GB300-4_GPUs-PyTorch-Post-Merge-2 ,DGX_B200-8_GPUs-PyTorch-1" |
|
PR_Github #53896 [ run ] completed with state
|
Signed-off-by: Pengbo Wang <221450789+pengbowang-nv@users.noreply.github.com>
Signed-off-by: Pengbo Wang <221450789+pengbowang-nv@users.noreply.github.com>
9275333 to
615aee7
Compare
|
/bot run --add-multi-gpu-test --disable-fail-fast --extra-stage "DGX_B200-8_GPUs-PyTorch-1, DGX_B200-8_GPUs-PyTorch-3, GB200-8_GPUs-2_Nodes-PyTorch-1, GB200-8_GPUs-2_Nodes-PyTorch-2, DGX_B300-4_GPUs-PyTorch-Post-Merge-1, GB300-4_GPUs-PyTorch-Post-Merge-2" |
|
PR_Github #54014 [ run ] triggered by Bot. Commit: |
|
PR_Github #54014 [ run ] completed with state
|
|
/bot run --add-multi-gpu-test --disable-fail-fast --extra-stage "DGX_B200-8_GPUs-PyTorch-1, DGX_B200-8_GPUs-PyTorch-3, GB200-8_GPUs-2_Nodes-PyTorch-1, GB200-8_GPUs-2_Nodes-PyTorch-2, DGX_B300-4_GPUs-PyTorch-Post-Merge-1, GB300-4_GPUs-PyTorch-Post-Merge-2" |
|
PR_Github #54137 [ run ] triggered by Bot. Commit: |
|
PR_Github #54137 [ run ] completed with state
|
|
/bot run --add-multi-gpu-test --disable-fail-fast --extra-stage "DGX_B200-8_GPUs-PyTorch-1, DGX_B200-8_GPUs-PyTorch-3, GB200-8_GPUs-2_Nodes-PyTorch-1, GB200-8_GPUs-2_Nodes-PyTorch-2, DGX_B300-4_GPUs-PyTorch-Post-Merge-1, GB300-4_GPUs-PyTorch-Post-Merge-2" |
|
PR_Github #54185 [ run ] triggered by Bot. Commit: |
|
PR_Github #54185 [ run ] completed with state |
…pace in KV cache estimation The fp8 context-MLA K/V dequant workspace scales with the summed attended KV length (total_kv_len) of the context requests in a forward step. The KV-cache memory estimator profiles with fresh-prefill dummy requests against an empty cache, so total_kv_len there is pinned near max_num_tokens and this workspace sits at its floor. With block reuse at serving time total_kv_len decouples from max_num_tokens and the workspace can grow far past the profiled floor, but the estimator has already handed that headroom to the KV pool, causing an OOM mid-forward (TestKimiK2::test_nvfp4[4gpus], exposed once NVIDIA#14852 sized the workspace correctly by total_kv_len). Fix the estimation rather than the memory fraction (a user co-tenancy knob): - Split the KV budget between the pool (k bytes/token, all layers) and the workspace (w bytes/token, one shared layer buffer) at a common token count, so max_tokens = budget / (k + w) and the reserved workspace covers exactly max_tokens tokens of attended KV. - Enforce it at admission: the scheduler trims scheduled context requests so their summed attended KV length stays within the pool token capacity, always keeping at least one request as a forward-progress guard. No-op for non-fp8-MLA models. The per-token workspace cost is a single source of truth in C++ (AttentionOp::contextMlaWorkspaceBytesPerToken), exposed via nanobind, so the estimator's reserve can never drift from the runtime allocation. Remove the waive for TestKimiK2::test_nvfp4[4gpus] (nvbugs/6368562) to re-enable the test. Co-Authored-By: Yueh-Ting Chen <yueh.ting.chen@gmail.com> Signed-off-by: Yueh-Ting Chen <yuehtingc@nvidia.com>
…pace in KV cache estimation TestKimiK2::test_nvfp4[4gpus] (NVFP4, TP4 + attention-DP, block reuse) OOMs mid-forward in the MLA context attention. The fp8 context-MLA K/V dequant workspace is one buffer shared across attention layers whose size scales with the summed attended KV length (total_kv_len) of the step's context requests. The KV-cache estimator profiles fresh-prefill dummies against an empty cache, so total_kv_len there sits near max_num_tokens and the workspace is at its floor; with block reuse at serving time total_kv_len decouples from max_num_tokens and the workspace grows past the floor, but the estimator has already handed that headroom to the KV pool. The under-reservation was latent until NVIDIA#14852 sized the workspace by total_kv_len. Reserve for the workspace during estimation instead of lowering free_gpu_memory_fraction (a user co-tenancy knob): - Reserve w * L_cap bytes, where w is the per-token workspace cost and L_cap is the never-stall worst-case summed attended KV per step, min(max_batch_size, max_num_tokens) * max_seq_len. Clamp the reserve to the per-token split budget * w / (k + w) so a memory-constrained node shares the budget at a common token count instead of starving the pool; equivalently the pool keeps max((budget - w*L_cap)/k, budget/(k + w)) tokens. - The estimator carries the exact cap it reserved for (min(L_cap, budget/(k+w))) onto the KV manager; the scheduler reads it directly and trims context requests whose summed attended total_kv_len would exceed it, always keeping one request as a forward-progress guard. It does not re-derive the cap from pool layout, which KV-cache-manager V2 overstates (blocks_in_primary_pool forwards get_page_index_upper_bound, not the available-page count). No-op for non-fp8-MLA models. - w counts the fp8 K/V dequant staging buffer. A sparse-MLA model normally stages nothing, but with the short-seq MHA fallback enabled (TRTLLM_MLA_SHORT_SEQ_MHA_THRESHOLD > 0) it routes short sequences through the dense context path that does stage the buffer, so reserve for that case too. - KvCacheConfig.fp8_context_mla_kv_len_cap (prototype) overrides L_cap to trade reserved workspace for KV pool; the scheduler enforces it. w is a single source of truth in C++ (AttentionOp::contextMlaWorkspaceBytesPerToken, guarded against getWorkspaceSizeForContext by a TLLM_CHECK) exposed via nanobind, so the reserve cannot drift from the runtime allocation. Re-enable TestKimiK2::test_nvfp4[4gpus] (remove the nvbugs/6368562 waive). Co-Authored-By: Yueh-Ting Chen <yueh.ting.chen@gmail.com> Signed-off-by: Yueh-Ting Chen <yuehtingc@nvidia.com>
…pace in KV cache estimation TestKimiK2::test_nvfp4[4gpus] (NVFP4, TP4 + attention-DP, block reuse) OOMs mid-forward in the MLA context attention. The fp8 context-MLA K/V dequant workspace is one buffer shared across attention layers whose size scales with the summed attended KV length (total_kv_len) of the step's context requests. The KV-cache estimator profiles fresh-prefill dummies against an empty cache, so total_kv_len there sits near max_num_tokens and the workspace is at its floor; with block reuse at serving time total_kv_len decouples from max_num_tokens and the workspace grows past the floor, but the estimator has already handed that headroom to the KV pool. The under-reservation was latent until NVIDIA#14852 sized the workspace by total_kv_len. Reserve for the workspace during estimation instead of lowering free_gpu_memory_fraction (a user co-tenancy knob): - Reserve w * L_cap bytes, where w is the per-token workspace cost and L_cap is the never-stall worst-case summed attended KV per step, min(max_batch_size, max_num_tokens) * max_seq_len. Clamp the reserve to the per-token split budget * w / (k + w) so a memory-constrained node shares the budget at a common token count instead of starving the pool; equivalently the pool keeps max((budget - w*L_cap)/k, budget/(k + w)) tokens. - The estimator carries the exact cap it reserved for (min(L_cap, budget/(k+w))) onto the KV manager; the scheduler reads it directly and trims context requests whose summed attended total_kv_len would exceed it, always keeping one request as a forward-progress guard. It does not re-derive the cap from pool layout, which KV-cache-manager V2 overstates (blocks_in_primary_pool forwards get_page_index_upper_bound, not the available-page count). No-op for non-fp8-MLA models. - w counts the fp8 K/V dequant staging buffer. A sparse-MLA model normally stages nothing, but with the short-seq MHA fallback enabled (TRTLLM_MLA_SHORT_SEQ_MHA_THRESHOLD > 0) it routes short sequences through the dense context path that does stage the buffer, so reserve for that case too. - KvCacheConfig.fp8_context_mla_kv_len_cap (prototype) overrides L_cap to trade reserved workspace for KV pool; the scheduler enforces it. w is a single source of truth in C++ (AttentionOp::contextMlaWorkspaceBytesPerToken, guarded against getWorkspaceSizeForContext by a TLLM_CHECK) exposed via nanobind, so the reserve cannot drift from the runtime allocation. Re-enable TestKimiK2::test_nvfp4[4gpus] (remove the nvbugs/6368562 waive). Co-Authored-By: Yueh-Ting Chen <yueh.ting.chen@gmail.com> Signed-off-by: Yueh-Ting Chen <yuehtingc@nvidia.com>
…pace in KV cache estimation TestKimiK2::test_nvfp4[4gpus] (NVFP4, TP4 + attention-DP, block reuse) OOMs mid-forward in the MLA context attention. The fp8 context-MLA K/V dequant workspace is one buffer shared across attention layers whose size scales with the summed attended KV length (total_kv_len) of the step's context requests. The KV-cache estimator profiles fresh-prefill dummies against an empty cache, so total_kv_len there sits near max_num_tokens and the workspace is at its floor; with block reuse at serving time total_kv_len decouples from max_num_tokens and the workspace grows past the floor, but the estimator has already handed that headroom to the KV pool. The under-reservation was latent until NVIDIA#14852 sized the workspace by total_kv_len. Reserve for the workspace during estimation instead of lowering free_gpu_memory_fraction (a user co-tenancy knob): - Only reserve when block reuse is enabled and chunked prefill is disabled -- the sole conditions under which total_kv_len can exceed the profiled floor. Without reuse the workspace is bounded by max_num_tokens (already profiled); with chunked prefill each attention launch is independently bounded by its own chunk buffer. Reserving in those configs would double-count and needlessly shrink the KV pool (up to ~37% for Kimi-K2 attention-DP). No-op there and for non-fp8-MLA models. - Reserve w * L_cap bytes, where w is the per-token workspace cost and L_cap is the never-stall worst-case summed attended KV per step, min(max_batch_size, max_num_tokens) * max_seq_len. Clamp the reserve to the per-token split budget * w / (k + w) so a memory-constrained node shares the budget at a common token count instead of starving the pool; equivalently the pool keeps max((budget - w*L_cap)/k, budget/(k + w)) tokens. - The estimator carries the exact cap it reserved for (min(L_cap, budget/(k+w))) onto the KV manager; the scheduler reads it directly and trims context requests whose summed attended total_kv_len would exceed it, always keeping one request as a forward-progress guard. It does not re-derive the cap from pool layout, which KV-cache-manager V2 overstates (blocks_in_primary_pool forwards get_page_index_upper_bound, not the available-page count). A carried cap of None (no reservation) applies no admission cap. - w counts the fp8 K/V dequant staging buffer. A sparse-MLA model normally stages nothing, but with the short-seq MHA fallback enabled (TRTLLM_MLA_SHORT_SEQ_MHA_THRESHOLD > 0) it routes short sequences through the dense context path that does stage the buffer, so reserve for that case too. - KvCacheConfig.fp8_context_mla_kv_len_cap (prototype) overrides L_cap to trade reserved workspace for KV pool; the scheduler enforces it. w is a single source of truth in C++ (AttentionOp::contextMlaWorkspaceBytesPerToken, guarded against getWorkspaceSizeForContext by a TLLM_CHECK) exposed via nanobind, so the reserve cannot drift from the runtime allocation. This accounts for the fp8 staging term only; the separate BF16 full-gather buffers on the reuse path are a follow-up, so the reserve bounds but does not by itself eliminate reuse-driven OOM. Re-enable TestKimiK2::test_nvfp4[4gpus] (remove the nvbugs/6368562 waive). Co-Authored-By: Yueh-Ting Chen <yueh.ting.chen@gmail.com> Signed-off-by: Yueh-Ting Chen <yuehtingc@nvidia.com>
…pace in KV cache estimation TestKimiK2::test_nvfp4[4gpus] (NVFP4, TP4 + attention-DP, block reuse) OOMs mid-forward in the MLA context attention. The fp8 context-MLA K/V dequant workspace is one buffer shared across attention layers whose size scales with the summed attended KV length (total_kv_len) of the step's context requests. The KV-cache estimator profiles fresh-prefill dummies against an empty cache, so total_kv_len there sits near max_num_tokens and the workspace is at its floor; with block reuse at serving time total_kv_len decouples from max_num_tokens and the workspace grows past the floor, but the estimator has already handed that headroom to the KV pool. The under-reservation was latent until NVIDIA#14852 sized the workspace by total_kv_len. Reserve for the workspace during estimation instead of lowering free_gpu_memory_fraction (a user co-tenancy knob): - Only reserve when block reuse is enabled and chunked prefill is disabled -- the sole conditions under which total_kv_len can exceed the profiled floor. Without reuse the workspace is bounded by max_num_tokens (already profiled); with chunked prefill each attention launch is independently bounded by its own chunk buffer. Reserving in those configs would double-count and needlessly shrink the KV pool (up to ~37% for Kimi-K2 attention-DP). No-op there and for non-fp8-MLA models. - Reserve w * L_cap bytes, where w is the per-token workspace cost and L_cap is the never-stall worst-case summed attended KV per step, min(max_batch_size, max_num_tokens) * max_seq_len. Clamp the reserve to the per-token split budget * w / (k + w) so a memory-constrained node shares the budget at a common token count instead of starving the pool; equivalently the pool keeps max((budget - w*L_cap)/k, budget/(k + w)) tokens. - The estimator carries the exact cap it reserved for (min(L_cap, budget/(k+w))) onto the KV manager; the scheduler reads it directly and trims context requests whose summed attended total_kv_len would exceed it, always keeping one request as a forward-progress guard. It does not re-derive the cap from pool layout, which KV-cache-manager V2 overstates (blocks_in_primary_pool forwards get_page_index_upper_bound, not the available-page count). A carried cap of None (no reservation) applies no admission cap. - w counts the fp8 K/V dequant staging buffer. A sparse-MLA model normally stages nothing, but with the short-seq MHA fallback enabled (TRTLLM_MLA_SHORT_SEQ_MHA_THRESHOLD > 0) it routes short sequences through the dense context path that does stage the buffer, so reserve for that case too. - KvCacheConfig.fp8_context_mla_kv_len_cap (prototype) overrides L_cap to trade reserved workspace for KV pool; the scheduler enforces it. w is a single source of truth in C++ (AttentionOp::contextMlaWorkspaceBytesPerToken, guarded against getWorkspaceSizeForContext by a TLLM_CHECK) exposed via nanobind, so the reserve cannot drift from the runtime allocation. This accounts for the fp8 staging term only; the separate BF16 full-gather buffers on the reuse path are a follow-up, so the reserve bounds but does not by itself eliminate reuse-driven OOM. Re-enable TestKimiK2::test_nvfp4[4gpus] (remove the nvbugs/6368562 waive). Co-Authored-By: Yueh-Ting Chen <yueh.ting.chen@gmail.com> Signed-off-by: Yueh-Ting Chen <yuehtingc@nvidia.com>
…pace in KV cache estimation TestKimiK2::test_nvfp4[4gpus] (NVFP4, TP4 + attention-DP, block reuse) OOMs mid-forward in the MLA context attention. The fp8 context-MLA K/V dequant workspace is one buffer shared across attention layers whose size scales with the summed attended KV length (total_kv_len) of the step's context requests. The KV-cache estimator profiles fresh-prefill dummies against an empty cache, so total_kv_len there sits near max_num_tokens and the workspace is at its floor; with block reuse at serving time total_kv_len decouples from max_num_tokens and the workspace grows past the floor, but the estimator has already handed that headroom to the KV pool. The under-reservation was latent until NVIDIA#14852 sized the workspace by total_kv_len. Reserve for the workspace during estimation instead of lowering free_gpu_memory_fraction (a user co-tenancy knob): - Only reserve when block reuse is enabled and chunked prefill is disabled -- the sole conditions under which total_kv_len can exceed the profiled floor. Without reuse the workspace is bounded by max_num_tokens (already profiled); with chunked prefill each attention launch is independently bounded by its own chunk buffer. Reserving in those configs would double-count and needlessly shrink the KV pool (up to ~37% for Kimi-K2 attention-DP). No-op there and for non-fp8-MLA models. - Reserve w * L_cap bytes, where w is the per-token workspace cost and L_cap is the never-stall worst-case summed attended KV per step, min(max_batch_size, max_num_tokens) * max_seq_len. Clamp the reserve to the per-token split budget * w / (k + w) so a memory-constrained node shares the budget at a common token count instead of starving the pool; equivalently the pool keeps max((budget - w*L_cap)/k, budget/(k + w)) tokens. - The estimator carries the exact cap it reserved for (min(L_cap, budget/(k+w))) onto the KV manager; the scheduler reads it directly and trims context requests whose summed attended total_kv_len would exceed it, always keeping one request as a forward-progress guard. It does not re-derive the cap from pool layout, which KV-cache-manager V2 overstates (blocks_in_primary_pool forwards get_page_index_upper_bound, not the available-page count). A carried cap of None (no reservation) applies no admission cap. - w counts the fp8 K/V dequant staging buffer. A sparse-MLA model normally stages nothing, but with the short-seq MHA fallback enabled (TRTLLM_MLA_SHORT_SEQ_MHA_THRESHOLD > 0) it routes short sequences through the dense context path that does stage the buffer, so reserve for that case too. - KvCacheConfig.fp8_context_mla_kv_len_cap (prototype) overrides L_cap to trade reserved workspace for KV pool; the scheduler enforces it. w is a single source of truth in C++ (AttentionOp::contextMlaWorkspaceBytesPerToken, guarded against getWorkspaceSizeForContext by a TLLM_CHECK) exposed via nanobind, so the reserve cannot drift from the runtime allocation. This accounts for the fp8 staging term only; the separate BF16 full-gather buffers on the reuse path are a follow-up, so the reserve bounds but does not by itself eliminate reuse-driven OOM. Re-enable TestKimiK2::test_nvfp4[4gpus] (remove the nvbugs/6368562 waive). Co-Authored-By: Yueh-Ting Chen <yueh.ting.chen@gmail.com> Signed-off-by: Yueh-Ting Chen <yuehtingc@nvidia.com>
…pace in KV cache estimation TestKimiK2::test_nvfp4[4gpus] (NVFP4, TP4 + attention-DP, block reuse) OOMs mid-forward in the MLA context attention. The fp8 context-MLA K/V dequant workspace is one buffer shared across attention layers whose size scales with the summed attended KV length (total_kv_len) of the step's context requests. The KV-cache estimator profiles fresh-prefill dummies against an empty cache, so total_kv_len there sits near max_num_tokens and the workspace is at its floor; with block reuse at serving time total_kv_len decouples from max_num_tokens and the workspace grows past the floor, but the estimator has already handed that headroom to the KV pool. The under-reservation was latent until NVIDIA#14852 sized the workspace by total_kv_len. Reserve for the workspace during estimation instead of lowering free_gpu_memory_fraction (a user co-tenancy knob): - Only reserve when block reuse is enabled and chunked prefill is disabled -- the sole conditions under which total_kv_len can exceed the profiled floor. Without reuse the workspace is bounded by max_num_tokens (already profiled); with chunked prefill each attention launch is independently bounded by its own chunk buffer. Reserving in those configs would double-count and needlessly shrink the KV pool (up to ~37% for Kimi-K2 attention-DP). No-op there and for non-fp8-MLA models. - Reserve w * L_cap bytes, where w is the per-token workspace cost and L_cap is the never-stall worst-case summed attended KV per step, min(max_batch_size, max_num_tokens) * max_seq_len. Clamp the reserve to the per-token split budget * w / (k + w) so a memory-constrained node shares the budget at a common token count instead of starving the pool; equivalently the pool keeps max((budget - w*L_cap)/k, budget/(k + w)) tokens. - The estimator carries the exact cap it reserved for (min(L_cap, budget/(k+w))) onto the KV manager; the scheduler reads it directly and trims context requests whose summed attended total_kv_len would exceed it, always keeping one request as a forward-progress guard. It does not re-derive the cap from pool layout, which KV-cache-manager V2 overstates (blocks_in_primary_pool forwards get_page_index_upper_bound, not the available-page count). A carried cap of None (no reservation) applies no admission cap. - w counts the fp8 K/V dequant staging buffer. A sparse-MLA model normally stages nothing, but with the short-seq MHA fallback enabled (TRTLLM_MLA_SHORT_SEQ_MHA_THRESHOLD > 0) it routes short sequences through the dense context path that does stage the buffer, so reserve for that case too. - KvCacheConfig.fp8_context_mla_kv_len_cap (prototype) overrides L_cap to trade reserved workspace for KV pool; the scheduler enforces it. w is a single source of truth in C++ (AttentionOp::contextMlaWorkspaceBytesPerToken, guarded against getWorkspaceSizeForContext by a TLLM_CHECK) exposed via nanobind, so the reserve cannot drift from the runtime allocation. This accounts for the fp8 staging term only; the separate BF16 full-gather buffers on the reuse path are a follow-up, so the reserve bounds but does not by itself eliminate reuse-driven OOM. Re-enable TestKimiK2::test_nvfp4[4gpus] (remove the nvbugs/6368562 waive). Co-Authored-By: Yueh-Ting Chen <yueh.ting.chen@gmail.com> Signed-off-by: Yueh-Ting Chen <yuehtingc@nvidia.com>
The attentionOp would use
max_num_tokento prepare for fp8 mla kvcache reuse workspace, but it would requiretotal_kv_lenwhen use, thus causing a potential OOB access to workspace. This has affected multiple CI tests.In terms of memory overhead, under extreme border condition (max possible number of sentence and max possible kvcache reuse), the workspace requires
prefill_batch_size×(max_num_token - 1)×hidden_sizeworth of space, whereprefill_batch_size=min(num_tokens,max_batch_size).For the case of bs=1024 and max_num_token=8192, the DS model would require about 9 GB of space. However, in practice, when KV cache reuse is enabled,
prefill_batch_sizeis usually in the range of 10 to 100. Under those conditions, the required space works out to only a few hundred MB, which is entirely acceptable.Summary by CodeRabbit
Description
Test Coverage
PR Checklist
Please review the following before submitting your PR:
PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.
PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.
Test cases are provided for new code paths (see test instructions)
If PR introduces API changes, an appropriate PR label is added - either
api-compatibleorapi-breaking. Forapi-breaking, includeBREAKINGin the PR title.Any new dependencies have been scanned for license and vulnerabilities
CODEOWNERS updated if ownership changes
Documentation updated as needed
Update tava architecture diagram if there is a significant design change in PR.
The reviewers assigned automatically/manually are appropriate for the PR.
Please check this after reviewing the above items as appropriate for this PR.
GitHub Bot Help
To see a list of available CI bot commands, please comment
/bot help.