Skip to content

[NVBUG-6280721][fix] Unwaive DeepSeek-V3.2 FP4 MTP perf-sanity test - #15783

Merged
hyukn merged 1 commit into
NVIDIA:mainfrom
xguannv:memorynew
Jul 3, 2026
Merged

[NVBUG-6280721][fix] Unwaive DeepSeek-V3.2 FP4 MTP perf-sanity test#15783
hyukn merged 1 commit into
NVIDIA:mainfrom
xguannv:memorynew

Conversation

@xguannv

@xguannv xguannv commented Jun 30, 2026

Copy link
Copy Markdown
Contributor

…s fixed by #14891

These 5 perf_sanity test_e2e cases were waived for nvbugs/6280721 (post-merge 2765/2769 on 06/08-06/09), a flaky CUDA illegal memory access hit during the generation CUDA-graph capture of DeepSeek-V3.2-Exp FP4 with MTP draft.

Root cause: the DSA indexer DSL atom-split path was taken for MTP-Eagle draft (i>=1) calls where next_n collapses to 1, while the cached (dsl_expand_factor, dsl_atom) and the pre-built expanded buffers still describe the verify-time next_n. The mismatch drives the paged-MQA-logits kernel with inflated kv-lens / block-table, producing out-of-range top-k indices that convertReqIndexToGlobal (no upper-bound check) maps to an unmapped KV-pool row. PR #14891 (5741389) guards the split on
next_n == dsl_expand_factor * dsl_atom, landing 06/10 -- after the waive was added, so the waive was never re-evaluated.

Local validation on clean main (no other change): B200 aggr configs v32_fp4_dep8_mtp1_8k1k and v32_fp4_tep8_mtp3_8k1k pass repeatedly (mtp3 4/4 in CI-faithful pytest runs). The remaining GB200 / disagg configs share the same config-agnostic root cause and fix; relying on CI to validate them.

Summary by CodeRabbit

  • Tests
    • Adjusted waived/skipped performance sanity test cases for certain DeepSeek v32 fp4 Blackwell scenarios.
    • Removed several no-longer-waived Blackwell variants and kept a smaller set of remaining waivers.
    • Updated waiver coverage for the upload-generation scenario by dropping one variant while preserving the others.

Description

These waived perf sanity tests are related to flaky illegal memory access during CUDA-graph generation. The root cause in one line: when drafter is decoding, it shares the same config with verifier and thus adopting the wrong path where the verifier should go. This path assumes a longer sequence length than the true sequence length and led to accessing not existing kv cache ids. This problem is fixed by NVBUG-6241842's fix.

Test Coverage

PR Checklist

Please review the following before submitting your PR:

  • PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.

  • PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.

  • Test cases are provided for new code paths (see test instructions)

  • If PR introduces API changes, an appropriate PR label is added - either api-compatible or api-breaking. For api-breaking, include BREAKING in the PR title.

  • Any new dependencies have been scanned for license and vulnerabilities

  • CODEOWNERS updated if ownership changes

  • Documentation updated as needed

  • Update tava architecture diagram if there is a significant design change in PR.

  • The reviewers assigned automatically/manually are appropriate for the PR.

  • Please check this after reviewing the above items as appropriate for this PR.

GitHub Bot Help

To see a list of available CI bot commands, please comment /bot help.

@coderabbitai

coderabbitai Bot commented Jun 30, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

📝 Walkthrough

Walkthrough

Removes 5 waiver entries from tests/integration/test_lists/waives.txt for perf/test_perf_sanity.py::test_e2e DeepSeek v32 fp4 on Blackwell targets: drops aggr_upload-deepseek_v32_fp4_blackwell entries, the grace_blackwell...tep4_mtp3_8k1k waiver, and the disagg_upload-gen_only con1_ctx1_dep4_gen1_tep8 variant.

Changes

Perf Sanity Waiver Removal

Layer / File(s) Summary
Remove DeepSeek v32 fp4 Blackwell waivers
tests/integration/test_lists/waives.txt
Removes aggr_upload-deepseek_v32_fp4_blackwell entries and the grace_blackwell...tep4_mtp3_8k1k waiver; removes the con1_ctx1_dep4_gen1_tep8 case from disagg_upload-gen_only waivers.

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~2 minutes

Possibly related PRs

  • NVIDIA/TensorRT-LLM#15764: Both PRs modify tests/integration/test_lists/waives.txt to adjust DeepSeek v32 fp4 perf sanity waiver entries for Blackwell targets.

Suggested reviewers

  • jieli-matrix
  • yingguo-trt
  • StanleySun639
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title matches the change: it names the NVBug, fix type, and unwaiving the DeepSeek-V3.2 FP4 MTP perf-sanity test.
Description check ✅ Passed The PR description explains the issue, root cause, and fix, and the checklist is filled; only Test Coverage is left blank.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@xguannv

xguannv commented Jun 30, 2026

Copy link
Copy Markdown
Contributor Author

/bot run --disable-fail-fast

@xguannv xguannv changed the title [NVBUG-6280721][infra] Unwaive DeepSeek-V3.2 FP4 MTP perf-sanity test [NVBUG-6280721][fix] Unwaive DeepSeek-V3.2 FP4 MTP perf-sanity test Jul 1, 2026
@xguannv

xguannv commented Jul 1, 2026

Copy link
Copy Markdown
Contributor Author

/bot --help

@github-actions

github-actions Bot commented Jul 1, 2026

Copy link
Copy Markdown

GitHub Bot Help

/bot [-h] ['run', 'kill', 'skip', 'reuse-pipeline'] ...

Provide a user friendly way for developers to interact with a Jenkins server.

Run /bot [-h|--help] to print this help message.

See details below for each supported subcommand.

Details

run [--reuse-test (optional)pipeline-id --disable-fail-fast --skip-test --stage-list "A10-PyTorch-1, xxx" --gpu-type "A30, H100_PCIe" --test-backend "pytorch, cpp" --add-multi-gpu-test --only-multi-gpu-test --disable-multi-gpu-test --post-merge --extra-stage "H100_PCIe-TensorRT-Post-Merge-1, xxx" --detailed-log --debug(experimental) --high-priority]

Launch build/test pipelines. All previously running jobs will be killed.

--reuse-test (optional)pipeline-id (OPTIONAL) : Allow the new pipeline to reuse build artifacts and skip successful test stages from a specified pipeline or the last pipeline if no pipeline-id is indicated. If the Git commit ID has changed, this option will be always ignored. The DEFAULT behavior of the bot is to reuse build artifacts and successful test results from the last pipeline.

--disable-reuse-test (OPTIONAL) : Explicitly prevent the pipeline from reusing build artifacts and skipping successful test stages from a previous pipeline. Ensure that all builds and tests are run regardless of previous successes.

--disable-fail-fast (OPTIONAL) : Disable fail fast on build/tests/infra failures.

--skip-test (OPTIONAL) : Skip all test stages, but still run build stages, package stages and sanity check stages. Note: Does NOT update GitHub check status.

--stage-list "A10-PyTorch-1, xxx" (OPTIONAL) : Only run the specified test stages. Supports wildcard * for pattern matching (e.g., "*PerfSanity*" matches all stages containing PerfSanity). Examples: "A10-PyTorch-1, xxx", "PerfSanity". Note: Does NOT update GitHub check status.

--gpu-type "A30, H100_PCIe" (OPTIONAL) : Only run the test stages on the specified GPU types. Examples: "A30, H100_PCIe". Note: Does NOT update GitHub check status.

--test-backend "pytorch, cpp" (OPTIONAL) : Skip test stages which don't match the specified backends. Only support [pytorch, cpp, tensorrt, triton]. Examples: "pytorch, cpp" (does not run test stages with tensorrt or triton backend). Note: Does NOT update GitHub pipeline status.

--only-multi-gpu-test (OPTIONAL) : Only run the multi-GPU tests. Note: Does NOT update GitHub check status.

--disable-multi-gpu-test (OPTIONAL) : Disable the multi-GPU tests. Note: Does NOT update GitHub check status.

--add-multi-gpu-test (OPTIONAL) : Force run the multi-GPU tests in addition to running L0 pre-merge pipeline.

--post-merge (OPTIONAL) : Run the L0 post-merge pipeline instead of the ordinary L0 pre-merge pipeline.

--extra-stage "H100_PCIe-TensorRT-Post-Merge-1, xxx" (OPTIONAL) : Run the ordinary L0 pre-merge pipeline and specified test stages. Supports wildcard * for pattern matching. Examples: --extra-stage "H100_PCIe-TensorRT-Post-Merge-1, xxx", --extra-stage "Post-Merge".

--detailed-log (OPTIONAL) : Enable flushing out all logs to the Jenkins console. This will significantly increase the log volume and may slow down the job.

--debug (OPTIONAL) : Experimental feature. Enable access to the CI container for debugging purpose. Note: Specify exactly one stage in the stage-list parameter to access the appropriate container environment. Note: Does NOT update GitHub check status.

--high-priority (OPTIONAL) : Run the pipeline with high priority. This option is restricted to authorized users only and will route the job to a high-priority queue.

kill

kill

Kill all running builds associated with pull request.

skip

skip --comment COMMENT

Skip testing for latest commit on pull request. --comment "Reason for skipping build/test" is required. IMPORTANT NOTE: This is dangerous since lack of user care and validation can cause top of tree to break.

reuse-pipeline

reuse-pipeline

Reuse a previous pipeline to validate current commit. This action will also kill all currently running builds associated with the pull request. IMPORTANT NOTE: This is dangerous since lack of user care and validation can cause top of tree to break.

@hyukn

hyukn commented Jul 1, 2026

Copy link
Copy Markdown
Collaborator

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #56804 [ run ] triggered by Bot. Commit: a87b4a2 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #56804 [ run ] completed with state SUCCESS. Commit: a87b4a2
/LLM/main/L0_MergeRequest_PR pipeline #45621 completed with status: 'SUCCESS'

CI Report

Link to invocation

@hyukn

hyukn commented Jul 2, 2026

Copy link
Copy Markdown
Collaborator

/bot run --post-merge --disable-fail-fast --stage-list "*PerfSanity*"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #57060 [ run ] triggered by Bot. Commit: 9fd01ed Link to invocation

@xguannv

xguannv commented Jul 2, 2026

Copy link
Copy Markdown
Contributor Author

/bot run --disable-fail-fast

2 similar comments
@xguannv

xguannv commented Jul 2, 2026

Copy link
Copy Markdown
Contributor Author

/bot run --disable-fail-fast

@hyukn

hyukn commented Jul 2, 2026

Copy link
Copy Markdown
Collaborator

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #57123 [ run ] triggered by Bot. Commit: ba04d93 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #57060 [ run ] completed with state ABORTED. Commit: 9fd01ed

Link to invocation

Signed-off-by: Xin Guan <294044352+xguannv@users.noreply.github.com>
@hyukn

hyukn commented Jul 2, 2026

Copy link
Copy Markdown
Collaborator

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #57169 [ run ] triggered by Bot. Commit: 9f895e7 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #57123 [ run ] completed with state ABORTED. Commit: ba04d93

Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #57169 [ run ] completed with state FAILURE. Commit: 9f895e7
/LLM/main/L0_MergeRequest_PR pipeline #45944 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@xguannv

xguannv commented Jul 2, 2026

Copy link
Copy Markdown
Contributor Author

/bot run --disable-fail-fast

1 similar comment
@hyukn

hyukn commented Jul 3, 2026

Copy link
Copy Markdown
Collaborator

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #57315 [ run ] triggered by Bot. Commit: 9f895e7 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #57315 [ run ] completed with state SUCCESS. Commit: 9f895e7
/LLM/main/L0_MergeRequest_PR pipeline #46073 completed with status: 'SUCCESS'

CI Report

Link to invocation

@hyukn
hyukn self-requested a review July 3, 2026 10:12
@hyukn
hyukn merged commit 3b23a9f into NVIDIA:main Jul 3, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants