[https://nvbugs/5948435][chore] Unwaive DeepSeekV3Lite test_nvfp4_4gpus CUTLASS ep4 fp8kv on RTXPro6000D - #16621
Conversation
…us CUTLASS ep4 fp8kv on RTXPro6000D
The test was waived as flaky ('Test terminated unexpectedly').
Local repro on RTX Pro 6000 Blackwell (SM120) x4 with a correctly-built
binary (both single-arch '120-real' and CI's multi-arch
'90-real;100-real;103-real;120-real' clean builds) passes deterministically
(GSM8K accuracy ~63.5-64.0, matching reference 63.71). The SM120 FP8 MLA
generation FMHA kernel is present in a correct build (381 sm_120 SASS
kernels). Repair Bot (2026-07-17) and OpenSearch (90+ days) also show no
reproduction. Removing the waive so post-merge CI can confirm.
Signed-off-by: XingFei Xi <xxi@nvidia.com>
|
/bot run --add-multi-gpu-test --stage-list "RTXPro6000D-4_GPUs-PyTorch-Post-Merge-1" |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (1)
💤 Files with no reviewable changes (1)
WalkthroughChangesDeepSeekV3Lite integration waiver
Estimated code review effort: 1 (Trivial) | ~2 minutes Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
|
PR_Github #60360 [ run ] triggered by Bot. Commit: |
|
PR_Github #60360 [ run ] completed with state
|
|
/bot run --extra-stage "RTXPro6000D-4_GPUs-PyTorch-Post-Merge-1" |
|
PR_Github #60502 [ run ] triggered by Bot. Commit: |
|
PR_Github #60502 [ run ] completed with state
|
|
/bot run --extra-stage "RTXPro6000D-4_GPUs-PyTorch-Post-Merge-1" |
|
PR_Github #60601 [ run ] triggered by Bot. Commit: |
|
PR_Github #60601 [ run ] completed with state
|
|
/bot run --extra-stage "RTXPro6000D-4_GPUs-PyTorch-Post-Merge-1" |
|
PR_Github #60676 [ run ] triggered by Bot. Commit: |
|
PR_Github #60676 [ run ] completed with state
|
|
/bot run --extra-stage "RTXPro6000D-4_GPUs-PyTorch-Post-Merge-1" |
|
PR_Github #60869 [ run ] triggered by Bot. Commit: |
|
PR_Github #60869 [ run ] completed with state
|
|
/bot run --disable-fail-fast --extra-stage "RTXPro6000D-4_GPUs-PyTorch-Post-Merge-1" |
|
PR_Github #60914 [ run ] triggered by Bot. Commit: |
|
PR_Github #60914 [ run ] completed with state
|
|
/bot run --stage-list "RTXPro6000D-4_GPUs-PyTorch-Post-Merge-1" |
|
PR_Github #61129 [ run ] triggered by Bot. Commit: |
|
PR_Github #61129 [ run ] completed with state |
|
/bot run --stage-list "RTXPro6000D-4_GPUs-PyTorch-Post-Merge-1, RTXPro6000D-4_GPUs-PyTorch-Post-Merge-2" |
|
/bot run --stage-list "RTXPro6000D-4_GPUs-PyTorch-Post-Merge-2" |
|
PR_Github #61576 [ run ] triggered by Bot. Commit: |
|
PR_Github #61576 [ run ] completed with state |
|
/bot run --stage-list "RTXPro6000D-4_GPUs-PyTorch-Post-Merge-1, RTXPro6000D-4_GPUs-PyTorch-Post-Merge-2" --disable-fail-fast |
|
PR_Github #61664 [ run ] triggered by Bot. Commit: |
|
PR_Github #61664 [ run ] completed with state |
|
/bot run --disable-fail-fast |
|
PR_Github #61668 [ run ] triggered by Bot. Commit: |
|
PR_Github #61668 [ run ] completed with state |
…us CUTLASS ep4 fp8kv on RTXPro6000D (NVIDIA#16621) Signed-off-by: XingFei Xi <xxi@nvidia.com>
…nnahz/dep-1083-port-flashinfer-stable-va-lifecycle-for-native-all-reduce * 'main' of https://github.com/NVIDIA/TensorRT-LLM: (54 commits) [NVIDIA#15673][fix] Enable CUDA core fast path for SM89/SM120/SM121 (NVIDIA#12705) [None][test] Adjust timeout cases in QA perf test (NVIDIA#16894) [https://nvbugs/6157892][fix] Mistral format refactor (NVIDIA#15123) [None][feat] Add kimi_k2/glm_5 grouped routing and fused router to bench_moe (NVIDIA#16830) [https://nvbugs/6501376][fix] Test-only fix — drop the `if hidden_size % 2 != 0: with pytest.raises(...)`… (NVIDIA#16844) [TRTLLM-13642][feat] Add perf sanity tests for Llama-3.1-8B and Gemma-3-1B and verify cache transceiver V2 support (NVIDIA#16355) [https://nvbugs/6433376][fix] Update the Dense test to mirror the MoE sibling — assert `bfloat16` under… (NVIDIA#16203) [None][fix] Resolve NVFP4 mixed-precision base layers for the DSpark draft (NVIDIA#16831) [https://nvbugs/6479324][test] Remove waiver for fixed qwen3_5_4b_fp8_stress disaggregated stress test (NVIDIA#16878) [https://nvbugs/6507109][infra] Split slow DGX B300 attention unit tests (NVIDIA#16838) [None][infra] Waive 21 failed cases for main in post-merge 2862 (NVIDIA#16882) [None][perf] prepare_inputs: avoid O(seq_len) get_tokens(0) marshalling on the host (NVIDIA#16791) [None][perf] Optimize Blackwell fused MHC half-MMA kernel (NVIDIA#16799) [None][infra] Auto-update test durations from OpenSearch (last 7 days) [None][perf] Skip DeepGEMM clean_logits in DSA indexer prefill on custom top-k path (NVIDIA#16789) [None][feat] Support DeepSeek-V4 in layer_wise_benchmarks (NVIDIA#16774) [https://nvbugs/6465993][fix] use attention cache dtype for disaggregated transfer (NVIDIA#16505) [https://nvbugs/6463822][fix] Fix LTX2 CUDA graph test leak issue (NVIDIA#16775) [https://nvbugs/5948435][chore] Unwaive DeepSeekV3Lite test_nvfp4_4gpus CUTLASS ep4 fp8kv on RTXPro6000D (NVIDIA#16621) [TRTLLM-14417][fix] Exclude ADP/cuda-graph dummy requests from speculative-decode acceptance stats (NVIDIA#16571) ... Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
…nnahz/dep-1082-shared-mnnvl-moe-lifecycle * 'main' of https://github.com/NVIDIA/TensorRT-LLM: (142 commits) [NVIDIA#15673][fix] Enable CUDA core fast path for SM89/SM120/SM121 (NVIDIA#12705) [None][test] Adjust timeout cases in QA perf test (NVIDIA#16894) [https://nvbugs/6157892][fix] Mistral format refactor (NVIDIA#15123) [None][feat] Add kimi_k2/glm_5 grouped routing and fused router to bench_moe (NVIDIA#16830) [https://nvbugs/6501376][fix] Test-only fix — drop the `if hidden_size % 2 != 0: with pytest.raises(...)`… (NVIDIA#16844) [TRTLLM-13642][feat] Add perf sanity tests for Llama-3.1-8B and Gemma-3-1B and verify cache transceiver V2 support (NVIDIA#16355) [https://nvbugs/6433376][fix] Update the Dense test to mirror the MoE sibling — assert `bfloat16` under… (NVIDIA#16203) [None][fix] Resolve NVFP4 mixed-precision base layers for the DSpark draft (NVIDIA#16831) [https://nvbugs/6479324][test] Remove waiver for fixed qwen3_5_4b_fp8_stress disaggregated stress test (NVIDIA#16878) [https://nvbugs/6507109][infra] Split slow DGX B300 attention unit tests (NVIDIA#16838) [None][infra] Waive 21 failed cases for main in post-merge 2862 (NVIDIA#16882) [None][perf] prepare_inputs: avoid O(seq_len) get_tokens(0) marshalling on the host (NVIDIA#16791) [None][perf] Optimize Blackwell fused MHC half-MMA kernel (NVIDIA#16799) [None][infra] Auto-update test durations from OpenSearch (last 7 days) [None][perf] Skip DeepGEMM clean_logits in DSA indexer prefill on custom top-k path (NVIDIA#16789) [None][feat] Support DeepSeek-V4 in layer_wise_benchmarks (NVIDIA#16774) [https://nvbugs/6465993][fix] use attention cache dtype for disaggregated transfer (NVIDIA#16505) [https://nvbugs/6463822][fix] Fix LTX2 CUDA graph test leak issue (NVIDIA#16775) [https://nvbugs/5948435][chore] Unwaive DeepSeekV3Lite test_nvfp4_4gpus CUTLASS ep4 fp8kv on RTXPro6000D (NVIDIA#16621) [TRTLLM-14417][fix] Exclude ADP/cuda-graph dummy requests from speculative-decode acceptance stats (NVIDIA#16571) ... Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
Summary
Unwaive
TestDeepSeekV3Lite::test_nvfp4_4gpus[moe_backend=CUTLASS-mtp_nextn=0-ep4-fp8kv=True-attention_dp=True-cuda_graph=True-overlap_scheduler=True-low_precision_combine=False-torch_compile=False]on RTXPro6000D, tracked by NVBug 5948435.Why
The test was waived with the generic symptom "Test terminated unexpectedly" (a crash, not an accuracy failure). Investigation shows it is not reproducible on a correctly-built binary:
-a 120-real→ PASS (GSM8K accuracy 63.99, ref 63.71)-a '90-real;100-real;103-real;120-real'→ PASS (GSM8K accuracy 63.50)sm_120SASS kernels), including the FP8 MLA-generation FMHA kernel — identical to the single-arch build, so multi-arch compilation does not drop SM120 kernels.The one crash I initially observed was traced to an incremental build that reused a stale
cpp/builddirectory pinned to100-real(B200), producing an SM120-incomplete.so— a local build artifact, not the product/CI behavior. CI builds include120-real(jenkins/Build.groovy,jenkins/L0_Test.groovy), so the CI binary is SM120-complete.The original 2026-03-03 failure was therefore most likely a genuine transient/flaky event (or has since been fixed). Removing the waive so post-merge CI can confirm on the real CI binary.
Test Coverage
Removing a waive; the test itself re-enters CI. Post-merge stage:
RTXPro6000D-4_GPUs-PyTorch-Post-Merge-1.PR Checklist
[JIRA/NVBUG][type] descriptionpre-commit runpasses (incl. waives.txt duplication + AST test-list validation)Summary by CodeRabbit
tests/integration/test_lists/waives.txtto remove thenvbugs/5948435SKIP waiver forfull:RTXPro6000D/accuracy/test_llm_api_pytorch.py::TestDeepSeekV3Lite::test_nvfp4_4gpus(CUTLASS, EP4, FP8KV,torch_compile=Falseparameterization).torch_compile=Truewaiver in place (still pointing atnvbugs/5961814), so the configuration can resume running inRTXPro6000D-4_GPUs-PyTorch-Post-Merge-1.Dev Engineer Review
nvbugs/5948435entry is no longer present intests/integration/test_lists/waives.txt.torch_compile=True(nvbugs/5961814), without broadening waiver scope to unrelated parameterizations.QA Engineer Review
tests/integration/test_lists/waives.txtfull:RTXPro6000D/accuracy/test_llm_api_pytorch.py::TestDeepSeekV3Lite::test_nvfp4_4gpus[...]torch_compile=FalseSKIP entry referencingnvbugs/5948435.torch_compile=TrueSKIP entry for the sameRTXPro6000Dtest/config (referencingnvbugs/5961814).