Skip to content

[None][infra] Waive 21 failed cases for main in post-merge 2780 - #15373

Merged
mzweilz merged 2 commits into
NVIDIA:mainfrom
ZhanruiSunCh:trtllm-ci-report/waive-20260615-095042
Jun 16, 2026
Merged

[None][infra] Waive 21 failed cases for main in post-merge 2780#15373
mzweilz merged 2 commits into
NVIDIA:mainfrom
ZhanruiSunCh:trtllm-ci-report/waive-20260615-095042

Conversation

@ZhanruiSunCh

@ZhanruiSunCh ZhanruiSunCh commented Jun 15, 2026

Copy link
Copy Markdown
Collaborator

Auto-generated Waive PR

Created by: TensorRT LLM CI Report (requested by @mzweilz)
Target branch: main
Bug(s): 6221055, 6323074, 6323889, 6324123, 6324131

Waive entries added

perf/test_perf_sanity.py::test_e2e[aggr_upload-deepseek_v32_fp4_grace_blackwell-v32_fp4_dep4_mtp1_8k1k] SKIP (https://nvbugs/6323889)
perf/test_perf_sanity.py::test_e2e[aggr_upload-deepseek_v32_fp4_grace_blackwell-v32_fp4_tep4_mtp3_1k1k] SKIP (https://nvbugs/6323889)
perf/test_perf_sanity.py::test_e2e[aggr_upload-glm5_fp4_2_nodes_grace_blackwell-glm5_fp4_dep8_mtp1_8k1k] SKIP (https://nvbugs/6324131)
perf/test_perf_sanity.py::test_e2e[aggr_upload-glm5_fp4_2_nodes_grace_blackwell-glm5_fp4_tep8_mtp3_8k1k] SKIP (https://nvbugs/6324131)
perf/test_perf_sanity.py::test_e2e[disagg_upload-e2e-gb300_deepseek-r1-fp4_128k8k_con256_ctx1_pp4_gen1_dep8_eplb0_mtp1_ccb-NIXL] SKIP (https://nvbugs/6323889)
perf/test_perf_sanity.py::test_e2e[disagg_upload-e2e-gb300_glm-5-fp4_1k1k_con4096_ctx1_dep2_gen1_dep8_eplb256_mtp1_ccb-NIXL] SKIP (https://nvbugs/6324131)
perf/test_perf_sanity.py::test_e2e[disagg_upload-e2e-gb300_glm-5-fp4_8k1k_con1024_ctx1_dep2_gen1_dep8_eplb256_mtp1_ccb-NIXL] SKIP (https://nvbugs/6324131)
perf/test_perf_sanity.py::test_e2e[disagg_upload-gen_only-gb200_deepseek-r1-fp4_1k1k_con1024_ctx1_dep4_gen1_dep8_eplb0_mtp0_ccb-NIXL] SKIP (https://nvbugs/6323889)
perf/test_perf_sanity.py::test_e2e[disagg_upload-gen_only-gb200_deepseek-r1-fp4_1k1k_con2048_ctx2_dep4_gen1_dep16_eplb288_mtp3_ccb-NIXL] SKIP (https://nvbugs/6323889)
perf/test_perf_sanity.py::test_e2e[disagg_upload-gen_only-gb200_deepseek-r1-fp4_1k1k_con3072_ctx1_dep4_gen1_dep4_eplb0_mtp1_ccb-NIXL] SKIP (https://nvbugs/6323889)
perf/test_perf_sanity.py::test_e2e[disagg_upload-gen_only-gb200_gpt-oss-120b-fp4_1k1k_con2048_ctx1_tp1_gen1_dep2_eplb0_mtp0_ccb-NIXL] SKIP (https://nvbugs/6324123)
perf/test_perf_sanity.py::test_e2e[disagg_upload-gen_only-gb200_kimi-k25-thinking-fp4_1k1k_con4096_ctx1_dep4_gen1_dep8_eplb0_mtp0_ccb-NIXL] SKIP (https://nvbugs/6323074)
perf/test_perf_sanity.py::test_e2e[disagg_upload-gen_only-gb200_kimi-k25-thinking-fp4_8k1k_con1024_ctx1_dep4_gen1_dep32_eplb416_mtp3_ccb-NIXL] SKIP (https://nvbugs/6323074)
perf/test_perf_sanity.py::test_e2e[disagg_upload-gen_only-gb300_deepseek-r1-fp4_1k1k_con3072_ctx1_dep4_gen1_dep4_eplb0_mtp1_ccb-NIXL] SKIP (https://nvbugs/6323889)
perf/test_perf_sanity.py::test_e2e[disagg_upload-gen_only-gb300_glm-5-fp4_1k1k_con4096_ctx1_dep2_gen1_dep8_eplb256_mtp1_ccb-NIXL] SKIP (https://nvbugs/6324131)
perf/test_perf_sanity.py::test_e2e[disagg_upload-gen_only-gb300_glm-5-fp4_8k1k_con1024_ctx1_dep2_gen1_dep8_eplb256_mtp1_ccb-NIXL] SKIP (https://nvbugs/6324131)
perf/test_perf_sanity.py::test_e2e[disagg_upload-gen_only-gb300_glm-5-fp4_8k1k_con1_ctx1_dep2_gen1_tep8_eplb0_mtp3_ccb-NIXL] SKIP (https://nvbugs/6324131)
perf/test_perf_sanity.py::test_e2e[disagg_upload-gen_only-gb300_kimi-k25-thinking-fp4_1k1k_con4096_ctx1_dep4_gen1_dep8_eplb0_mtp0_ccb-NIXL] SKIP (https://nvbugs/6323074)
perf/test_perf_sanity.py::test_e2e[disagg_upload-gen_only-gb300_kimi-k25-thinking-fp4_1k1k_con4_ctx1_dep4_gen1_tep4_eplb0_mtp0_ccb-NIXL] SKIP (https://nvbugs/6323074)
perf/test_perf_sanity.py::test_e2e[disagg_upload-gen_only-gb300_kimi-k25-thinking-fp4_8k1k_con1024_ctx1_dep4_gen1_dep32_eplb416_mtp3_ccb-NIXL] SKIP (https://nvbugs/6323074)
unittest/_torch/thop/parallel SKIP (https://nvbugs/6221055)

This PR was auto-generated by TensorRT LLM CI Report. Please review the waive entries before merging.

Summary by CodeRabbit

  • Tests
    • Updated test waiver configurations and related references.

Note: This release includes internal test infrastructure updates with no user-facing changes.

Bug(s): 6221055, 6323074, 6323889, 6324123, 6324131
Requested by: @mzweilz

Signed-off-by: ZhanruiSunCh <184402041+ZhanruiSunCh@users.noreply.github.com>
@coderabbitai

coderabbitai Bot commented Jun 15, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 25f2fe3f-fbf2-4689-a640-f7aed3931921

📥 Commits

Reviewing files that changed from the base of the PR and between 870f9b5 and c0f59b4.

📒 Files selected for processing (1)
  • tests/integration/test_lists/waives.txt

📝 Walkthrough

Walkthrough

Updates tests/integration/test_lists/waives.txt by replacing existing perf/test_perf_sanity.py::test_e2e SKIP entries with new aggr_upload, disagg_upload, and gen_only waivers for Blackwell v32 FP4, GB200, and GB300 model variants, and adds a new SKIP entry for unittest/_torch/thop/parallel.

Changes

Test Waiver List Updates

Layer / File(s) Summary
Perf sanity and unittest waiver additions
tests/integration/test_lists/waives.txt
Replaces perf/test_perf_sanity.py::test_e2e SKIP entries (lines 278–297, 308–319) with new aggr_upload/disagg_upload/gen_only waivers for Blackwell v32 FP4 grace, GB200, and GB300 variants with updated nvbugs references; adds a new SKIP for unittest/_torch/thop/parallel (nvbugs/6221055).

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~3 minutes

Possibly related PRs

  • NVIDIA/TensorRT-LLM#15253: Modifies the same waives.txt file by adding/updating SKIP entries for perf/e2e tests with NV bug references.
  • NVIDIA/TensorRT-LLM#15250: Directly related — updates the same perf/test_perf_sanity.py::test_e2e SKIP waivers for disagg_upload/gen_only FP4 configurations including DeepSeek/V32.
  • NVIDIA/TensorRT-LLM#15269: Modifies the same waives.txt by updating/removing perf/test_perf_sanity.py::test_e2e waiver entries.

Suggested reviewers

  • mzweilz
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title follows the required format [None][infra] and clearly summarizes the main change: adding waived test cases for the main branch post-merge.
Description check ✅ Passed The description clearly explains the auto-generated waive PR with referenced bugs, listed waive entries, and relevant context. It includes all essential information needed to understand the changes.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands and usage tips.

@mzweilz
mzweilz requested a review from chenfeiz0326 June 15, 2026 10:08
@mzweilz

mzweilz commented Jun 15, 2026

Copy link
Copy Markdown
Collaborator

/bot run --stage-list "A10-Build_Docs"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #54282 [ run ] triggered by Bot. Commit: c0f59b4 Link to invocation

Signed-off-by: Abby Wei <18545893+mzweilz@users.noreply.github.com>
@mzweilz

mzweilz commented Jun 15, 2026

Copy link
Copy Markdown
Collaborator

/bot run --stage-list "A10-Build_Docs"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #54286 [ run ] triggered by Bot. Commit: c2386cf Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #54282 [ run ] completed with state ABORTED. Commit: c0f59b4

Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #54286 [ run ] completed with state SUCCESS. Commit: c2386cf
/LLM/main/L0_MergeRequest_PR pipeline #43357 (Partly Tested) completed with status: 'SUCCESS'

CI Report

Link to invocation

@mzweilz

mzweilz commented Jun 16, 2026

Copy link
Copy Markdown
Collaborator

/bot skip --comment "Tests passed for waive PR created via CI report"

@mzweilz
mzweilz enabled auto-merge (squash) June 16, 2026 02:13
@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #54418 [ skip ] triggered by Bot. Commit: c2386cf Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #54418 [ skip ] completed with state SUCCESS. Commit: c2386cf
Skipping testing for commit c2386cf

Link to invocation

@mzweilz
mzweilz merged commit f49d09f into NVIDIA:main Jun 16, 2026
8 checks passed
chenfeiz0326 added a commit to chenfeiz0326/TensorRT-LLM that referenced this pull request Jun 17, 2026
…num_tokens, perf-sanity timeout

Group 3 (KimiK25ForConditionalGeneration not in moe_model_arch_list):
  Add 'KimiK25ForConditionalGeneration' to moe_model_arch_list so
  maybe_create_moe_load_balancer calls MoeLoadBalancerConfig.setup() for
  Kimi K2.5 MoE+EPLB configs. Without it, the wrapped DeepseekV3 backbone
  raised 'Cannot calculate num_local_slots' on every rank during model
  load (gb200_kimi-k25-thinking-fp4_8k1k_con1024_dep32_eplb416_mtp3
  failures, OCI verify Slurm 3320245).

Group 4 (glm5 8k1k tep8_mtp3 RequestError):
  max_num_tokens=256 < ISL=8192 in glm5_fp4_tep8_mtp3_8k1k config; bump
  to 8192 to match every other 8k1k aggregated config in the repo. OCI
  verify (Slurm 3320259) shows output_token_throughput=238.94 tok/s.

Group 6 (long-context 128k8k > 5400s):
  DEFAULT_TIMEOUT 5400 -> 10800 in test_perf_sanity.py so the in-test
  wait_for_server_config / wait_for_benchmark_ready give the engine
  enough budget for 128k context fills. Re-balance pytest TIMEOUT
  markers across the perf-sanity test-db YAMLs:
  - 128k8k entries: TIMEOUT (120) -> TIMEOUT (180), matches the new 3h
    server budget;
  - all other entries: TIMEOUT (120) -> TIMEOUT (90), tightening the
    outer cap on tests that don't need 2h.
  AWS GB300 verify of e2e-gb300_deepseek-r1-fp4_128k8k_con256
  (Slurm 560881) reached output_token_throughput=820.77 tok/s.

Signed-off-by: Chenfei Zhang <chenfeiz@nvidia.com>
chenfeiz0326 added a commit to chenfeiz0326/TensorRT-LLM that referenced this pull request Jun 17, 2026
…ve fixed cases

local/submit.py was missing the GB300 UCX_TLS path that the CI submit.py
already has. Without it, AWS GB300 disagg tests crashed with
'Failed to create NIXL backend: UCX' at startup. Mirror the CI logic
(jenkins/scripts/perf/submit.py:543-548):
  is_gb300 -> export UCX_TLS=cuda_copy,cuda_ipc,sm,self,tcp
  is_b200  -> export UCX_TLS=^ib  (unchanged)
  default  -> unset UCX_TLS UCX_NET_DEVICES  (was: only UCX_TLS)
AWS GB300 verify of e2e-gb300_deepseek-r1-fp4_128k8k_con256
(Slurm 560881) reached output_token_throughput=820.77 tok/s after the
fix; before it died at 5min with the NIXL backend error.

Drop waivers for the four cases that the fixes in this PR repair:
  - aggr_upload-glm5_fp4_2_nodes_grace_blackwell-glm5_fp4_tep8_mtp3_8k1k
    (was nvbugs/6324131; fixed by max_num_tokens 256->8192)
  - disagg_upload-gen_only-gb200_kimi-k25-thinking-fp4_8k1k_con1024_dep32_eplb416_mtp3
    (was nvbugs/6323074; fixed by adding KimiK25 to moe_model_arch_list)
  - disagg_upload-gen_only-gb300_kimi-k25-thinking-fp4_8k1k_con1024_dep32_eplb416_mtp3
    (was nvbugs/6323074; same fix)
  - disagg_upload-e2e-gb300_deepseek-r1-fp4_128k8k_con256_pp4_dep8_mtp1
    (was nvbugs/6323889; fixed by DEFAULT_TIMEOUT 5400->10800 + UCX_TLS)

Signed-off-by: Chenfei Zhang <chenfeiz@nvidia.com>
xinhe-nv pushed a commit to tensorrt-cicd/TensorRT-LLM that referenced this pull request Jun 23, 2026
…IA#15373)

Signed-off-by: ZhanruiSunCh <184402041+ZhanruiSunCh@users.noreply.github.com>
Signed-off-by: Abby Wei <18545893+mzweilz@users.noreply.github.com>
Co-authored-by: Abby Wei <18545893+mzweilz@users.noreply.github.com>
Signed-off-by: GitLab CI Bot <gitlab-ci@nvidia.com>
chenfeiz0326 added a commit to chenfeiz0326/TensorRT-LLM that referenced this pull request Jun 23, 2026
…num_tokens, perf-sanity timeout

Group 3 (KimiK25ForConditionalGeneration not in moe_model_arch_list):
  Add 'KimiK25ForConditionalGeneration' to moe_model_arch_list so
  maybe_create_moe_load_balancer calls MoeLoadBalancerConfig.setup() for
  Kimi K2.5 MoE+EPLB configs. Without it, the wrapped DeepseekV3 backbone
  raised 'Cannot calculate num_local_slots' on every rank during model
  load (gb200_kimi-k25-thinking-fp4_8k1k_con1024_dep32_eplb416_mtp3
  failures, OCI verify Slurm 3320245).

Group 4 (glm5 8k1k tep8_mtp3 RequestError):
  max_num_tokens=256 < ISL=8192 in glm5_fp4_tep8_mtp3_8k1k config; bump
  to 8192 to match every other 8k1k aggregated config in the repo. OCI
  verify (Slurm 3320259) shows output_token_throughput=238.94 tok/s.

Group 6 (long-context 128k8k > 5400s):
  DEFAULT_TIMEOUT 5400 -> 10800 in test_perf_sanity.py so the in-test
  wait_for_server_config / wait_for_benchmark_ready give the engine
  enough budget for 128k context fills. Re-balance pytest TIMEOUT
  markers across the perf-sanity test-db YAMLs:
  - 128k8k entries: TIMEOUT (120) -> TIMEOUT (180), matches the new 3h
    server budget;
  - all other entries: TIMEOUT (120) -> TIMEOUT (90), tightening the
    outer cap on tests that don't need 2h.
  AWS GB300 verify of e2e-gb300_deepseek-r1-fp4_128k8k_con256
  (Slurm 560881) reached output_token_throughput=820.77 tok/s.

Signed-off-by: Chenfei Zhang <chenfeiz@nvidia.com>
chenfeiz0326 added a commit to chenfeiz0326/TensorRT-LLM that referenced this pull request Jun 23, 2026
…ve fixed cases

local/submit.py was missing the GB300 UCX_TLS path that the CI submit.py
already has. Without it, AWS GB300 disagg tests crashed with
'Failed to create NIXL backend: UCX' at startup. Mirror the CI logic
(jenkins/scripts/perf/submit.py:543-548):
  is_gb300 -> export UCX_TLS=cuda_copy,cuda_ipc,sm,self,tcp
  is_b200  -> export UCX_TLS=^ib  (unchanged)
  default  -> unset UCX_TLS UCX_NET_DEVICES  (was: only UCX_TLS)
AWS GB300 verify of e2e-gb300_deepseek-r1-fp4_128k8k_con256
(Slurm 560881) reached output_token_throughput=820.77 tok/s after the
fix; before it died at 5min with the NIXL backend error.

Drop waivers for the four cases that the fixes in this PR repair:
  - aggr_upload-glm5_fp4_2_nodes_grace_blackwell-glm5_fp4_tep8_mtp3_8k1k
    (was nvbugs/6324131; fixed by max_num_tokens 256->8192)
  - disagg_upload-gen_only-gb200_kimi-k25-thinking-fp4_8k1k_con1024_dep32_eplb416_mtp3
    (was nvbugs/6323074; fixed by adding KimiK25 to moe_model_arch_list)
  - disagg_upload-gen_only-gb300_kimi-k25-thinking-fp4_8k1k_con1024_dep32_eplb416_mtp3
    (was nvbugs/6323074; same fix)
  - disagg_upload-e2e-gb300_deepseek-r1-fp4_128k8k_con256_pp4_dep8_mtp1
    (was nvbugs/6323889; fixed by DEFAULT_TIMEOUT 5400->10800 + UCX_TLS)

Signed-off-by: Chenfei Zhang <chenfeiz@nvidia.com>
chenfeiz0326 added a commit to chenfeiz0326/TensorRT-LLM that referenced this pull request Jun 24, 2026
…num_tokens, perf-sanity timeout

Group 3 (KimiK25ForConditionalGeneration not in moe_model_arch_list):
  Add 'KimiK25ForConditionalGeneration' to moe_model_arch_list so
  maybe_create_moe_load_balancer calls MoeLoadBalancerConfig.setup() for
  Kimi K2.5 MoE+EPLB configs. Without it, the wrapped DeepseekV3 backbone
  raised 'Cannot calculate num_local_slots' on every rank during model
  load (gb200_kimi-k25-thinking-fp4_8k1k_con1024_dep32_eplb416_mtp3
  failures, OCI verify Slurm 3320245).

Group 4 (glm5 8k1k tep8_mtp3 RequestError):
  max_num_tokens=256 < ISL=8192 in glm5_fp4_tep8_mtp3_8k1k config; bump
  to 8192 to match every other 8k1k aggregated config in the repo. OCI
  verify (Slurm 3320259) shows output_token_throughput=238.94 tok/s.

Group 6 (long-context 128k8k > 5400s):
  DEFAULT_TIMEOUT 5400 -> 10800 in test_perf_sanity.py so the in-test
  wait_for_server_config / wait_for_benchmark_ready give the engine
  enough budget for 128k context fills. Re-balance pytest TIMEOUT
  markers across the perf-sanity test-db YAMLs:
  - 128k8k entries: TIMEOUT (120) -> TIMEOUT (180), matches the new 3h
    server budget;
  - all other entries: TIMEOUT (120) -> TIMEOUT (90), tightening the
    outer cap on tests that don't need 2h.
  AWS GB300 verify of e2e-gb300_deepseek-r1-fp4_128k8k_con256
  (Slurm 560881) reached output_token_throughput=820.77 tok/s.

Signed-off-by: Chenfei Zhang <chenfeiz@nvidia.com>
chenfeiz0326 added a commit to chenfeiz0326/TensorRT-LLM that referenced this pull request Jun 24, 2026
…ve fixed cases

local/submit.py was missing the GB300 UCX_TLS path that the CI submit.py
already has. Without it, AWS GB300 disagg tests crashed with
'Failed to create NIXL backend: UCX' at startup. Mirror the CI logic
(jenkins/scripts/perf/submit.py:543-548):
  is_gb300 -> export UCX_TLS=cuda_copy,cuda_ipc,sm,self,tcp
  is_b200  -> export UCX_TLS=^ib  (unchanged)
  default  -> unset UCX_TLS UCX_NET_DEVICES  (was: only UCX_TLS)
AWS GB300 verify of e2e-gb300_deepseek-r1-fp4_128k8k_con256
(Slurm 560881) reached output_token_throughput=820.77 tok/s after the
fix; before it died at 5min with the NIXL backend error.

Drop waivers for the four cases that the fixes in this PR repair:
  - aggr_upload-glm5_fp4_2_nodes_grace_blackwell-glm5_fp4_tep8_mtp3_8k1k
    (was nvbugs/6324131; fixed by max_num_tokens 256->8192)
  - disagg_upload-gen_only-gb200_kimi-k25-thinking-fp4_8k1k_con1024_dep32_eplb416_mtp3
    (was nvbugs/6323074; fixed by adding KimiK25 to moe_model_arch_list)
  - disagg_upload-gen_only-gb300_kimi-k25-thinking-fp4_8k1k_con1024_dep32_eplb416_mtp3
    (was nvbugs/6323074; same fix)
  - disagg_upload-e2e-gb300_deepseek-r1-fp4_128k8k_con256_pp4_dep8_mtp1
    (was nvbugs/6323889; fixed by DEFAULT_TIMEOUT 5400->10800 + UCX_TLS)

Signed-off-by: Chenfei Zhang <chenfeiz@nvidia.com>
xinhe-nv pushed a commit to tensorrt-cicd/TensorRT-LLM that referenced this pull request Jun 24, 2026
…IA#15373)

Signed-off-by: ZhanruiSunCh <184402041+ZhanruiSunCh@users.noreply.github.com>
Signed-off-by: Abby Wei <18545893+mzweilz@users.noreply.github.com>
Co-authored-by: Abby Wei <18545893+mzweilz@users.noreply.github.com>
Signed-off-by: GitLab CI Bot <gitlab-ci@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants