Skip to content

[None][test] Unwaive some Perf Tests - #14664

Merged
chenfeiz0326 merged 3 commits into
NVIDIA:mainfrom
chenfeiz0326:chenfeiz/unwaive-perf-tests-528
May 30, 2026
Merged

[None][test] Unwaive some Perf Tests#14664
chenfeiz0326 merged 3 commits into
NVIDIA:mainfrom
chenfeiz0326:chenfeiz/unwaive-perf-tests-528

Conversation

@chenfeiz0326

@chenfeiz0326 chenfeiz0326 commented May 28, 2026

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • Tests
    • Adjusted test waiver entries for performance and integration testing.

Review Change Stack

Description

Test Coverage

PR Checklist

Please review the following before submitting your PR:

  • PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.

  • PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.

  • Test cases are provided for new code paths (see test instructions)

  • If PR introduces API changes, an appropriate PR label is added - either api-compatible or api-breaking. For api-breaking, include BREAKING in the PR title.

  • Any new dependencies have been scanned for license and vulnerabilities

  • CODEOWNERS updated if ownership changes

  • Documentation updated as needed

  • Update tava architecture diagram if there is a significant design change in PR.

  • The reviewers assigned automatically/manually are appropriate for the PR.

  • Please check this after reviewing the above items as appropriate for this PR.

GitHub Bot Help

To see a list of available CI bot commands, please comment /bot help.

@coderabbitai

coderabbitai Bot commented May 28, 2026

Copy link
Copy Markdown
Contributor

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 20f12fd7-57b8-4906-8500-faadafca4df4

📥 Commits

Reviewing files that changed from the base of the PR and between 1b8c739 and 82a334c.

📒 Files selected for processing (1)
  • tests/integration/test_lists/waives.txt
💤 Files with no reviewable changes (1)
  • tests/integration/test_lists/waives.txt

📝 Walkthrough

Walkthrough

Updated CI test waiver configuration by removing three existing performance test skip entries for deepseek and k25-thinking scenarios, and adding new skip entries for additional k25-thinking, dynamo, llama3, super_ad, and deepseek configurations across aggregated and disaggregated upload test paths.

Changes

Test Waiver Updates

Layer / File(s) Summary
Performance test case waivers
tests/integration/test_lists/waives.txt
Waiver entries for perf/test_perf_sanity.py::test_e2e[...] tests are updated: removed two aggr_upload deepseek_v32_fp4_blackwell and one disagg_upload gb200_kimi-k25-thinking-fp4 entries; added multiple new aggr_upload entries for k25_thinking_fp4_blackwell, dynamo_k25_thinking_fp4_blackwell, llama3_1_8b_fp8_ad_hopper, and super_ad, plus two new disagg_upload gb200 deepseek-r1-fp4 and deepseek-v32-fp4 entries.

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~5 minutes

Possibly related PRs

  • NVIDIA/TensorRT-LLM#14393: Both PRs update tests/integration/test_lists/waives.txt, including replacing/adding perf/test_perf_sanity.py::test_e2e waiver entries for aggr_upload/disagg_upload DeepSeek/GPU-specific cases.
  • NVIDIA/TensorRT-LLM#14461: Both PRs update tests/integration/test_lists/waives.txt by removing/adjusting perf/test_perf_sanity.py::test_e2e waiver entries for K2.5/K2.5-thinking scenarios in the same area.
  • NVIDIA/TensorRT-LLM#13790: Both PRs modify tests/integration/test_lists/waives.txt to add/remove specific waived CI test cases for performance sanity tests.

Suggested reviewers

  • mzweilz
  • ZhanruiSunCh
  • jieli-matrix
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The PR description is incomplete. While the template structure is present, all key sections (Description, Test Coverage) are empty with only comments, and the checklist is not properly filled out. Add a concrete description of which tests are being unwaived and why, document the test coverage, and provide clear rationale for the changes in the PR description body.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly indicates the main change: removing waiver entries from a performance test configuration file, making certain tests 'unwaived' to run again.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands and usage tips.

@chenfeiz0326
chenfeiz0326 enabled auto-merge (squash) May 28, 2026 04:26
@chenfeiz0326
chenfeiz0326 requested a review from EmmaQiaoCh May 28, 2026 04:34
@chenfeiz0326

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB200-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-2,DGX_B200-8_GPUs-PyTorch-PerfSanity-Post-Merge-1,DGX_B200-8_GPUs-PyTorch-PerfSanity-Post-Merge-4"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #50728 [ run ] triggered by Bot. Commit: 82a334c Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #50728 [ run ] completed with state SUCCESS. Commit: 82a334c
/LLM/main/L0_MergeRequest_PR pipeline #40210 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@chenfeiz0326

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-3,GB200-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-4,GB200-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-2,DGX_B200-8_GPUs-PyTorch-PerfSanity-Post-Merge-1,DGX_B200-8_GPUs-PyTorch-PerfSanity-Post-Merge-4"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #50957 [ run ] triggered by Bot. Commit: 0a0af05 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #50957 [ run ] completed with state SUCCESS. Commit: 0a0af05
/LLM/main/L0_MergeRequest_PR pipeline #40414 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@chenfeiz0326

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-3"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #51031 [ run ] triggered by Bot. Commit: a201fb9 Link to invocation

@chenfeiz0326
chenfeiz0326 force-pushed the chenfeiz/unwaive-perf-tests-528 branch from a201fb9 to c016008 Compare May 29, 2026 10:03
@chenfeiz0326

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-3"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #51035 [ run ] triggered by Bot. Commit: 4a12b77 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #51035 [ run ] completed with state SUCCESS. Commit: 4a12b77
/LLM/main/L0_MergeRequest_PR pipeline #40482 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@chenfeiz0326

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-3"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #51053 [ run ] triggered by Bot. Commit: 4a12b77 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #51053 [ run ] completed with state SUCCESS. Commit: 4a12b77
/LLM/main/L0_MergeRequest_PR pipeline #40500 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

Signed-off-by: Chenfei Zhang <chenfeiz@nvidia.com>
Signed-off-by: Chenfei Zhang <chenfeiz@nvidia.com>
@chenfeiz0326
chenfeiz0326 force-pushed the chenfeiz/unwaive-perf-tests-528 branch from 4a12b77 to 8b8e0a0 Compare May 30, 2026 01:24
@chenfeiz0326

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-3"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #51142 [ run ] triggered by Bot. Commit: 8b8e0a0 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #51142 [ run ] completed with state SUCCESS. Commit: 8b8e0a0
/LLM/main/L0_MergeRequest_PR pipeline #40578 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

Signed-off-by: Chenfei Zhang <chenfeiz@nvidia.com>
@chenfeiz0326

Copy link
Copy Markdown
Collaborator Author

/bot skip --comment "Only unwaive perf tests, no need to run the whole CI pipeline"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #51170 [ skip ] triggered by Bot. Commit: e34ad28 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #51170 [ skip ] completed with state SUCCESS. Commit: e34ad28
Skipping testing for commit e34ad28

Link to invocation

@chenfeiz0326
chenfeiz0326 merged commit 32eda52 into NVIDIA:main May 30, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants