Skip to content

[None][test] Waive 11 failed cases for main in post-merge - #14854

Merged
xinhe-nv merged 4 commits into
NVIDIA:mainfrom
tensorrt-cicd:trtllm-ci-report/waive-20260602-070758
Jun 2, 2026
Merged

[None][test] Waive 11 failed cases for main in post-merge#14854
xinhe-nv merged 4 commits into
NVIDIA:mainfrom
tensorrt-cicd:trtllm-ci-report/waive-20260602-070758

Conversation

@tensorrt-cicd

@tensorrt-cicd tensorrt-cicd commented Jun 2, 2026

Copy link
Copy Markdown
Collaborator

Auto-generated Waive PR

Created by: TensorRT LLM CI Report (requested by qa@nvidia.com)
Target branch: main
Bug(s): 6059036, 6075533, 6181383, 6211191, 6211193, 6211441, 6215689, 6224636, 6224637, 6240420, 6245279, 6245317, 6245389, 6245394

Waive entries added

accuracy/test_llm_api_pytorch.py::TestDeepSeekV3Lite::test_nvfp4_4gpus[moe_backend=CUTLASS-mtp_nextn=0-pp4-fp8kv=False-attention_dp=False-cuda_graph=False-overlap_scheduler=False-low_precision_combine=False-torch_compile=True] SKIP (https://nvbugs/6211191)
accuracy/test_llm_api_pytorch.py::TestDeepSeekV3Lite::test_fp8_block_scales_4gpus[pp4-mtp_nextn=0-fp8kv=False-attention_dp=True-cuda_graph=True-overlap_scheduler=True-torch_compile=False-sampler_async_worker=False] SKIP (https://nvbugs/6211191)
accuracy/test_llm_api_pytorch.py::TestLlama3_1_8BInstruct::test_fp8_4gpus[pp4-fp8kv=True-attn_backend=FLASHINFER-torch_compile=False] SKIP (https://nvbugs/6211191)
accuracy/test_disaggregated_serving.py::TestDeepSeekV3Lite::test_guided_decoding[llguidance-mtp_nextn=2] SKIP (https://nvbugs/6075533)
accuracy/test_llm_api_pytorch.py::TestLlama3_3_70BInstruct::test_fp4_tp2pp2[torch_compile=True-enable_gemm_allreduce_fusion=True] SKIP (https://nvbugs/6211441)
accuracy/test_llm_api_pytorch.py::TestDeepSeekV3Lite::test_bfloat16_4gpus[pp4-mtp_nextn=0-attention_dp=True-cuda_graph=True-overlap_scheduler=True-torch_compile=False] SKIP (https://nvbugs/6211191)
accuracy/test_llm_api_pytorch.py::TestDeepSeekV3Lite::test_cute_dsl_bf16_gemm_4gpus[tp4-cuda_graph=False] SKIP (https://nvbugs/6224636)
accuracy/test_llm_api_pytorch.py::TestLlama3_1_8BInstruct::test_bfloat16_4gpus[pp4-attn_backend=FLASHINFER-torch_compile=False] SKIP (https://nvbugs/6224637)
disaggregated/test_disaggregated.py::test_disaggregated_gpt_oss_120b_harmony[gpt_oss/gpt-oss-120b] SKIP (https://nvbugs/6245317)
accuracy/test_llm_api_pytorch.py::TestLlama3_1_8BInstruct::test_fp8_4gpus[tp4-fp8kv=False-attn_backend=TRTLLM-torch_compile=True] SKIP (https://nvbugs/6245389)
accuracy/test_llm_api_pytorch.py::TestDeepSeekV3Lite::test_nvfp4_4gpus[moe_backend=CUTLASS-mtp_nextn=2-pp4-fp8kv=True-attention_dp=True-cuda_graph=True-overlap_scheduler=True-low_precision_combine=False-torch_compile=False] SKIP (https://nvbugs/6245394)

Already waived (skipped)

  • test_e2e.py::test_openai_disagg_multi_nodes_completion_service_discovery[etcd]
  • accuracy/test_llm_api_pytorch_multimodal.py::TestGemma3_27BInstruct::test_fp8_prequantized
  • accuracy/test_llm_api_autodeploy.py::TestModelRegistryAccuracy::test_autodeploy_from_registry[google_gemma-3-1b-it-False]
  • accuracy/test_llm_api_autodeploy.py::TestModelRegistryAccuracy::test_autodeploy_from_registry[mistralai_Ministral-8B-Instruct-2410-False]
  • accuracy/test_llm_api_autodeploy.py::TestModelRegistryAccuracy::test_autodeploy_from_registry[nvidia_Llama-3.1-8B-Instruct-NVFP4-True]
  • accuracy/test_llm_api_pytorch.py::TestMistralLarge3_675B::test_nvfp4_4gpus[latency_moe_trtllm]

This PR was auto-generated by TensorRT LLM CI Report. Please review the waive entries before merging.

Summary by CodeRabbit

  • Tests
    • Updated test waiver configurations to optimize test execution for specific model configurations, adjusting skip lists for improved testing efficiency across different hardware and quantization settings.

Bug(s): 6059036, 6075533, 6181383, 6211191, 6211193, 6211441, 6215689, 6224636, 6224637, 6240420, 6245279, 6245317, 6245389, 6245394
Requested by: qa@nvidia.com

Signed-off-by: tensorrt-cicd <90828364+tensorrt-cicd@users.noreply.github.com>
@xinhe-nv
xinhe-nv self-requested a review June 2, 2026 07:08
@xinhe-nv
xinhe-nv enabled auto-merge (squash) June 2, 2026 07:08
@xinhe-nv

xinhe-nv commented Jun 2, 2026

Copy link
Copy Markdown
Collaborator

/bot run --stage-list ""

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator Author

PR_Github #51544 [ run ] triggered by Bot. Commit: 3641784 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator Author

PR_Github #51544 [ run ] completed with state SUCCESS. Commit: 3641784
/LLM/main/L0_MergeRequest_PR pipeline #40940 (Partly Tested) completed with status: 'SUCCESS'

CI Report

Link to invocation

@xinhe-nv

xinhe-nv commented Jun 2, 2026

Copy link
Copy Markdown
Collaborator

/bot reuse-pipeline

@coderabbitai

coderabbitai Bot commented Jun 2, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

📝 Walkthrough

Walkthrough

Updated tests/integration/test_lists/waives.txt to add 11 test waiver entries for failing integration tests. Changes include waivers for DeepSeekV3Lite guided decoding, quantization, and torch compilation configurations; Llama3 variants with specific attention backends and fusion settings; and disaggregated GPT-OSS serving.

Changes

Test waiver configuration updates

Layer / File(s) Summary
DeepSeekV3Lite test waivers
tests/integration/test_lists/waives.txt
Added SKIP waivers for DeepSeekV3Lite guided decoding with llguidance, extended bfloat16_4gpus and fp8_block_scales_4gpus configurations with new parallel strategy combinations, and added nvfp4_4gpus waivers with torch_compile, fp8kv, and attention_dp settings.
Llama3 and GPT-OSS test waivers
tests/integration/test_lists/waives.txt
Added Llama3_1_8BInstruct waivers for bfloat16 with FLASHINFER attention and fp8 configurations with torch_compile, inserted Llama3_3_70BInstruct fp4 tensor-parallel/pipeline-parallel waiver with gemm allreduce fusion, and added disaggregated GPT-OSS 120B harmony test waiver.

Possibly related PRs

  • NVIDIA/TensorRT-LLM#13986: Adds overlapping DeepSeekV3Lite and Llama3 test waivers for the same bfloat16/fp8/nvfp4 GPU configurations.
  • NVIDIA/TensorRT-LLM#14789: Updates the same waives.txt file with additional test SKIP entries for failing CI tests.
  • NVIDIA/TensorRT-LLM#13200: Adds related Llama3_1_8BInstruct and GPT-OSS test waivers for fp8 and attention/compilation configurations.

Suggested reviewers

  • jieli-matrix
  • StanleySun639
  • xinhe-nv

🎯 1 (Trivial) | ⏱️ ~3 minutes

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 inconclusive)

Check name Status Explanation Resolution
Description check ❓ Inconclusive The description is missing the structured sections (Description, Test Coverage) required by the template, though it includes auto-generated content with bug links and waive entries. Add explicit Description and Test Coverage sections explaining why the waivers are needed and what tests are impacted, even if auto-generated.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly describes the main change: waiving 11 failed test cases for the main branch post-merge, with the exact number and purpose specified.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/integration/test_lists/waives.txt`:
- Line 100: The waiver removes FLASHINFER multi-GPU coverage for the test
accuracy/test_llm_api_pytorch.py::TestLlama3_1_8BInstruct::test_bfloat16_4gpus
(and the fp8kv+FLASHINFER combination), creating a coverage gap; reinstate or
replace these skips by either re-enabling the skipped test(s) for pp4 FLASHINFER
or add equivalent tests that exercise FLASHINFER on other parallel configs
(e.g., tp4 or different sharding) so FLASHINFER backend gets multi-GPU
validation, and ensure the test names
(TestLlama3_1_8BInstruct::test_bfloat16_4gpus and the fp8kv+FLASHINFER case) are
present in the test matrix and not skipped without a tracked bug reference.
- Line 67: The waivers file currently skips multiple tests where
torch_compile=True for quantized models (e.g.,
accuracy/test_llm_api_pytorch.py::TestDeepSeekV3Lite::test_nvfp4_4gpus[...torch_compile=True],
the Llama3_1_8BInstruct and Llama3_3_70BInstruct entries) without triage info;
update tests/integration/test_lists/waives.txt to consolidate these
torch_compile-related waivers into a single clearly labeled section, add the
root cause or ticket reference (bug/feature ID) and one of: planned ETA if work
is active, explicit "unsupported" note if not supported, or a temporary waiver
expiration date if low priority, and include the exact test identifiers (the
three failing test names) and an owner assigned for follow-up so QA knows
whether/when to re-enable them.
- Line 47: The waiver entry grouping NVBug 6211191 covers multiple PP4 test
variants that do not share the same feature combination; update the waives so
NVBug 6211191 only applies to the exact failing combos (e.g.,
accuracy/test_llm_api_pytorch.py::TestDeepSeekV3Lite::test_bfloat16_4gpus[...]
and ::test_fp8_block_scales_4gpus[...] with
attention_dp=True,cuda_graph=True,overlap_scheduler=True), and split or add
distinct NVBug references for the other PP4 entries that have different flags
(e.g., ::test_nvfp4_4gpus[...] where attention_dp/cuda_graph/overlap_scheduler
are False and ::TestLlama3_1_8BInstruct::test_fp8_4gpus[...] with
fp8kv=True/attn_backend=FLASHINFER); ensure each waived test line in waives.txt
references the correct NVBug URL matching the exact feature combo.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 4af13a6b-e888-4800-82f0-9fbb70fd952c

📥 Commits

Reviewing files that changed from the base of the PR and between 209f371 and 62a429c.

📒 Files selected for processing (1)
  • tests/integration/test_lists/waives.txt

Comment thread tests/integration/test_lists/waives.txt
Comment thread tests/integration/test_lists/waives.txt
Comment thread tests/integration/test_lists/waives.txt
@tensorrt-cicd

Copy link
Copy Markdown
Collaborator Author

PR_Github #51570 [ reuse-pipeline ] triggered by Bot. Commit: 62a429c Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator Author

PR_Github #51570 [ reuse-pipeline ] completed with state SUCCESS. Commit: 62a429c
Reusing PR_Github #51544 (Partly Tested) for commit 62a429c

Link to invocation

@xinhe-nv

xinhe-nv commented Jun 2, 2026

Copy link
Copy Markdown
Collaborator

/bot reuse-pipeline

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator Author

PR_Github #51577 [ reuse-pipeline ] triggered by Bot. Commit: 5325337 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator Author

PR_Github #51577 [ reuse-pipeline ] completed with state SUCCESS. Commit: 5325337
Reusing PR_Github #51544 (Partly Tested) for commit 5325337

Link to invocation

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants