[https://nvbugs/6401921][fix] Stabilize single-GPU Wan2.2/LTX2/Wan2.1 LPIPS test#15854
Conversation
Signed-off-by: Chang Liu <9713593+chang-l@users.noreply.github.com>
461dc9c to
247f5e1
Compare
📝 WalkthroughWalkthroughThis change updates the golden video test metadata (adding torch_version, updating TensorRT-LLM version/commit and container image), wraps the Wan 2.2 LPIPS video generation call in a forced eager compiler stance, and removes the corresponding test waiver from waives.txt. ChangesWan22 T2V LPIPS golden test fix
Estimated code review effort: 2 (Simple) | ~10 minutes Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
|
/bot run --disable-fail-fast --extra-stage "DGX_B200-PyTorch-Post-Merge-1, DGX_B200-PyTorch-Post-Merge-2" |
|
PR_Github #57100 [ run ] triggered by Bot. Commit: |
|
PR_Github #57100 [ run ] completed with state
|
|
/bot run --disable-fail-fast |
|
PR_Github #57140 [ run ] triggered by Bot. Commit: |
|
PR_Github #57140 [ run ] completed with state
|
|
/bot run --disable-fail-fast |
Signed-off-by: Chang Liu <9713593+chang-l@users.noreply.github.com>
…ed during conflict resolution Signed-off-by: Chang Liu <9713593+chang-l@users.noreply.github.com>
|
/bot run --disable-fail-fast --extra-stage "DGX_B200-PyTorch-Post-Merge-1, DGX_B200-PyTorch-Post-Merge-2" |
|
PR_Github #57221 [ run ] triggered by Bot. Commit: |
|
PR_Github #57221 [ run ] completed with state
|
|
/bot run --stage-list "DGX_B200-PyTorch-Post-Merge-1" |
…eration and preserve failing candidates Wan2.1 and LTX-2 LPIPS goldens were generated on the 26.02 container; the CI container moved to 26.04 and both tests now fail deterministically in B200 post-merge (wan21 0.0956, ltx2 0.1513 vs the 0.05 threshold). Cross-machine eager variance on the same container measures ~0.04 LPIPS for the 1-step Wan2.1 config, so goldens regenerated on a dev machine leave no reliable margin. - Run Wan2.1 and LTX-2 LPIPS generation under torch.compiler.set_stance (force_eager), matching the Wan2.2 fix; the LTX-2 wrap covers the golden fixture and both sides of the cuda-graph-vs-eager comparison. - On an LPIPS threshold failure, copy the generated candidate into pytest's --output-dir (archived per-stage by CI) so the golden can be refreshed with CI-generated media instead of dev-machine approximations. Signed-off-by: Chang Liu <9713593+chang-l@users.noreply.github.com>
|
/bot run --disable-fail-fast --extra-stage "DGX_B200-PyTorch-Post-Merge-1, DGX_B200-PyTorch-Post-Merge-2" |
|
PR_Github #57279 [ run ] triggered by Bot. Commit: |
…the 26.04 stack The previous goldens were generated on the 26.02 container and fail deterministically on 26.04 CI (wan21 0.0956, ltx2 0.1513 vs 0.05). Regenerated on B200 with the CI devel image (pytorch-26.04, tag -15694), the CI-built 1.3.0rc21 wheel, torch 2.12.0a0+0291f960b6.nv26.04, and force_eager generation. Only the wan21/ltx2 zip members changed; both tests score LPIPS 0.000000 against these goldens on the generating host. Signed-off-by: Chang Liu <9713593+chang-l@users.noreply.github.com>
|
/bot run --disable-fail-fast --extra-stage "DGX_B200-PyTorch-Post-Merge-1, DGX_B200-PyTorch-Post-Merge-2" |
Signed-off-by: Chang Liu <9713593+chang-l@users.noreply.github.com>
|
/bot run --disable-fail-fast --extra-stage "DGX_B200-PyTorch-Post-Merge-1, DGX_B200-PyTorch-Post-Merge-2" |
1 similar comment
|
/bot run --disable-fail-fast --extra-stage "DGX_B200-PyTorch-Post-Merge-1, DGX_B200-PyTorch-Post-Merge-2" |
|
PR_Github #57351 [ run ] triggered by Bot. Commit: |
|
PR_Github #57352 [ run ] triggered by Bot. Commit: |
|
PR_Github #57527 [ run ] completed with state |
|
PR_Github #57534 [ run ] completed with state
|
…a-graph LPIPS test test_ltx2_cuda_graph_lpips_matches_eager never requested _visual_gen_deps (which installs ffmpeg), so when it ran first in a stage it previously passed only via the silent cv2/mp4v fallback on both sides of the comparison. With that fallback now a hard failure, declare the fixture so ffmpeg is installed before encoding. Signed-off-by: Chang Liu <9713593+chang-l@users.noreply.github.com>
|
/bot run --disable-fail-fast --extra-stage "DGX_B200-PyTorch-Post-Merge-1, DGX_B200-PyTorch-Post-Merge-2" |
|
PR_Github #57552 [ run ] triggered by Bot. Commit: |
|
PR_Github #57552 [ run ] completed with state
|
|
/bot run --disable-fail-fast --extra-stage "DGX_B200-PyTorch-Post-Merge-1, DGX_B200-PyTorch-Post-Merge-2" |
|
PR_Github #57559 [ run ] triggered by Bot. Commit: |
|
PR_Github #57559 [ run ] completed with state
|
…en by NVIDIA#14827 Same root cause as the already-waived T2V variant: NVIDIA#14827 changed the effective Cosmos3 generation parameters (_resolve_t2i_default rewrites the test's pinned steps/guidance/resolution because they equal the video defaults), so the T2I output no longer matches its golden (LPIPS 0.608, deterministic across three runs: pipelines 46107/46265/46287). Signed-off-by: Chang Liu <9713593+chang-l@users.noreply.github.com>
|
/bot run --disable-fail-fast --extra-stage "DGX_B200-PyTorch-Post-Merge-1, DGX_B200-PyTorch-Post-Merge-2" |
|
/bot run |
|
PR_Github #57577 [ run ] triggered by Bot. Commit: |
|
PR_Github #57578 [ run ] triggered by Bot. Commit: |
|
PR_Github #57577 [ run ] completed with state |
|
PR_Github #57578 [ run ] completed with state
|
|
/bot run |
|
PR_Github #57617 [ run ] triggered by Bot. Commit: |
|
PR_Github #57617 [ run ] completed with state
|
|
/bot run --disable-fail-fast |
|
PR_Github #57624 [ run ] triggered by Bot. Commit: |
|
PR_Github #57624 [ run ] completed with state |
|
✅ LFS objects already in storage (1 file) — no sync needed. These LFS-tracked files are already present in this repository's LFS storage:
|
…orrt_llm site dir Address review: replace the loose truthiness check on tllm_site with a validator that requires the directory to contain the installed tensorrt_llm package with compiled bindings, applied both before spawning and inside each worker so a bad environment fails loudly. Also record that the single-GPU wan22 golden is now generated under force_eager (post NVIDIA#15854). Signed-off-by: Chang Liu <9713593+chang-l@users.noreply.github.com>
… the 26.04 container Main moved the CI base container from 26.02 to 26.04 (and NVIDIA#15854 refreshed the single-GPU goldens accordingly), so the multi-GPU FA4 fully-eager reference is regenerated in the current nightly staging release image. Two independent generations produced bit-identical MP4s (single SHA-256) with zero dynamo compile counters. The media zip is main's post-NVIDIA#15854 archive plus the new FA4 reference; sidecar provenance now records the image digest, torch/diffusers versions, and the image's main commit. Signed-off-by: Chang Liu <9713593+chang-l@users.noreply.github.com>
…VIDIA#15854) Signed-off-by: Chang Liu <9713593+chang-l@users.noreply.github.com>
Summary by CodeRabbit
Bug Fixes
Tests
Description
Fix NVBug 6401921, where
test_wan22_t2v_lpips_against_goldenregressed to LPIPS 0.251536 after thePyTorch CI container moved from 26.02 to 26.04.
TorchCompileConfig(enable=False)skips VisualGen's configured transformercompilation, but it does not suppress nested or unconditional
@torch.compilecall sites. The resulting execution trajectory changed acrossthe PyTorch upgrade. A controlled A/B showed that forcing both containers fully
eager makes their final pre-VAE latents bit-identical and keeps cross-container
LPIPS below the existing 0.05 threshold.
This change:
torch.compiler.set_stance("force_eager");26.04 CI image and records the exact runtime provenance;
The scope is intentionally limited to the failing Wan2.2 single-GPU golden.
Other VisualGen goldens retain their existing execution policy. This is the
single-GPU counterpart to #15730 and does not depend on it.
Test Coverage
Exact B200 reproduction with PyTorch 26.04 image
sha256:dad31c0b5290d836033c96d8b91f6524bdc7cc5b4d1000b4abcc57c6868ffdc0and Jenkins post-merge build 2814 (
tensorrt-llm==1.3.0rc21, commit539ee226c4df7ab15802911083fe501e9d64c66e).The regenerated force-eager video reproduced byte-for-byte across two fresh
runs: SHA-256
52828186f44b82a9f686f177d635b9f3cb0050f41c8d3ae55dade01d30a00b28.Targeted integration test:
zip -Tpassed, the archive still contains eight unique members, and onlywan22_t2v_lpips_golden_video.mp4differs from the previous archive.All pre-commit hooks passed on the four changed files.
PR Checklist
GitHub Bot Help
To see a list of available CI bot commands, please comment
/bot help.