Skip to content

[None][fix] make FA4 proper pip dependency - #13788

Merged
o-stoner merged 1 commit into
NVIDIA:mainfrom
o-stoner:user/o-stoner/visual-gen-update-fa4
May 27, 2026
Merged

[None][fix] make FA4 proper pip dependency#13788
o-stoner merged 1 commit into
NVIDIA:mainfrom
o-stoner:user/o-stoner/visual-gen-update-fa4

Conversation

@o-stoner

@o-stoner o-stoner commented May 6, 2026

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

Release Notes

  • Chores

    • Updated Flash Attention implementation to use external library, improving kernel compatibility and maintainability.
  • Refactor

    • Consolidated internal attention kernel utilities and removed legacy helper implementations.

Description

Test Coverage

PR Checklist

Please review the following before submitting your PR:

  • PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.

  • PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.

  • Test cases are provided for new code paths (see test instructions)

  • Any new dependencies have been scanned for license and vulnerabilities

  • CODEOWNERS updated if ownership changes

  • Documentation updated as needed

  • Update tava architecture diagram if there is a significant design change in PR.

  • The reviewers assigned automatically/manually are appropriate for the PR.

  • Please check this after reviewing the above items as appropriate for this PR.

GitHub Bot Help

To see a list of available CI bot commands, please comment /bot help.

@o-stoner

o-stoner commented May 6, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --add-multi-gpu-test

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #46890 [ run ] triggered by Bot. Commit: de0a859 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #46890 [ run ] completed with state SUCCESS. Commit: de0a859
/LLM/main/L0_MergeRequest_PR pipeline #36900 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

Link to invocation

@o-stoner
o-stoner force-pushed the user/o-stoner/visual-gen-update-fa4 branch from de0a859 to 0c0f333 Compare May 6, 2026 17:50
@o-stoner

o-stoner commented May 6, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --add-multi-gpu-test

2 similar comments
@o-stoner

o-stoner commented May 6, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --add-multi-gpu-test

@o-stoner

o-stoner commented May 6, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --add-multi-gpu-test

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #47045 [ run ] triggered by Bot. Commit: 0c0f333 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #47045 [ run ] completed with state SUCCESS. Commit: 0c0f333
/LLM/main/L0_MergeRequest_PR pipeline #37018 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@o-stoner

o-stoner commented May 7, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --add-multi-gpu-test

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #47230 [ run ] triggered by Bot. Commit: 0c0f333 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #47230 [ run ] completed with state SUCCESS. Commit: 0c0f333
/LLM/main/L0_MergeRequest_PR pipeline #37186 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

Link to invocation

@o-stoner

o-stoner commented May 8, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --add-multi-gpu-test

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #47431 [ run ] triggered by Bot. Commit: c9b3ed9 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #47431 [ run ] completed with state FAILURE. Commit: c9b3ed9
/LLM/main/L0_MergeRequest_PR pipeline #37354 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@o-stoner

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --add-multi-gpu-test

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #47622 [ run ] triggered by Bot. Commit: c9b3ed9 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #47622 [ run ] completed with state SUCCESS. Commit: c9b3ed9
/LLM/main/L0_MergeRequest_PR pipeline #37527 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@o-stoner
o-stoner force-pushed the user/o-stoner/visual-gen-update-fa4 branch from c9b3ed9 to aaa667a Compare May 12, 2026 20:59
@o-stoner

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --add-multi-gpu-test

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #48027 [ run ] triggered by Bot. Commit: a98e045 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #48027 [ run ] completed with state SUCCESS. Commit: a98e045
/LLM/main/L0_MergeRequest_PR pipeline #37864 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@o-stoner

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --add-multi-gpu-test

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #48212 [ run ] triggered by Bot. Commit: a98e045 Link to invocation

@o-stoner
o-stoner marked this pull request as ready for review May 13, 2026 20:05
@o-stoner
o-stoner requested review from a team as code owners May 13, 2026 20:05
@coderabbitai

coderabbitai Bot commented May 13, 2026

Copy link
Copy Markdown
Contributor

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: c8a5d0b4-7966-4728-8bb5-2cf9637eedf2

📥 Commits

Reviewing files that changed from the base of the PR and between ca876e0 and a98e045.

📒 Files selected for processing (43)
  • requirements.txt
  • tensorrt_llm/_torch/visual_gen/attention_backend/flash_attn4.py
  • tensorrt_llm/_torch/visual_gen/attention_backend/parallel.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/__init__.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/__init__.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/.flake8
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/AUTHORS
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/LICENSE
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/README.md
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/__init__.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/ampere_helpers.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/barrier.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/benchmark.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/blackwell_helpers.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/block_info.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/block_sparse_utils.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/block_sparsity.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/compute_block_sparsity.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/copy_utils.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/cute_dsl_utils.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/fast_math.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/flash_bwd.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/flash_bwd_postprocess.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/flash_bwd_preprocess.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/flash_bwd_sm100.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/flash_bwd_sm90.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/flash_fwd.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/flash_fwd_combine.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/flash_fwd_sm100.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/hopper_helpers.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/interface.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/mask.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/mma_sm100_desc.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/named_barrier.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/pack_gqa.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/paged_kv.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/pipeline.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/pyproject.toml
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/seqlen_info.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/softmax.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/testing.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/tile_scheduler.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/utils.py
💤 Files with no reviewable changes (34)
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/LICENSE
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/AUTHORS
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/init.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/softmax.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/pack_gqa.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/pyproject.toml
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/flash_fwd_combine.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/block_info.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/seqlen_info.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/barrier.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/paged_kv.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/benchmark.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/flash_bwd.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/fast_math.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/hopper_helpers.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/tile_scheduler.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/testing.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/ampere_helpers.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/named_barrier.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/flash_bwd_preprocess.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/compute_block_sparsity.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/flash_bwd_sm90.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/block_sparsity.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/interface.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/.flake8
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/blackwell_helpers.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/cute_dsl_utils.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/mma_sm100_desc.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/flash_bwd_postprocess.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/mask.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/copy_utils.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/block_sparse_utils.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/pipeline.py
  • tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/utils.py

📝 Walkthrough

Walkthrough

This PR migrates from a vendored Flash Attention 4 implementation to an external dependency. Requirements now pin flash-attn-4==4.0.0b11. Import paths in attention backends are updated to use flash_attn.cute.interface instead of project-local paths. The entire vendored tensorrt_llm._torch.visual_gen.jit_kernels.flash_attention.cute module is removed.

Changes

Flash Attention External Dependency Migration

Layer / File(s) Summary
Dependency declaration
requirements.txt
Added explicit pinned requirement for flash-attn-4==4.0.0b11 to supply the Flash Attention kernels externally.
Import path updates to external package
tensorrt_llm/_torch/visual_gen/attention_backend/flash_attn4.py, tensorrt_llm/_torch/visual_gen/attention_backend/parallel.py
Updated _flash_attn_fwd and _flash_attn_combine imports to source from flash_attn.cute.interface instead of the removed local tensorrt_llm._torch.visual_gen.jit_kernels...cute.interface paths. Try/except error handling and ImportError fallback behavior remain unchanged.
Removal of vendored Flash Attention implementation
tensorrt_llm/_torch/visual_gen/jit_kernels/flash_attention/cute/*
Deleted entire vendored implementation directory including CUDA/CUTE JIT kernels, block-sparsity utilities, tile scheduling, softmax, masking, backward kernels, and supporting infrastructure (~40 files, ~11,000+ lines). Module docstring reference to upstream source also removed from flash_attn4.py.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Suggested reviewers

  • laikhtewari
  • kaiyux
  • yuxianq
  • zhenhuaw-me
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The PR description is empty except for the template structure and checklist; it lacks any actual explanation of the changes, test coverage, or implementation details. Add a clear description explaining the rationale for converting FA4 to a pip dependency, list relevant test coverage, and document any breaking changes or migration steps.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly and specifically describes the main objective: making FA4 (Flash Attention 4) a proper pip dependency, which aligns with the changeset.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Tip

💬 Introducing Slack Agent: The best way for teams to turn conversations into code.

Slack Agent is built on CodeRabbit's deep understanding of your code, so your team can collaborate across the entire SDLC without losing context.

  • Generate code and open pull requests
  • Plan features and break down work
  • Investigate incidents and troubleshoot customer tickets together
  • Automate recurring tasks and respond to alerts with triggers
  • Summarize progress and report instantly

Built for teams:

  • Shared memory across your entire org—no repeating context
  • Per-thread sandboxes to safely plan and execute work
  • Governance built-in—scoped access, auditability, and budget controls

One agent for your entire SDLC. Right inside Slack.

👉 Get started


Comment @coderabbitai help to get the list of available commands and usage tips.

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #48212 [ run ] completed with state ABORTED. Commit: a98e045

Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #48943 [ run ] triggered by Bot. Commit: 7e40ce4 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #48943 [ run ] completed with state FAILURE. Commit: 7e40ce4
/LLM/main/L0_MergeRequest_PR pipeline #38689 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@o-stoner

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --add-multi-gpu-test

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #48956 [ run ] triggered by Bot. Commit: 094fc63 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #48956 [ run ] completed with state FAILURE. Commit: 094fc63
/LLM/main/L0_MergeRequest_PR pipeline #38702 completed with status: 'ABORTED'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

Link to invocation

@o-stoner

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --add-multi-gpu-test --disable-artifact-copy

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #48961 Bot args parsing error: usage: /bot [-h]
{run,kill,skip,submit,reviewers,reuse-pipeline,reuse-review} ...
/bot: error: unrecognized arguments: --disable-artifact-copy

Link to invocation

@o-stoner

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --add-multi-gpu-test

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #48974 [ run ] triggered by Bot. Commit: 094fc63 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #48974 [ run ] completed with state SUCCESS. Commit: 094fc63
/LLM/main/L0_MergeRequest_PR pipeline #38719 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@o-stoner

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --add-multi-gpu-test

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #49236 [ run ] triggered by Bot. Commit: 094fc63 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #49236 [ run ] completed with state SUCCESS. Commit: 094fc63
/LLM/main/L0_MergeRequest_PR pipeline #38908 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@o-stoner

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --add-multi-gpu-test

1 similar comment
@o-stoner

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --add-multi-gpu-test

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #49476 [ run ] triggered by Bot. Commit: 0c1b24a Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #49476 [ run ] completed with state SUCCESS. Commit: 0c1b24a
/LLM/main/L0_MergeRequest_PR pipeline #39118 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

Signed-off-by: Olivia Stoner <245287810+o-stoner@users.noreply.github.com>
@o-stoner
o-stoner force-pushed the user/o-stoner/visual-gen-update-fa4 branch from 0c1b24a to 9049c01 Compare May 22, 2026 21:13
@o-stoner

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --add-multi-gpu-test

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #49988 [ run ] triggered by Bot. Commit: 9049c01 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #49988 [ run ] completed with state FAILURE. Commit: 9049c01
/LLM/main/L0_MergeRequest_PR pipeline #39553 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@o-stoner

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --add-multi-gpu-test

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #50349 [ run ] triggered by Bot. Commit: 9049c01 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #50349 [ run ] completed with state SUCCESS. Commit: 9049c01
/LLM/main/L0_MergeRequest_PR pipeline #39876 completed with status: 'SUCCESS'

CI Report

Link to invocation

@o-stoner
o-stoner merged commit 9fa1358 into NVIDIA:main May 27, 2026
7 checks passed
bmarimuthu-nv pushed a commit to nv-auto-deploy/TensorRT-LLM that referenced this pull request May 28, 2026
Signed-off-by: Olivia Stoner <245287810+o-stoner@users.noreply.github.com>
luyiyun1021 added a commit to luyiyun1021/TensorRT-LLM that referenced this pull request May 28, 2026
…ysses

Per reviewer feedback, remove the ``parallel.audio_pad_for_ulysses``
ParallelConfig knob and derive audio padding behavior internally based on
the runtime context. With Ulysses active, audio is always padded to make
``T_a`` divisible by ``ulysses_size`` and a ``[B, T_a_padded]`` validity
mask is attached so attention zeros pad slots; with Ulysses inactive
nothing changes. The opt-out mode the knob exposed was a silent
performance downgrade (disengaged the v2a Ulysses wrapper, fell back to
plain attention on non-divisible T_a) and not something users should be
deciding.

Source changes (all behavior-equivalent to the previous True default):
* Drop ``ParallelConfig.audio_pad_for_ulysses``.
* ``LTX2Attention._init_audio_modules``: TRTLLM→VANILLA downgrade for
  ``audio_attn1`` now keys on ``vgm.ulysses_size > 1`` (the condition
  that actually drives the mask requirement) instead of the dropped flag.
* ``LTXModel.__init__``: drop the cached ``self._audio_pad_for_ulysses``.
* ``LTXModel.configure_audio_ulysses``: always pad when the sharder is
  active. Removes the dead "no-pad mode" branch and the CP-without-
  Ulysses ``ValueError`` (now unreachable because padding always makes
  audio shardable).
* Docstrings + the ``audio_padding_mask`` field comment updated.

Tests adjusted: drop the parameter from the three call sites in
test_ltx2_ulysses.py / test_ulysses_attention.py / test_fa4_key_padding_mask.py.
Verified locally with the multi-GPU e2e LTXModel parity test
(VANILLA backend, ws=2): both ``test_av_ulysses_no_audio_pad`` and
``test_av_ulysses_audio_pad`` PASS. The FA4 backend cases failed in this
container due to an unrelated env regression already fixed on main
(PR NVIDIA#13788, ``nvvm.fmax`` API change in ``nvidia_cutlass_dsl``); CI on
a fresh container picks up the fix automatically.

Signed-off-by: Yiyun Lu <55233584+luyiyun1021@users.noreply.github.com>
luyiyun1021 added a commit to luyiyun1021/TensorRT-LLM that referenced this pull request May 29, 2026
…ysses

Per reviewer feedback, remove the ``parallel.audio_pad_for_ulysses``
ParallelConfig knob and derive audio padding behavior internally based on
the runtime context. With Ulysses active, audio is always padded to make
``T_a`` divisible by ``ulysses_size`` and a ``[B, T_a_padded]`` validity
mask is attached so attention zeros pad slots; with Ulysses inactive
nothing changes. The opt-out mode the knob exposed was a silent
performance downgrade (disengaged the v2a Ulysses wrapper, fell back to
plain attention on non-divisible T_a) and not something users should be
deciding.

Source changes (all behavior-equivalent to the previous True default):
* Drop ``ParallelConfig.audio_pad_for_ulysses``.
* ``LTX2Attention._init_audio_modules``: TRTLLM→VANILLA downgrade for
  ``audio_attn1`` now keys on ``vgm.ulysses_size > 1`` (the condition
  that actually drives the mask requirement) instead of the dropped flag.
* ``LTXModel.__init__``: drop the cached ``self._audio_pad_for_ulysses``.
* ``LTXModel.configure_audio_ulysses``: always pad when the sharder is
  active. Removes the dead "no-pad mode" branch and the CP-without-
  Ulysses ``ValueError`` (now unreachable because padding always makes
  audio shardable).
* Docstrings + the ``audio_padding_mask`` field comment updated.

Tests adjusted: drop the parameter from the three call sites in
test_ltx2_ulysses.py / test_ulysses_attention.py / test_fa4_key_padding_mask.py.
Verified locally with the multi-GPU e2e LTXModel parity test
(VANILLA backend, ws=2): both ``test_av_ulysses_no_audio_pad`` and
``test_av_ulysses_audio_pad`` PASS. The FA4 backend cases failed in this
container due to an unrelated env regression already fixed on main
(PR NVIDIA#13788, ``nvvm.fmax`` API change in ``nvidia_cutlass_dsl``); CI on
a fresh container picks up the fix automatically.

Signed-off-by: Yiyun Lu <55233584+luyiyun1021@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants