Skip to content

[None][perf] Fold q/k/v quantization into qknorm_rope_fused kernel & remove contiguous - #16699

Merged
brb-nv merged 1 commit into
NVIDIA:feat/m3_with_msafrom
brb-nv:user/brb/qk-quant-contiguous
Jul 22, 2026
Merged

[None][perf] Fold q/k/v quantization into qknorm_rope_fused kernel & remove contiguous#16699
brb-nv merged 1 commit into
NVIDIA:feat/m3_with_msafrom
brb-nv:user/brb/qk-quant-contiguous

Conversation

@brb-nv

@brb-nv brb-nv commented Jul 22, 2026

Copy link
Copy Markdown
Collaborator

Description

This MR does the following:

  • Folds q/k/v quantization into fused qk norm + rope kernel.
  • Removes .contiguous() calls on main q/k as well as index q/k.

Test Coverage

$ pytest tests/integration/defs/accuracy/test_llm_api_pytorch.py::TestMiniMaxM3::test_nvfp4[use_msa=True] -s -v

PR Checklist

Please review the following before submitting your PR:

  • PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.

  • PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.

  • Test cases are provided for new code paths (see test instructions)

  • If PR introduces API changes, an appropriate PR label is added - either api-compatible or api-breaking. For api-breaking, include BREAKING in the PR title.

  • Any new dependencies have been scanned for license and vulnerabilities

  • CODEOWNERS updated if ownership changes

  • Documentation updated as needed

  • Update tava architecture diagram if there is a significant design change in PR.

  • The reviewers assigned automatically/manually are appropriate for the PR.

  • Please check this after reviewing the above items as appropriate for this PR.

GitHub Bot Help

To see a list of available CI bot commands, please comment /bot help.

@brb-nv
brb-nv requested review from a team as code owners July 22, 2026 01:44
@brb-nv
brb-nv requested review from kaiyux, kris1025 and yunruis July 22, 2026 01:44
@brb-nv
brb-nv requested review from pcicotti, peihu-nv and zheyuf July 22, 2026 01:45
…remove contiguous calls

Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>
@brb-nv
brb-nv force-pushed the user/brb/qk-quant-contiguous branch from 786e4d9 to f008ad0 Compare July 22, 2026 19:51
@brb-nv
brb-nv merged commit 6c5ec86 into NVIDIA:feat/m3_with_msa Jul 22, 2026
6 checks passed
brb-nv added a commit to brb-nv/TensorRT-LLM that referenced this pull request Jul 23, 2026
…remove contiguous (NVIDIA#16699)

Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>
ZhanruiSunCh pushed a commit that referenced this pull request Jul 29, 2026
…remove contiguous (#16699)

Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>
brb-nv added a commit to brb-nv/TensorRT-LLM that referenced this pull request Aug 1, 2026
…remove contiguous (NVIDIA#16699)

Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>
brb-nv added a commit to brb-nv/TensorRT-LLM that referenced this pull request Aug 3, 2026
…remove contiguous (NVIDIA#16699)

Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants