Skip to content

[None][Feat] WIP DO NOT REVIEW Nemotron moe fp8 - #8598

Closed
nvchenghaoz wants to merge 16 commits into
NVIDIA:mainfrom
nv-auto-deploy:nemotron-moe-fp8
Closed

[None][Feat] WIP DO NOT REVIEW Nemotron moe fp8#8598
nvchenghaoz wants to merge 16 commits into
NVIDIA:mainfrom
nv-auto-deploy:nemotron-moe-fp8

Conversation

@nvchenghaoz

@nvchenghaoz nvchenghaoz commented Oct 23, 2025

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • New Features

    • Added fused MoE (Mixture of Experts) alignment kernel for optimized token distribution
    • Introduced Triton-backed fused MoE with FP8 quantization support
    • Added gated RMSNorm with optimized Triton backend
    • Enabled chunked prefill configuration for better memory management
    • Extended MoE flexibility with multiple MLP styles and activation function options
  • Improvements

    • Enhanced grouped attention pattern matching with optimized repeat_kv variants
    • Improved KV cache size tracking for better memory awareness
    • Simplified executor initialization by consolidating configuration sources

Description

Test Coverage

PR Checklist

Please review the following before submitting your PR:

  • PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.

  • PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.

  • Test cases are provided for new code paths (see test instructions)

  • Any new dependencies have been scanned for license and vulnerabilities

  • CODEOWNERS updated if ownership changes

  • Documentation updated as needed

  • The reviewers assigned automatically/manually are appropriate for the PR.

  • Please check this after reviewing the above items as appropriate for this PR.

GitHub Bot Help

/bot [-h] ['run', 'kill', 'skip', 'reuse-pipeline'] ...

Provide a user friendly way for developers to interact with a Jenkins server.

Run /bot [-h|--help] to print this help message.

See details below for each supported subcommand.

Details

run [--reuse-test (optional)pipeline-id --disable-fail-fast --skip-test --stage-list "A10-PyTorch-1, xxx" --gpu-type "A30, H100_PCIe" --test-backend "pytorch, cpp" --add-multi-gpu-test --only-multi-gpu-test --disable-multi-gpu-test --post-merge --extra-stage "H100_PCIe-TensorRT-Post-Merge-1, xxx" --detailed-log --debug(experimental)]

Launch build/test pipelines. All previously running jobs will be killed.

--reuse-test (optional)pipeline-id (OPTIONAL) : Allow the new pipeline to reuse build artifacts and skip successful test stages from a specified pipeline or the last pipeline if no pipeline-id is indicated. If the Git commit ID has changed, this option will be always ignored. The DEFAULT behavior of the bot is to reuse build artifacts and successful test results from the last pipeline.

--disable-reuse-test (OPTIONAL) : Explicitly prevent the pipeline from reusing build artifacts and skipping successful test stages from a previous pipeline. Ensure that all builds and tests are run regardless of previous successes.

--disable-fail-fast (OPTIONAL) : Disable fail fast on build/tests/infra failures.

--skip-test (OPTIONAL) : Skip all test stages, but still run build stages, package stages and sanity check stages. Note: Does NOT update GitHub check status.

--stage-list "A10-PyTorch-1, xxx" (OPTIONAL) : Only run the specified test stages. Examples: "A10-PyTorch-1, xxx". Note: Does NOT update GitHub check status.

--gpu-type "A30, H100_PCIe" (OPTIONAL) : Only run the test stages on the specified GPU types. Examples: "A30, H100_PCIe". Note: Does NOT update GitHub check status.

--test-backend "pytorch, cpp" (OPTIONAL) : Skip test stages which don't match the specified backends. Only support [pytorch, cpp, tensorrt, triton]. Examples: "pytorch, cpp" (does not run test stages with tensorrt or triton backend). Note: Does NOT update GitHub pipeline status.

--only-multi-gpu-test (OPTIONAL) : Only run the multi-GPU tests. Note: Does NOT update GitHub check status.

--disable-multi-gpu-test (OPTIONAL) : Disable the multi-GPU tests. Note: Does NOT update GitHub check status.

--add-multi-gpu-test (OPTIONAL) : Force run the multi-GPU tests in addition to running L0 pre-merge pipeline.

--post-merge (OPTIONAL) : Run the L0 post-merge pipeline instead of the ordinary L0 pre-merge pipeline.

--extra-stage "H100_PCIe-TensorRT-Post-Merge-1, xxx" (OPTIONAL) : Run the ordinary L0 pre-merge pipeline and specified test stages. Examples: --extra-stage "H100_PCIe-TensorRT-Post-Merge-1, xxx".

--detailed-log (OPTIONAL) : Enable flushing out all logs to the Jenkins console. This will significantly increase the log volume and may slow down the job.

--debug (OPTIONAL) : Experimental feature. Enable access to the CI container for debugging purpose. Note: Specify exactly one stage in the stage-list parameter to access the appropriate container environment. Note: Does NOT update GitHub check status.

For guidance on mapping tests to stage names, see docs/source/reference/ci-overview.md
and the scripts/test_to_stage_mapping.py helper.

kill

kill

Kill all running builds associated with pull request.

skip

skip --comment COMMENT

Skip testing for latest commit on pull request. --comment "Reason for skipping build/test" is required. IMPORTANT NOTE: This is dangerous since lack of user care and validation can cause top of tree to break.

reuse-pipeline

reuse-pipeline

Reuse a previous pipeline to validate current commit. This action will also kill all currently running builds associated with the pull request. IMPORTANT NOTE: This is dangerous since lack of user care and validation can cause top of tree to break.

nvchenghaoz and others added 16 commits October 17, 2025 18:27
Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>
Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>
Signed-off-by: junq <22017000+QiJune@users.noreply.github.com>
Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>
Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>
Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>
Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>
Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>
… --> dtype

Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>
Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>
Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>
Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>
Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>
Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>
Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>
Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>
@coderabbitai

coderabbitai Bot commented Oct 23, 2025

Copy link
Copy Markdown
Contributor

Caution

Review failed

The pull request is closed.

📝 Walkthrough

Walkthrough

This pull request introduces fused Mixture-of-Experts (MoE) infrastructure with CUDA and Triton kernels, adds gated RMSNorm support, refactors attention pattern matching, improves KV-cache tracking, and simplifies executor APIs by removing lora_config and kv_connector_config from public signatures.

Changes

Cohort / File(s) Summary
MoE CUDA & Triton Implementation
setup.py, tensorrt_llm/_torch/auto_deploy/custom_ops/fused_moe/moe_align_kernel.cu, tensorrt_llm/_torch/auto_deploy/custom_ops/fused_moe/load_moe_align.py, tensorrt_llm/_torch/auto_deploy/custom_ops/fused_moe/triton_moe.py
Added CUDA kernel for MoE token alignment (moe_align_kernel), Python wrapper for JIT compilation and invocation (load_moe_align), and comprehensive Triton-based fused MoE operations for both unquantized and FP8-quantized paths. Included CUDA source in package_data for runtime availability.
MoE Configuration & Python Wrappers
tensorrt_llm/_torch/auto_deploy/custom_ops/fused_moe/torch_moe.py, tensorrt_llm/_torch/auto_deploy/custom_ops/fused_moe/load_moe_align.py
Added mlp_style ("gated_mlp" vs "mlp") and act_fn ("silu", "relu2") parameters to torch_moe, torch_quant_fp8_moe variants. Introduced activation resolver and support for gated MLP with W1, W2, W3 weight decomposition.
Gated RMSNorm
tensorrt_llm/_torch/auto_deploy/custom_ops/rms_norm.py, tensorrt_llm/_torch/auto_deploy/transform/library/rms_norm.py
Implemented torch_rmsnorm_gated custom op with Triton-backed _layer_norm_fwd; added FuseGatedRMSNorm transform to match and fuse NemotronH-style gated RMSNorm subgraphs.
Attention Pattern Refactoring
tensorrt_llm/_torch/auto_deploy/transform/library/attention.py
Split match_grouped_attention into MatchGroupedAttentionWithRepeatKV and MatchGroupedAttentionWithoutRepeatKV with only_repeat_kv gating; updated pattern generation and registration.
MoE Fusion & Quantization
tensorrt_llm/_torch/auto_deploy/transform/library/fused_moe.py, tensorrt_llm/_torch/auto_deploy/transform/library/quantize_moe.py
Added FuseFP8Moe transform with _stack_fp8_moe_weights function; updated quantize_moe_node to propagate mlp_style and act_fn into quantized ops.
KV Cache Management
tensorrt_llm/_torch/auto_deploy/shim/interface.py, tensorrt_llm/_torch/auto_deploy/transform/library/kvcache.py
Added current_kv_cache_size_bytes method to CachedSequenceInterface; updated cache resizing logic to use KV-only cache metrics instead of total cache.
Chunked Prefill & Context Wiring
tensorrt_llm/_torch/auto_deploy/shim/ad_executor.py, tensorrt_llm/_torch/auto_deploy/llm_args.py
Removed frozen constraint from enable_chunked_prefill field; added ContextChunkingPolicy/Config imports and wiring; truncate cache indices to active blocks; wire ctx_chunk_config into scheduler.
Executor API Simplification
tensorrt_llm/executor/{executor.py,proxy.py,ray_executor.py,ray_gpu_worker.py,rpc_proxy.py,rpc_worker.py,worker.py,base_worker.py}
Removed lora_config and kv_connector_config from public signatures across all executor classes; these now source from llm_args internally.
PyExecutor Refactoring
tensorrt_llm/_torch/pyexecutor/py_executor_creator.py, tensorrt_llm/llmapi/llm.py
Simplified create_py_executor signature; removed lora_config/kv_connector_config parameters; internal extraction from llm_args.
Configuration & Model Loading
tensorrt_llm/_torch/auto_deploy/models/hf.py
Moved torch_dtype string-to-dtype conversion from \_\_init\_\_ into _recursive_update_config; updated build_and_load_model to use "dtype": "auto".
NemotronH MOE Patching
tensorrt_llm/_torch/auto_deploy/models/patches/nemotron_h.py
Added _nemotron_h_moe_forward specialized MOE path; registered NemotronHMOE in CUSTOM_MODULE_PATCHES.
Config Updates
tensorrt_llm/_torch/auto_deploy/config/default.yaml
Replaced match_grouped_attention with two variants (with/without_repeat_kv); added fuse_fp8_moe and fuse_gated_rmsnorm transforms.
Custom Ops Module
tensorrt_llm/_torch/auto_deploy/custom_ops/__init__.py
Extended recursive module import to include nested sub-packages for op registrations.
Mxfp4 MOE Import
tensorrt_llm/_torch/auto_deploy/transform/library/mxfp4_moe.py
Updated IS_TRITON_KERNELS_AVAILABLE import path from ...custom_ops.mxfp4_moe to ...custom_ops.fused_moe.mxfp4_moe.
Test Suite Expansions
tests/integration/defs/accuracy/test_llm_api_autodeploy.py, tests/integration/test_lists/test-db/{l0_*.yml,*}, tests/unittest/_torch/auto_deploy/...
Added enable_chunked_prefill parameterization; expanded test variants; added unit tests for custom_ops (RMS norm, Triton MoE), MoE align kernel validation, NemotronH MOE patches, and ADEngine chunked prefill equivalence. Updated test configs to use "dtype" instead of "torch_dtype".

Sequence Diagram(s)

sequenceDiagram
    participant App as Application
    participant Executor as Executor API
    participant Worker as Worker
    participant LLMArgs as LLMArgs
    
    rect rgb(200, 220, 255)
    Note over Executor,Worker: Old API (Removed)
    App->>Executor: create(..., lora_config, kv_connector_config)
    Executor->>Worker: init(..., lora_config, kv_connector_config)
    Worker->>Worker: store lora_config, kv_connector_config
    end
    
    rect rgb(200, 255, 200)
    Note over Executor,LLMArgs: New API (Updated)
    App->>Executor: create(..., llm_args)
    Executor->>Worker: init(..., llm_args)
    Worker->>LLMArgs: extract lora_config, kv_connector_config from llm_args
    LLMArgs->>Worker: return configs from llm_args
    end
Loading
sequenceDiagram
    participant Input as Token Input
    participant Router as Router
    participant MoEAlign as MoE Align Kernel
    participant Triton as Triton Fused MoE
    participant Output as Output
    
    Input->>Router: topk_ids, routing_weights
    Router->>MoEAlign: topk_ids, num_experts, block_size
    MoEAlign->>MoEAlign: compute per-expert counts
    MoEAlign->>MoEAlign: prefix-sum for token placement
    MoEAlign-->>Triton: sorted_token_ids, expert_ids, num_tokens_post_pad
    Triton->>Triton: select path (mlp_style: gated_mlp vs mlp)
    alt gated_mlp: W1, W3, W2 with gate
    Triton->>Triton: act(W1*x) * W3*x
    Triton->>Triton: W2 * (gated output)
    else mlp: W_up, W_down with activation
    Triton->>Triton: act(W_up * x)
    Triton->>Triton: W_down * (activated)
    end
    Triton->>Output: fused MoE output
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

The diff introduces substantial new infrastructure (CUDA kernel, Triton kernels, gated RMSNorm) requiring careful verification of numerical correctness and performance assumptions. The MoE alignment logic, Triton fused operations, and attention pattern splitting are logic-dense. Executor API simplification is repetitive across many files but must be verified for consistency. Cross-cutting concerns (KV cache tracking, chunked prefill wiring) and multiple interconnected features (configuration handling, test parameterization) increase cognitive load despite some pattern repetition.

Possibly related PRs

Suggested labels

AutoDeploy, MoE, CUDA, Triton, Quantization, KVCache

Suggested reviewers

  • suyoggupta
✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

📜 Recent review details

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between e5865de and a8d99c3.

📒 Files selected for processing (42)
  • setup.py (1 hunks)
  • tensorrt_llm/_torch/auto_deploy/config/default.yaml (3 hunks)
  • tensorrt_llm/_torch/auto_deploy/custom_ops/__init__.py (1 hunks)
  • tensorrt_llm/_torch/auto_deploy/custom_ops/fused_moe/load_moe_align.py (1 hunks)
  • tensorrt_llm/_torch/auto_deploy/custom_ops/fused_moe/moe_align_kernel.cu (1 hunks)
  • tensorrt_llm/_torch/auto_deploy/custom_ops/fused_moe/torch_moe.py (6 hunks)
  • tensorrt_llm/_torch/auto_deploy/custom_ops/fused_moe/triton_moe.py (1 hunks)
  • tensorrt_llm/_torch/auto_deploy/custom_ops/rms_norm.py (2 hunks)
  • tensorrt_llm/_torch/auto_deploy/llm_args.py (1 hunks)
  • tensorrt_llm/_torch/auto_deploy/models/hf.py (2 hunks)
  • tensorrt_llm/_torch/auto_deploy/models/patches/nemotron_h.py (3 hunks)
  • tensorrt_llm/_torch/auto_deploy/shim/ad_executor.py (3 hunks)
  • tensorrt_llm/_torch/auto_deploy/shim/interface.py (1 hunks)
  • tensorrt_llm/_torch/auto_deploy/transform/library/attention.py (5 hunks)
  • tensorrt_llm/_torch/auto_deploy/transform/library/fused_moe.py (3 hunks)
  • tensorrt_llm/_torch/auto_deploy/transform/library/kvcache.py (2 hunks)
  • tensorrt_llm/_torch/auto_deploy/transform/library/mxfp4_moe.py (1 hunks)
  • tensorrt_llm/_torch/auto_deploy/transform/library/quantize_moe.py (1 hunks)
  • tensorrt_llm/_torch/auto_deploy/transform/library/rms_norm.py (2 hunks)
  • tensorrt_llm/_torch/pyexecutor/py_executor_creator.py (2 hunks)
  • tensorrt_llm/executor/base_worker.py (2 hunks)
  • tensorrt_llm/executor/executor.py (8 hunks)
  • tensorrt_llm/executor/proxy.py (1 hunks)
  • tensorrt_llm/executor/ray_executor.py (2 hunks)
  • tensorrt_llm/executor/ray_gpu_worker.py (1 hunks)
  • tensorrt_llm/executor/rpc_proxy.py (0 hunks)
  • tensorrt_llm/executor/rpc_worker.py (1 hunks)
  • tensorrt_llm/executor/worker.py (1 hunks)
  • tensorrt_llm/llmapi/llm.py (1 hunks)
  • tests/integration/defs/accuracy/test_llm_api_autodeploy.py (6 hunks)
  • tests/integration/test_lists/test-db/l0_b200.yml (1 hunks)
  • tests/integration/test_lists/test-db/l0_dgx_h100.yml (1 hunks)
  • tests/integration/test_lists/test-db/l0_dgx_h200.yml (1 hunks)
  • tests/integration/test_lists/test-db/l0_h100.yml (1 hunks)
  • tests/unittest/_torch/auto_deploy/_utils_test/_model_test_utils.py (2 hunks)
  • tests/unittest/_torch/auto_deploy/unit/singlegpu/custom_ops/test_mamba_rms_norm.py (1 hunks)
  • tests/unittest/_torch/auto_deploy/unit/singlegpu/custom_ops/triton_kernels/test_triton_moe.py (1 hunks)
  • tests/unittest/_torch/auto_deploy/unit/singlegpu/models/test_hybrid_patches.py (1 hunks)
  • tests/unittest/_torch/auto_deploy/unit/singlegpu/models/test_nemotron_h_patches.py (1 hunks)
  • tests/unittest/_torch/auto_deploy/unit/singlegpu/shim/test_engine.py (2 hunks)
  • tests/unittest/_torch/auto_deploy/unit/singlegpu/transformations/library/test_attention_matcher.py (1 hunks)
  • tests/unittest/_torch/auto_deploy/unit/singlegpu/transformations/library/test_attention_matcher_hf.py (1 hunks)

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants