Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
119 commits
Select commit Hold shift + click to select a range
316430f
[None] [waive] Waive the failed step3p7 test case due to ckpt update …
kaiyux Jun 5, 2026
d5de55e
[https://nvbugs/6210714][fix] Fix mamba block calculation (#14524)
VALLIS-NERIA Jun 5, 2026
6818233
[https://nvbugs/5546507][https://nvbugs/5612313][test] Remove obsolet…
xinhe-nv Jun 5, 2026
fdcdcb3
[None][fix] AutoDeploy: Move hf_id_to_local_model_dir to function for…
bmarimuthu-nv Jun 5, 2026
6387eac
[TRTLLM-12893][infra] Parallelize post stages: Rerun Report, Test Cov…
ZhanruiSunCh Jun 5, 2026
2336e47
[None][infra] Waive 11 failed cases for main in post-merge 2760 (#15003)
ZhanruiSunCh Jun 5, 2026
58da60a
[None][fix] Uncomment Qwen3.5 and DSR1 from model registry so that th…
taylor-yb-lee Jun 5, 2026
73c824c
[TRTLLM-11410][feat] Cosmos3 Support (#14824)
NVShreyas Jun 5, 2026
fb5bd44
[https://nvbugs/5859886][fix] Remove the waiver (#14948)
ziyixiong-nv Jun 5, 2026
501b5c2
[https://nvbugs/6248744][fix] Added `trust_remote_code=True` to the `…
tensorrt-cicd Jun 5, 2026
37ece3f
[https://nvbugs/6160629][fix] AutoDeploy: Fix manual seed setting for…
galagam Jun 5, 2026
86f9602
[TRTLLM-12714][feat] KVCacheManagerV2: wire PyExecutor rebalance hook…
thorjohnsen Jun 5, 2026
3e17560
[None][feat] add Wan I2V generation example (#14981)
o-stoner Jun 5, 2026
52ba2bb
[TRTLLM-12527][feat] Parallelize multi-shard visual-gen checkpoint lo…
yibinl-nvidia Jun 5, 2026
d639c57
[https://nvbugs/6250866][fix] Fix deep ep partial warp sync for gptos…
dongfengy Jun 6, 2026
3b21093
[https://nvbugs/6272668][infra] Unwaive DSR1 and Qwen3.5 again (#15010)
taylor-yb-lee Jun 6, 2026
d7a5872
[None][feat] Afmoe trinity support (#13148)
alyosha-swamy Jun 6, 2026
4279e5b
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 6, 2026
e47f26e
[TRTLLM-13027][ci] Relocate under-using tests to right-sized stages (…
QiJune Jun 6, 2026
520262d
[None][feat] Add LTX-2 visual generation example (#14976)
yibinl-nvidia Jun 6, 2026
ec6b284
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 7, 2026
bedad85
[None][feat] AutoDeploy: Fix hardcoded configs (#14943)
taylor-yb-lee Jun 7, 2026
47666de
[#13718][feat] AutoDeploy MoE all-to-all: cache + runtime max-tokens …
greg-kwasniewski1 Jun 7, 2026
428cc3e
[TRTLLM-13177][doc] Add Nemotron 3 Ultra doc (#14964)
nv-guomingz Jun 7, 2026
dcd4e90
[#10710][feat] Make explicit CLI flags take precedence over --config …
marinayanov Jun 7, 2026
8be182d
[https://nvbugs/6260907][fix] unwaive test (#15058)
bo-nv Jun 8, 2026
b8d17d7
[None][chore] Increase GB200-4_GPUs-PyTorch shards (#14836)
tburt-nv Jun 8, 2026
71debd5
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 8, 2026
0e0ee27
[TRTLLM-12648][test] implement disagg cancellation canary thread (#15…
chienchunhung Jun 8, 2026
98a88f7
[TRTLLM-12507][feat] Cudagraph support for routed-expert MoE LoRA wit…
brb-nv Jun 8, 2026
86f33e6
[https://nvbugs/6245317][test] set Harmony tiktoken env for GPT-OSS d…
dongfengy Jun 8, 2026
b4d44d3
[https://nvbugs/6153955][test] unwaive GPT-OSS w4 DP4 CUTLASS (#14884)
dongfengy Jun 8, 2026
ca2bc5e
[None][perf] kv_cache_manager_v2: batch block-key SHA-256 hashing (#1…
lancelly Jun 8, 2026
2cad6db
[TRTLLM-13259][ci] Merge DGX_H100 DeepSeek and GptOss stages (#15035)
QiJune Jun 8, 2026
5fa68a4
[None][infra] Waive 11 failed cases for main in post-merge 2765 (#15080)
ZhanruiSunCh Jun 8, 2026
2632530
[None][infra] Waive 3 failed cases for main in post-merge 2765 (#15082)
ZhanruiSunCh Jun 8, 2026
7e49baa
[None][test] waive weekly qa ci failure cases (#15077)
crazydemo Jun 8, 2026
02f6b2f
[None][feat] AutoDeploy: propagate layer_type hint across pattern-mat…
greg-kwasniewski1 Jun 8, 2026
28dc25e
[None][test] Waive 15 failed cases for main in QA CI (#15056)
tensorrt-cicd Jun 8, 2026
2febb37
[None][infra] Waive 1 failed cases for main in pre-merge 41894 (#15089)
ZhanruiSunCh Jun 8, 2026
09c21b6
[TRTLLM-13262][ci] Move non-default-feature tests to post merge (#15038)
QiJune Jun 8, 2026
c93c63d
[None][feat] Enable disk cache config for KV cache v2 (#14845)
reasonsolo Jun 8, 2026
6dee167
[https://nvbugs/6185446][fix] Add warmup for trtllm-gen fmha JIT kern…
pengbowang-nv Jun 8, 2026
b14794c
[https://nvbugs/6162940][chore] Unwaive fixed test (#15078)
longlee0622 Jun 8, 2026
2bf4d3d
[None][perf] Support Gemma RMSNorm + interleaved mRoPE in fused_qk_no…
nv-guomingz Jun 8, 2026
9eaa468
[None][test] Half K25 Agg Multi Round to Solve Timeout Issue (#15083)
chenfeiz0326 Jun 8, 2026
9af8a16
[None][infra] Reduce Docker image layer count in release stage (#14972)
tburt-nv Jun 8, 2026
cb01607
[#14828][feat] AutoDeploy: support multi KV cache memory pool in trtl…
MrGeva Jun 8, 2026
15d06c0
[None][doc] Refine Nemotron Ultra doc (#15113)
nv-guomingz Jun 8, 2026
8036cde
[None][infra] Waive TestQwen3NextInstruct nvfp4 cases (#15086)
mzweilz Jun 8, 2026
1998324
[https://nvbugs/6248757][fix] Avoid running all reduce in aux stream …
tensorrt-cicd Jun 8, 2026
900d069
[https://nvbugs/6221483][fix] AutoDeploy: Fix Eagle metadata host syn…
govind-ramnarayan Jun 8, 2026
9827c21
[None][feat] add FLUX visual generation examples (#14987)
karljang Jun 8, 2026
b222246
[https://nvbugs/6261164][fix] In the kvcache insert transform (`_Inse…
tensorrt-cicd Jun 8, 2026
c1e9b00
[https://nvbugs/6211189][fix] Lower the reference to 46.5 (matching c…
tensorrt-cicd Jun 9, 2026
bfb4537
[None][refactor] split VisualGen pipeline and model configs (#14956)
bobboli Jun 9, 2026
5e3af40
[TRTLLM-11457][feat] Async Ulysses pipeline (Enabled for LTX-2 + WAN)…
luyiyun1021 Jun 9, 2026
09ebc59
[TRTLLM-11548][doc] Add Qwen3.5 deployment guide doc (#15111)
nv-guomingz Jun 9, 2026
a33dec7
[https://nvbugs/6181383][fix] Build inner text/vision/audio sub-confi…
tensorrt-cicd Jun 9, 2026
041ed83
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 9, 2026
2490441
[https://nvbugs/6273850][chore] waive TestQwen3_5_4B::test_bf16 for a…
tburt-nv Jun 9, 2026
64497e2
[None][doc] Add docs for AutoDeploy transforms (#15122)
bmarimuthu-nv Jun 9, 2026
9349fcc
[None][infra] Waive 4 failed cases for main in post-merge 2769 (#15140)
ZhanruiSunCh Jun 9, 2026
28845dd
[https://nvbugs/6227203][fix] Remove redundant TikTokenTokenizer shim…
tianyuxbear Jun 9, 2026
a197a5e
[None][fix] tunable_fp4_quantize: rename misnamed kwarg + add real SF…
luyiyun1021 Jun 9, 2026
e9402ab
[None][test] Fix gen_only missing prev_device_step_time race in perf …
tensorrt-cicd Jun 9, 2026
6bf3e49
[None][test] Fix disagg test result dir (#14864)
fredricz-20070104 Jun 9, 2026
a7e4a9b
[TRTLLM-13332][test] Remove TestLlama4ScoutInstruct tests (#15144)
QiJune Jun 9, 2026
6f7aea5
[https://nvbugs/6266705][fix] Gate FlashInfer GDN kernels to supporte…
nv-guomingz Jun 9, 2026
6254f3a
[https://nvbugs/6255037][fix] Count DSA indexer K-cache correctly as …
eopXD Jun 9, 2026
a90fd15
[https://nvbugs/6194812][test] Update llm_perf_core.yml to require a …
yufeiwu-nv Jun 9, 2026
34a94ee
[TRTLLMINF-112][infra] Reduce the waiting time between check node is …
EmmaQiaoCh Jun 9, 2026
b852703
[None][infra] Waive 1 failed cases for main in pre-merge 41821 (#15135)
ZhanruiSunCh Jun 9, 2026
178f4e6
[None][infra] CBTS Layer 3: pass test-db via Artifactory instead of e…
crazydemo Jun 9, 2026
45e2523
[TRTLLM-13264][feat] Add native bias epilogue to NVFP4 GEMM (#15053)
luyiyun1021 Jun 9, 2026
2ee96cf
[https://nvbugs/6278380][unwaive] unwaive ad cases (#15148)
crazydemo Jun 9, 2026
104b9d7
[https://nvbugs/6244474][fix] AutoDeploy: Remove llama perf test from…
MrGeva Jun 9, 2026
ba6ba1f
[https://nvbugs/6212252][fix] Select CUTLASS MoE backend on non-Black…
xxi-nv Jun 9, 2026
d620851
[TRTLLM-13302][feat] Register NVIDIA Wan2.2-T2V quantized checkpoints…
zhenhuaw-me Jun 9, 2026
484e6c9
[None][chore] add VisualGen team as the codeowner of the VisualGen At…
zhenhuaw-me Jun 9, 2026
487330e
[None][feat] Default on FlashInferTrtllmGenAttention (#14618)
yihwang-nv Jun 9, 2026
58fbfb9
[None][infra] Test DFW with BSL branch (#14597)
yuanjingx87 Jun 9, 2026
451dbb8
[TRTLLM-12214][perf] customMoeRoutingKernel: lower BLOCK_SIZE to 128,…
xwang233 Jun 9, 2026
f0ba8c7
[TRTLLM-12214][perf] DeepGemmFusedMoE: skip redundant data expand via…
xwang233 Jun 9, 2026
736dc22
[TRTLLM-12648][test] implement disagg cancellation load thread (#15124)
chienchunhung Jun 9, 2026
f1d39ea
[None][fix] Fix regression from SageAttention kernel: Use static sche…
xrq-phys Jun 9, 2026
680c6c4
[TRTLLM-12467][feat] EPD improvements (#13864)
venkywonka Jun 9, 2026
0edbbfe
[None][feat] Expose stored block-hash chain to KV cache connector (#1…
jthomson04 Jun 9, 2026
358505c
[#12805][fix] Fall back to local cache when loading tokenizer for gat…
1MrazorT1 Jun 9, 2026
3ddef66
[None][feat] Support partial RoPE fusion for Hopper kernels in XQA fo…
DomBrown Jun 9, 2026
b0216c6
[None][infra] Add nv-xtf, rahul-steiger-nv, tedzhouhk, tensorrt-cicd …
ZhanruiSunCh Jun 9, 2026
98393f3
[None][feat] Add Prometheus metrics for prompt cache, speculative dec…
vedularaghu Jun 9, 2026
f39a79c
[None][chore] Unwaive DSV32 helix tests (#14871)
brb-nv Jun 9, 2026
e0a909a
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 9, 2026
ddef2d0
[None][fix] unset UCX_TLS=tcp (#15008)
tburt-nv Jun 9, 2026
884520c
[None][feat] Port 13 AutoDeploy custom models to sharding IR + opt th…
greg-kwasniewski1 Jun 9, 2026
48d2b89
[None][chore] Make image paths absolute in blog22 (#15177)
brb-nv Jun 9, 2026
0f7e1db
Fix PyExecutor FPM iteration timing (#14922)
tedzhouhk Jun 9, 2026
ffcd8e6
[#13816][feat] AutoDeploy: Optimize gpt-oss-120b perf (#14202)
taylor-yb-lee Jun 9, 2026
9a7f76f
[None][fix] Register Multimodal Placeholders for Qwen3.5 MoE VLM Serv…
anurags25 Jun 9, 2026
edfc667
[None][feat] Weight trtllm-bench AR/AL averages by output length (#14…
zhaoyangwang-nvidia Jun 10, 2026
3b945f7
[TRTLLM-13052][feat] Enable TRTLLM moe backend for nemotron-h BF16 ck…
Wanli-Jiang Jun 10, 2026
8e40515
[None][fix] Fix and unwaive nemotron related bugs (#15085)
Wanli-Jiang Jun 10, 2026
9c6cb35
[https://nvbugs/6140226][test] Add DFlash coverage for Qwen3.5 MoE va…
yingguo-trt Jun 10, 2026
2763557
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 10, 2026
9c100bb
[None][test] temporarily waive Cosmos3 B200 failures (#15195)
bobboli Jun 10, 2026
5741389
[NVBUG-6241842][fix] DSA DSL atom-split: guard against MTP draft next…
limin2021 Jun 10, 2026
2148a3e
[#11423][feat] AutoDeploy: Basic Disagg Support (#14057)
govind-ramnarayan Jun 10, 2026
9bc4321
[https://nvbugs/6280060][fix] Scope disagg-ctx cache-transfer quorum …
tensorrt-cicd Jun 10, 2026
27b52b3
[None][test] Add e2e example tests for flux1/2, ltx2, wan_i2v, and co…
chang-l Jun 10, 2026
2878b30
[#12632][feat] Add pipeline cache support for AutoDeploy (#13729)
nvchenghaoz Jun 10, 2026
31e730a
[None][test] Add support for nemotron_3_ultra_550b_nvfp4 model in per…
yufeiwu-nv Jun 10, 2026
b206f68
[None][feat] Indexer TopK: single-block / multi-pass radix (#14268)
dcampora Jun 10, 2026
3b46728
[None][fix] Clear workspace in run_mla_generation to avoid potential …
yihwang-nv Jun 10, 2026
90cb7ff
[None][chore] Unwaive AutoDeploy accuracy tests (#14971)
bmarimuthu-nv Jun 10, 2026
2d196f7
[None][test] Increase kv_transfer_timeout_ms for b200 deepseek-r1 dis…
tensorrt-cicd Jun 10, 2026
62c6521
[None][feat] Enable MTP for Step-3.7 NVFP4 and port Step-3.7VL vision…
kaiyux Jun 10, 2026
49837b5
[None][fix] Revert "Merge Eagle3 and MTP-eagle one-model workers (#12…
chenfeiz0326 Jun 10, 2026
628ff8a
[None][fix] Unwaive deepseek-v32 perf-sanity cases after #12353 revert
chenfeiz0326 Jun 10, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
48 changes: 19 additions & 29 deletions .claude/skills/ad-sharding-ir-port/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,8 +35,9 @@ You MAY introduce ONLY the following changes:
- `torch.split(...)` / `torch.split_with_sizes(...)` → `torch.ops.auto_deploy.split_with_sizes(...)`
- **A2. Sharding-hint kwargs added** to call sites of: `torch_moe`, `torch_ssm`, `torch_gated_delta_rule`, `torch_causal_conv1d`, `torch_rmsnorm_gated`, `torch_mla`, `torch_linear_simple`, `auto_deploy.split_with_sizes`, `auto_deploy.view`. Allowed kwargs: `tp_mode`, `layer_type`, `output_sizes`, `tp_min_local_shape`, `tp_scaled_dim`, `shardable`, `enable_sharding`.
- **A3. Inserting `torch.ops.auto_deploy.all_reduce(..., layer_type=...)`** after rowwise projections / at MoE merge points (single all_reduce after routed + shared sums).
- **A4. Adding `import tensorrt_llm._torch.auto_deploy.custom_ops # noqa: F401`** side-effect import at the top if not already present.
- **A5. The module docstring update** describing the sharding strategy.
- **A4. Docstring updates:**
- Module-level: a single-line header noting the file uses sharding IR, followed by the existing source-of-truth / HF link block. Example: `"""Llama 3 model (sharding IR)."""`.
- Per-class (MLP, Attention, MoE block, etc.): a short `Sharding strategy:` block listing what each projection maps to (`colwise` / `rowwise` / `all_reduce` / `tp_scaled_dim`).

**FORBIDDEN (everything else, including but not limited to):**

Expand Down Expand Up @@ -71,45 +72,35 @@ Before editing, ensure the file is committed so you can diff against the origina
git stash # or commit — ensure a clean baseline to diff against
```

### Step 2: Add the custom_ops side-effect import

If not already present at the top of the file:

```python
import tensorrt_llm._torch.auto_deploy.custom_ops # noqa: F401 -- register all ops
```

Do **not** add global `SHARD_*` flags. Layer-level control uses the `layer_type` hint on each op and `shard_layers` in YAML.

### Step 3: Replace linear projections
### Step 2: Replace linear projections

For every `self.proj(x)` or `nn.Linear` call, use `torch.ops.auto_deploy.torch_linear_simple` with explicit `tp_mode` and `layer_type`. Always set `tp_mode` unconditionally (no `if _s else "none"`). **Rules:** opening projections (Q/K/V/gate/up/in_proj) → `"colwise"`; closing (O/down/out_proj) → `"rowwise"`; tiny outputs (e.g. `shared_expert_gate` dim 1) → `"none"`; MLA latent projections (q_a, kv_a) → `"none"`. For fused weights split later, pass `output_sizes=[...]`. For GQA, use `tp_min_local_shape=self.head_dim` on K/V colwise lines.

### Step 4: Replace split / chunk after fused colwise projections
### Step 3: Replace split / chunk after fused colwise projections

Use `torch.ops.auto_deploy.split_with_sizes` with `shardable` / `layer_type` where sizes scale with TP.

### Step 5: Replace view / reshape with concrete head counts
### Step 4: Replace view / reshape with concrete head counts

During `torch.export`, `-1` becomes concrete; after TP, wrong values break. Any reshape whose dimension is a head count that scales with TP must use `torch.ops.auto_deploy.view` with `tp_scaled_dim` set appropriately. Safe cases: flat-to-2D, or `[B,S,-1]` when the input is already correctly sharded.

### Step 6: Insert `all_reduce`
### Step 5: Insert `all_reduce`

After every rowwise projection, add `torch.ops.auto_deploy.all_reduce(..., layer_type=...)`. **Parallel branch rule:** when branches merge by addition, use a **single** `all_reduce` after the sum (e.g. MoE routed + shared expert; parallel attention + MLP residual branches).

### Step 7: Special ops (Conv1d, SSM, GatedDeltaNet, gated RMSNorm)
### Step 6: Special ops (Conv1d, SSM, GatedDeltaNet, gated RMSNorm)

Add sharding hints on `torch_causal_conv1d`, `torch_ssm`, `torch_gated_delta_rule`, `torch_rmsnorm_gated` per docstrings—typically `shardable` / `output_sizes` / `tp_mode` as required.

### Step 8: MoE
### Step 7: MoE

Pass `layer_type="moe"` into `torch_moe`; `apply_sharding_hints` handles EP/TP.

### Step 9: Verify registration
### Step 8: Verify registration

The model's existing registration (`AutoModelForCausalLMFactory.register_custom_model_cls` at the bottom of the file and its import in `__init__.py`) stays unchanged. No new registration is needed — sharding hints do not change the model identity.

### Step 10: YAML — enable hint-driven sharding
### Step 9: YAML — enable hint-driven sharding

Add `enable_sharder_ir.yaml` to the model's `yaml_extra` list in `examples/auto_deploy/model_registry/models.yaml` (if not already present). This composable fragment disables legacy sharding passes and enables `apply_sharding_hints`. Registry fragments are deep-merged in `yaml_extra` order (see `DynamicYamlMixInForSettings` in `tensorrt_llm/_torch/auto_deploy/utils/_config.py`).

Expand All @@ -136,11 +127,11 @@ transforms:
enabled: true
```

Set `world_size` once, to the **maximum number of GPUs available on the machine**, auto-detected with `python -c 'import torch; print(torch.cuda.device_count())'` (or `nvidia-smi --list-gpus | wc -l`). Do **not** hardcode `world_size: 8` (or any other literal) — porting agents run on heterogeneous hardware and an 8-GPU literal will simply fail to launch on a 2- or 4-GPU machine. If the model's `num_attention_heads` (and, for GQA, `num_key_value_heads`) does not divide the detected GPU count, fall back to the largest power-of-two divisor that does (e.g. 4 on an 8-GPU machine if `num_attention_heads = 12`). Run the end-to-end command exactly once at that size — there is no value in repeating it at multiple smaller sizes, because the offline sharding equivalence test (Step 11b) already exercises 2- and 4-GPU dist configs cheaply.
Set `world_size` once, to the **maximum number of GPUs available on the machine**, auto-detected with `python -c 'import torch; print(torch.cuda.device_count())'` (or `nvidia-smi --list-gpus | wc -l`). Do **not** hardcode `world_size: 8` (or any other literal) — porting agents run on heterogeneous hardware and an 8-GPU literal will simply fail to launch on a 2- or 4-GPU machine. If the model's `num_attention_heads` (and, for GQA, `num_key_value_heads`) does not divide the detected GPU count, fall back to the largest power-of-two divisor that does (e.g. 4 on an 8-GPU machine if `num_attention_heads = 12`). Run the end-to-end command exactly once at that size — there is no value in repeating it at multiple smaller sizes, because the offline sharding equivalence test (Step 10b) already exercises 2- and 4-GPU dist configs cheaply.

Optional `shard_layers` limits which `layer_type` hints are processed; unset means shard all shardable nodes.

### Step 11a — End-to-end run
### Step 10a — End-to-end run

Do not report success until a run completes successfully.

Expand All @@ -151,7 +142,7 @@ Do not report success until a run completes successfully.

**Layer type strings** (for `layer_type` / `shard_layers`): use `"mha"`, `"mla"`, `"mlp"`, `"moe"`, `"ssm"`, `"delta"`, or `"unknown"` (default; skipped when `shard_layers` is set). Match the conventions used in `apply_sharding_hints` and project enums.

### Step 11b — Sharding equivalence test (MANDATORY)
### Step 10b — Sharding equivalence test (MANDATORY)

Run the offline sharding-IR equivalence test ([`tests/unittest/auto_deploy/multigpu/transformations/library/test_sharding_ir_equivalence.py`](tests/unittest/auto_deploy/multigpu/transformations/library/test_sharding_ir_equivalence.py)) against the modeling file you just edited, under **every** parallelism configuration the test exposes. The port is **not** complete until every configuration passes. Skipping this step or treating a partial pass (e.g. only `tep`) as success is not allowed.

Expand Down Expand Up @@ -191,10 +182,10 @@ done
**Failure handling:**

- A cell failing with `KeyError`, `AttributeError`, `ValueError: You must specify exactly one of input_ids or inputs_embeds`, or any exception *before* `[sharding-ir-eq]` prints means the **modeling code itself** does not yet build / export on a tiny config — fix the modeling code (within the Step 0 allowlist) before proceeding. Do not silently skip the cell.
- A cell where `[sharding-ir-eq]` prints `rel_rmse >= tol` (from the same log line) means a **sharding-hint bug**: a missing `all_reduce`, a wrong `tp_mode`, a `view` without `tp_scaled_dim`, a `split_with_sizes` whose sizes do not scale, etc. Re-read Step 6 (all_reduce), Step 3 (tp_mode), Step 5 (view), Step 4 (split_with_sizes) and the layer-specific patterns. Iterate on the hints until clean. If the failure is small (rel_rmse just slightly above tol) and you have reason to believe it is real numerical noise from the specific layer mix of this model rather than a sharding-hint bug, raise it with the parent agent rather than silently bumping `SHARDING_IR_REL_RMSE_TOL`.
- A cell where `[sharding-ir-eq]` prints `rel_rmse >= tol` (from the same log line) means a **sharding-hint bug**: a missing `all_reduce`, a wrong `tp_mode`, a `view` without `tp_scaled_dim`, a `split_with_sizes` whose sizes do not scale, etc. Re-read Step 5 (all_reduce), Step 2 (tp_mode), Step 4 (view), Step 3 (split_with_sizes) and the layer-specific patterns. Iterate on the hints until clean. If the failure is small (rel_rmse just slightly above tol) and you have reason to believe it is real numerical noise from the specific layer mix of this model rather than a sharding-hint bug, raise it with the parent agent rather than silently bumping `SHARDING_IR_REL_RMSE_TOL`.
- A cell that the modeling file legitimately does not support (e.g. `ep-only` on a dense model with no MoE) is acceptable only if the failure is a documented `pytest.skip(...)` from the test infrastructure. A silent `FAIL` is **not** acceptable.

### Step 12 — Pre-finalization self-audit (MANDATORY)
### Step 11 — Pre-finalization self-audit (MANDATORY)

Before reporting the file as done, you MUST diff your changes against the git baseline:

Expand All @@ -209,8 +200,7 @@ Then classify every hunk into one of the following categories (defined in Step 0
| **A1** | yes | Op substitution (`linear` / `view` / `split`) |
| **A2** | yes | Sharding-hint kwarg added (`tp_mode`, `layer_type`, `output_sizes`, `tp_min_local_shape`, `tp_scaled_dim`, `shardable`, `enable_sharding`) |
| **A3** | yes | `auto_deploy.all_reduce` insertion |
| **A4** | yes | `custom_ops` side-effect import added |
| **A5** | yes | Module docstring update describing sharding strategy |
| **A4** | yes | Docstring updates: one-line module header + per-class `Sharding strategy:` blocks |
| **F1** | NO | `torch.ops.trtllm.*` replaced with vanilla PyTorch |
| **F2** | NO | Input contract change (asserts, fallbacks added/removed) |
| **F3** | NO | Module hierarchy / parameter / buffer / load-hook change |
Expand Down Expand Up @@ -261,8 +251,8 @@ You are NOT done until every row in the table is a yes-allowed category.

## Validation checklist (human review)

- All four configurations of the **sharding equivalence test** (Step 11b) pass with the parsed `rel_rmse` strictly below the parsed `tol` from the same rank-0 log line. Report the per-cell `rel_rmse` and `tol` pair.
- All four configurations of the **sharding equivalence test** (Step 10b) pass with the parsed `rel_rmse` strictly below the parsed `tol` from the same rank-0 log line. Report the per-cell `rel_rmse` and `tol` pair.
- `world_size=1`: unsharded path; hints should not break correctness.
- `world_size=<max-available>`: end-to-end run (Step 11a) at the maximum GPU count auto-detected on the machine (head-divisibility permitting; see Step 11).
- `world_size=<max-available>`: end-to-end run (Step 10a) at the maximum GPU count auto-detected on the machine (head-divisibility permitting; see Step 10).
- `apply_sharding_hints` node count vs expectation.
- Optional: `shard_layers: ['moe']` to verify selective sharding.
8 changes: 4 additions & 4 deletions .claude/skills/trtllm-model-onboard-multimodal/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -83,7 +83,7 @@ metadata:
When `@support_multimodal_disaggregated` is set and the deployment uses `TLLM_MULTIMODAL_DISAGGREGATED=1`:

- **Encoder worker:** runs as a standalone `MultimodalEncoder` (`mm_encoder_only=True`). It executes only the multimodal encoder and ships `mm_embeddings` (+ mRoPE position ids/deltas) to prefill+decode workers as shared-tensor handles.
- **Prefill+decode worker:** the model's `__init__` skips constructing `self.mm_encoder` when `_is_disagg()` is true; the input processor's `_attach_multimodal_embeddings_impl()` override binds the encoder handles into the request (the base `attach_multimodal_embeddings` wrapper detokenizes tokenized inputs for non-fast-path VLMs, then delegates to your impl). For context-only requests, the engine re-clones mrope tensors so IPC handles outlive the encoder worker's freed memory — replicate that pattern for any new GPU-resident mm tensors.
- **Prefill+decode worker:** the model's `__init__` skips constructing `self.mm_encoder` when `_is_mm_disagg()` is true; the input processor's `attach_multimodal_embeddings()` override binds the encoder handles into the request. For context-only requests, the engine re-clones mrope tensors so IPC handles outlive the encoder worker's freed memory — replicate that pattern for any new GPU-resident mm tensors.

### Templates to study

Expand Down Expand Up @@ -229,7 +229,7 @@ class {Name}Model(PreTrainedModel):
if hasattr(self, "llm"):
return # idempotency guard — re-entry from `post_config` etc.

if not _is_disagg():
if not _is_mm_disagg():
self.mm_encoder = {Name}VisionModel(model_config)
else:
self.mm_encoder = None
Expand Down Expand Up @@ -269,7 +269,7 @@ class {Name}Model(PreTrainedModel):

multimodal_params = kwargs.get("multimodal_params", [])
mm_embeds = []
if len(multimodal_params) > 0 and not _is_disagg():
if len(multimodal_params) > 0 and not _is_mm_disagg():
mm_embeds = get_multimodal_embeddings(
encoder_forward_fn=self.mm_encoder.forward,
multimodal_params=multimodal_params[:num_context_requests],
Expand Down Expand Up @@ -341,7 +341,7 @@ class {Name}Model(PreTrainedModel): ...

```python
def load_weights(self, weights, weight_mapper):
if not _is_disagg():
if not _is_mm_disagg():
self.mm_encoder.load_weights(weights)
# Release mmap pages backing the encoder weights as soon as we're done.
if hasattr(weights, "mark_consumed"):
Expand Down
4 changes: 2 additions & 2 deletions .github/CODEOWNERS
Original file line number Diff line number Diff line change
Expand Up @@ -45,8 +45,8 @@

## TensorRT-LLM Pytorch - VisualGen
/tensorrt_llm/_torch/visual_gen @NVIDIA/trt-llm-torch-visual-gen-devs
/tensorrt_llm/_torch/visual_gen/attention_backend @NVIDIA/trt-llm-torch-attention-devs
/tensorrt_llm/_torch/visual_gen/modules/attention.py @NVIDIA/trt-llm-torch-attention-devs
/tensorrt_llm/_torch/visual_gen/attention_backend @NVIDIA/trt-llm-torch-attention-devs @NVIDIA/trt-llm-torch-visual-gen-devs
/tensorrt_llm/_torch/visual_gen/modules/attention.py @NVIDIA/trt-llm-torch-attention-devs @NVIDIA/trt-llm-torch-visual-gen-devs
/tensorrt_llm/visual_gen @NVIDIA/trt-llm-llmapi-devs
/tests/integration/defs/examples/test_visual_gen.py @NVIDIA/trt-llm-torch-visual-gen-devs
/tests/integration/defs/visual_gen @NVIDIA/trt-llm-torch-visual-gen-devs
Expand Down
4 changes: 4 additions & 0 deletions .github/workflows/blossom-ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,7 @@ jobs:
"arekay",
"arysef",
"aswinvisva",
"athena-nv",
"atrifex",
"Autumn1998",
"baize97",
Expand Down Expand Up @@ -234,6 +235,7 @@ jobs:
"nv-anants",
"nv-guomingz",
"nv-lschneider",
"nv-xtf",
"nv-yilinf",
"nv-yna",
"nvamyt",
Expand Down Expand Up @@ -267,6 +269,7 @@ jobs:
"qsang-nv",
"raayandhar",
"rabiel",
"rahul-steiger-nv",
"rakib-hasan",
"RayenTian",
"raymochen",
Expand Down Expand Up @@ -310,6 +313,7 @@ jobs:
"taylor-yb-lee",
"tburt-nv",
"tcherckez-nvidia",
"tedzhouhk",
"tfogal",
"thorjohnsen",
"tianyuxbear",
Expand Down
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -65,6 +65,9 @@ tensorrt_llm/scripts
docs/source/**/*.rst
!docs/source/examples/index.rst
!docs/source/_includes/note_sections.rst
!docs/source/features/auto_deploy/transforms.rst
!docs/source/features/auto_deploy/transforms/
!docs/source/features/auto_deploy/transforms/*.rst
*.swp
.nfs*

Expand Down
3 changes: 2 additions & 1 deletion 3rdparty/fetch_content.json
Original file line number Diff line number Diff line change
Expand Up @@ -26,7 +26,8 @@
"display_name": "deep_ep",
"git_repository": "${github_base_url}/deepseek-ai/DeepEP",
"git_tag": "5be51b228a7c82dbdb213ea58e77bffd12b38af8",
"use_url": true
"use_url": true,
"patch_file": "patches/deep_ep_intranode_combine_fix.patch"
},
{
"name": "deepgemm",
Expand Down
35 changes: 35 additions & 0 deletions 3rdparty/patches/deep_ep_intranode_combine_fix.patch
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
--- a/csrc/kernels/intranode.cu
+++ b/csrc/kernels/intranode.cu
@@ -844,9 +844,15 @@

#ifndef DISABLE_SM90_FEATURES
// Wait TMA arrival
+ // hidden_int4 is not always divisible by a warp. The final tile can have
+ // only a subset of lanes active, so synchronize only participating lanes.
+ auto const tile_start = i - lane_id;
+ auto const active_lanes = min(32, hidden_int4 - tile_start);
+ auto const sync_mask = active_lanes == 32 ? 0xffffffffu : ((1u << active_lanes) - 1u);
+
if (lane_id == 0)
tma_store_wait<kNumStages - 1>();
- __syncwarp();
+ __syncwarp(sync_mask);

// Write into TMA buffer
auto tma_stage_idx = (i / 32) % kNumStages;
@@ -854,13 +860,13 @@

// Issue TMA
tma_store_fence();
- __syncwarp();
+ __syncwarp(sync_mask);
if (lane_id == 0) {
auto tma_bytes = min(32, hidden_int4 - i) * static_cast<int>(sizeof(int4));
tma_store_1d(reinterpret_cast<int4*>(tma_buffer) + tma_stage_idx * 32,
recv_int4 + token_idx * hidden_int4 + i, tma_bytes, false);
}
- __syncwarp();
+ __syncwarp(sync_mask);
#else
recv_int4[token_idx * hidden_int4 + i] = out_int4;
#endif
Loading
Loading