Skip to content

[TRTLLM-14810][chore] Catch-up merge: main into feat/kimi_k3 - #17143

Closed
brnguyen2 wants to merge 372 commits into
NVIDIA:feat/kimi_k3from
brnguyen2:merge/main-into-kimi-k3
Closed

[TRTLLM-14810][chore] Catch-up merge: main into feat/kimi_k3#17143
brnguyen2 wants to merge 372 commits into
NVIDIA:feat/kimi_k3from
brnguyen2:merge/main-into-kimi-k3

Conversation

@brnguyen2

Copy link
Copy Markdown
Collaborator

@coderabbitai summary

Description

Catch-up merge of main into feat/kimi_k3 (TRTLLM-14810): merges main @ e34d3d4 (362 commits since the last sync point b602fa6) into the feature branch.

Draft status / how to use this branch: conflict resolution is complete and reviewed; build + unit-test validation has passed (details below), and the wider integration/accuracy qualification is still in progress. Teams blocked on the merge can base work on this branch now, accepting that qualification may still produce small fixups on it.

Note: this branch includes the commits of #17088 (TRTLLM-14703), which was open at merge time. #17088 should merge before or together with this PR.

Conflict resolution summary

13 files had textual conflicts. The notable resolutions:

All ~39 auto-merged files touched by both sides were hand-reviewed: branch-side deltas survived intact; requirements.txt and sa_worker.py intentionally resolve to main.

Test Coverage

  • Build of the merged branch (RelWithDebInfo): passing. The one semantic collision the build caught — the MNNVL allreduce MoE-scale additions vs main's nvinfer1::DataTypetensorrt_llm::DataType migration — is fixed in a dedicated commit on this branch.
  • KDA / SA / MoE unit suites on GB300: 115 passed, 3 skipped, 0 failed (tests/unittest/_torch/modeling/test_kda_mtp_decode_cute_parity.py, test_kimi_kda_fused_verify_parity.py, test_kimi_kda_verify_parity.py, test_kimi_kda_fp8_packed_prefill.py, tests/unittest/_torch/modules/kimi_kda/, tests/unittest/_torch/modules/moe/test_kimi_k3_situ_moe.py, tests/torch/speculative/test_suffix_automaton.py).
  • In progress (will be reported on this PR before undrafting): 4-GPU SA-vs-baseline logits-parity integration run, 16-GPU smoke with and without KV block reuse, GSM8K accuracy sweep (baseline / reuse / SA), serving performance sweep vs a matched-commit baseline, and disaggregated-serving suites.

PR Checklist

  • PR title and description follow the repo conventions
  • Test coverage listed above

xinhe-nv and others added 30 commits July 21, 2026 05:03
Signed-off-by: Xin He (SW-GPU) <200704525+xinhe-nv@users.noreply.github.com>
Signed-off-by: Ivy Zhang <25222398+crazydemo@users.noreply.github.com>
…A#16675)

Signed-off-by: trtllm-agent <296075020+trtllm-agent@users.noreply.github.com>
…ies (NVIDIA#16551)

Signed-off-by: Derek Pitman <dpitman@nvidia.com>
…A#16682)

Signed-off-by: trtllm-agent <296075020+trtllm-agent@users.noreply.github.com>
Signed-off-by: TensorRT LLM <90828364+tensorrt-cicd@users.noreply.github.com>
…IA#16700)

Signed-off-by: trtllm-agent <296075020+trtllm-agent@users.noreply.github.com>
…rRT SDK into images (NVIDIA#16608)

Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>
…A#16698)

Signed-off-by: trtllm-agent <296075020+trtllm-agent@users.noreply.github.com>
Co-authored-by: Mingyang Hao <mingyangh@nvidia.com>
…nput (NVIDIA#16569)

Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>
Signed-off-by: TensorRT LLM <90828364+tensorrt-cicd@users.noreply.github.com>
…rt odd top_k (NVIDIA#16546)

Signed-off-by: tianruih <tianruih@nvidia.com>
…ests for gpt-oss-120b (NVIDIA#16564)

Signed-off-by: Chenfei Zhang <chenfeiz@nvidia.com>
…ic IPC executor path (NVIDIA#16523)

Signed-off-by: Lance Liao <108499334+lancelly@users.noreply.github.com>
…session shutdown test (NVIDIA#16630)

Signed-off-by: JunyiXu-nv <219237550+JunyiXu-nv@users.noreply.github.com>
…undant-warp sync reduction (NVIDIA#16424)

Signed-off-by: siyidNV <297196620+siyidNV@users.noreply.github.com>
Co-authored-by: siyidNV <297196620+siyidNV@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
…VIDIA#16709)

Signed-off-by: Yuxian Qiu <142763828+yuxianq@users.noreply.github.com>
…me to satisfy flashinfer's contract… (NVIDIA#16069)

Signed-off-by: trtllm-agent <296075020+trtllm-agent@users.noreply.github.com>
…nd CI plumbing (NVIDIA#16610)

Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>
…ised] (NVIDIA#16291)

Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>
…nal metric scrapers (NVIDIA#12596)

Signed-off-by: BenjaminBraunDev <benjaminbraun@google.com>
…VIDIA#16468)

Signed-off-by: peihengh <259410613+peihu-nv@users.noreply.github.com>
…d to _forward_impl (NVIDIA#16724)

Signed-off-by: qgai <qgai@nvidia.com>
…IDIA#16696)

Signed-off-by: Pietro Cicotti <5833013+pcicotti@users.noreply.github.com>
…ifier (NVIDIA#16557)

Signed-off-by: Derek Pitman <dpitman@nvidia.com>
litaotju and others added 12 commits July 31, 2026 09:40
…R100 dispatch wiring)

Signed-off-by: Tao Li <tali@nvidia.com>
Signed-off-by: Michal Guzek <mguzek@nvidia.com>
…ed-epilogue batch gate, hardening

Signed-off-by: Michal Guzek <mguzek@nvidia.com>
… checkpoint (NVIDIA#16690)

Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>
Signed-off-by: Michal Guzek <mguzek@nvidia.com>
…uff E402

Signed-off-by: Michal Guzek <mguzek@nvidia.com>
Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>
Signed-off-by: William Zhang <133824995+2ez4bz@users.noreply.github.com>
Catch-up merge (TRTLLM-14810): main @ e34d3d4 into feat/kimi_k3 @
7003910 (= ff8360e + PR NVIDIA#17088 cherry-picks). 13 textual conflicts
resolved per the TRTLLM-14810 resolution plan; K3 stays on the V1
compatibility cache managers (V2 port tracked as TRTLLM-14769).

Signed-off-by: Brian Nguyen <brnguyen@nvidia.com>
…to-kimi-k3

Signed-off-by: Brian Nguyen <brnguyen@nvidia.com>

# Conflicts:
#	tensorrt_llm/_torch/modules/fused_moe/fused_moe_trtllm_gen.py
#	tensorrt_llm/_torch/modules/fused_moe/moe_scheduler.py
…-scale params

Main migrated mnnvlAllreduceKernels.h from nvinfer1::DataType to
tensorrt_llm::DataType (tllmDataType.h); the expert-scale-factor
additions merged from NVIDIA#17088 still referenced nvinfer1. Rename the
scaleDType member and kernel-dispatch comparisons to the new enum
(identical enumerator names/values).

Signed-off-by: Brian Nguyen <brnguyen@nvidia.com>
Signed-off-by: Brian Nguyen <brnguyen@nvidia.com>
@brnguyen2

Copy link
Copy Markdown
Collaborator Author

/bot run

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63130 [ run ] triggered by Bot. Commit: 78ae85b Link to invocation

… trtllm-gen backend

flashinfer 0.6.15 raises ValueError when multi_ctas_kv_counter_buffer is
supplied with any runner other than trtllm-gen; the Kimi K3 cute-dsl MLA
decode path must pass None.

Signed-off-by: Brian Nguyen <brnguyen@nvidia.com>
@brnguyen2

Copy link
Copy Markdown
Collaborator Author

/bot run

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63138 [ run ] triggered by Bot. Commit: ff30e9f Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github/17143-78ae85b #63130 was force-killed by a newer pipeline run.
L0 job information not available (job may not have been triggered yet).

Link to superseding invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63138 [ run ] completed with state FAILURE. Commit: ff30e9f
/LLM/main/L0_MergeRequest_PR pipeline #51226 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@brnguyen2 brnguyen2 added the api-compatible Accepted LLM API contract change that is backwards-compatible label Aug 1, 2026
@brnguyen2

Copy link
Copy Markdown
Collaborator Author

/bot run

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63192 [ run ] triggered by Bot. Commit: 78e0c47 Link to invocation

…erval, spec-decoding KV-estimation workaround

Follow-ups from qualifying the catch-up merge:

- Run pre-commit formatters over files touched by the merge whose
  internal formatting diverged from the repo hooks, and fix the
  docstring style (D205/D209) violations they flagged.
- eval_extra_llm_options_reuse.yaml: the merge renamed
  mamba_state_cache_interval (default 256) to
  mamba_state_config.periodic_snapshot_interval (default 0 = disabled);
  set it explicitly so the reuse eval keeps taking periodic snapshots.
- perf_sweep.sbatch: export PYTHONPATH in the server block (the client
  block already did), so the served code always resolves to the
  checkout under test.
- Workaround for TRTLLM-14903: with speculative decoding enabled and a
  self-spawned MPI session, the KV cache size estimation executor's
  warmup hangs indefinitely while exercising the q>1 generation-path
  attention kernels that only its spec-mode dummy requests reach (the
  pre-merge branch tip ec52c64 passes the identical run). When
  speculative_config is set, skip the estimation phase and size the
  cache analytically via configure_kv_cache_capacity() (the
  KVCacheManagerV2-validated path); non-speculative runs keep the
  normal estimation behavior and the TRTLLM_SKIP_KV_CACHE_ESTIMATION
  gate. With the workaround, an SA-vs-baseline logits-parity
  integration run passes with statistics identical to the pre-merge
  tip (52 prompts, 11 non-tie divergences, zero drift). Remove once
  TRTLLM-14903 is fixed.
- examples/kimi_k3/README: document the TRTLLM-14903 workaround and a
  known performance regression at DEP16 saturation (TRTLLM-14904; 8K/1K
  serving sweep: reproducibly ~15% lower output throughput at
  concurrency 1024 and 4-5% at 128-256 vs the pre-merge tip;
  concurrency <= 64 and the TEP16/TEP8 latency recipes are at parity).
  Developers can A/B against ec52c64, the last commit before this
  merge.

Signed-off-by: Brian Nguyen <brnguyen@nvidia.com>
@brnguyen2
brnguyen2 force-pushed the merge/main-into-kimi-k3 branch from 78e0c47 to fc38923 Compare August 1, 2026 12:35
@brnguyen2

Copy link
Copy Markdown
Collaborator Author

Validation summary

Accuracy — GSM8K with the three serving recipes (baseline decode, KV-block-reuse, suffix-automaton speculative decoding): 96.74 / 96.89 / 96.66 vs a 97.01 pre-merge reference (all within run-to-run noise). The reuse recipe requires the periodic_snapshot_interval setting restored in this PR (the merge renamed the option and its default changed to disabled).

Speculative-decoding parity — SA-vs-baseline logits-parity integration run on a truncated model: post-merge statistics are identical to the pre-merge branch tip (52/52 prompts diverge only at benign near-ties, 11 non-tie divergences, zero drift; same numbers on both trees).

Serving performance — 17-point 8K/1K benchmark sweep across three recipes, merge vs pre-merge branch tip, single run per point plus repeat runs on outliers:

  • Low/mid concurrency: parity within ±2%.
  • Attention-DP EP16 at extreme saturation (concurrency 1024): reproducible ~15% output-throughput deficit vs the pre-merge tip (and −4–5% at concurrency 128–256). Tracked in TRTLLM-14904; low/mid-concurrency operating points are unaffected.

Disaggregated smoke — context/generation split with MTP on the V2 transceiver path: server healthy, benchmark completes.

Known issue (introduced upstream, not K3-specific code) — speculative decoding with a self-spawned MPI session and fraction-based KV-cache sizing hangs during the KV-cache-size-estimation phase (the estimation executor's warmup never completes; the pre-merge tip passes the identical run). This branch carries a temporary workaround that skips the estimation phase whenever a speculative config is set (cache sized analytically instead; non-speculative runs keep normal estimation) — to be removed when the underlying hang is fixed. Production launch paths (trtllm-llmapi-launch, trtllm-serve) are unaffected. Tracked in TRTLLM-14903.

@brnguyen2

Copy link
Copy Markdown
Collaborator Author

Closing this PR without merging via the web UI: the repository is configured for squash merges, and squashing a true merge commit would linearize main's history into feat/kimi_k3 and break every future mainfeat/kimi_k3 sync. The validated branch is instead direct-pushed (fast-forward) to feat/kimi_k3, preserving the merge topology. Final content = the merge plus one squashed follow-up commit (post-merge formatting, the reuse snapshot-interval restore, and the TRTLLM-14903 workaround documented above).

Rollback reference: the feat/kimi_k3 tip prior to this merge is ec52c6418b801b2c444f8ee2bedc35b914351725. Anyone needing the pre-merge behavior (e.g. to A/B the TRTLLM-14904 performance regression or bisect) should use that commit; it remains an ancestor of the pushed branch, so the push is a pure fast-forward with no history rewritten.

@brnguyen2 brnguyen2 closed this Aug 1, 2026
@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63192 [ run ] completed with state FAILURE. Commit: 78e0c47
/LLM/main/L0_MergeRequest_PR pipeline #51276 completed with status: 'ABORTED'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

Link to invocation

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

api-compatible Accepted LLM API contract change that is backwards-compatible VisualGen

Projects

None yet

Development

Successfully merging this pull request may close these issues.