[None][test] CI verify L0_PostMerge 2767 (b4d44d3): disagg-v32-fp4 NIXL gen_only IMA - #15241
Merged
tensorrt-cicd merged 41 commits intoJun 15, 2026
Conversation
…#12353) Signed-off-by: ZhaoyangWang <zhaoyangw@nvidia.com>
…ee_gpu_memory_fraction=0.6)` to TestQwen3 (#13852) Signed-off-by: tensorrt-cicd <90828364+tensorrt-cicd@users.noreply.github.com>
…l attention in AutoDeploy (#14906) Signed-off-by: Yueh-Ting Chen <yuehtingc@nvidia.com>
…en (#14905) Signed-off-by: Yueh-Ting Chen <yuehtingc@nvidia.com>
Signed-off-by: TensorRT LLM <90828364+tensorrt-cicd@users.noreply.github.com>
… issue to avoid HF Model not found (#14989) Signed-off-by: FredricZ-2007 <226039983+fredricz-20070104@users.noreply.github.com>
Signed-off-by: Ruodi Lu <ruodil@users.noreply.github.com> Co-authored-by: Ruodi Lu <ruodil@users.noreply.github.com>
…14920) Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Balamurugan Marimuthu <246387390+bmarimuthu-nv@users.noreply.github.com>
…14997) Signed-off-by: Kaiyu Xie <26294424+kaiyux@users.noreply.github.com>
Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> Signed-off-by: xiweny <13230610+VALLIS-NERIA@users.noreply.github.com>
#14995) Signed-off-by: Xin He (SW-GPU) <200704525+xinhe-nv@users.noreply.github.com>
… GLM4.7 Flash test (#14999) Signed-off-by: Balamurugan Marimuthu <246387390+bmarimuthu-nv@users.noreply.github.com>
…erage, and AI Failure Analysis (#14528) Signed-off-by: ZhanruiSunCh <184402041+ZhanruiSunCh@users.noreply.github.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Signed-off-by: ZhanruiSunCh <184402041+ZhanruiSunCh@users.noreply.github.com> Signed-off-by: Emma Qiao <qqiao@nvidia.com> Co-authored-by: Emma Qiao <qqiao@nvidia.com>
…ey can run f… (#15001) Signed-off-by: Taylor Yeonbok Lee <249374542+taylor-yb-lee@users.noreply.github.com>
Signed-off-by: Shreyas Misra <shreyasm@nvidia.com>
Signed-off-by: ziyixiong-nv <219238287+ziyixiong-nv@users.noreply.github.com>
…LLM(...)` constructor and removed the… (#14892) Signed-off-by: tensorrt-cicd <90828364+tensorrt-cicd@users.noreply.github.com>
… standalone tests (#14954) Signed-off-by: Gal Hubara Agam <96368689+galagam@users.noreply.github.com>
… (single GPU, aggregated for now) (#14578) Signed-off-by: Thor Johnsen <41591019+thorjohnsen@users.noreply.github.com>
Signed-off-by: Olivia Stoner <245287810+o-stoner@users.noreply.github.com>
…ading and pre-fetch checking (#14021) Signed-off-by: Yibin Li <109242046+yibinl-nvidia@users.noreply.github.com>
…s shapes (#14977) Signed-off-by: Dongfeng Yu <dongfengy@nvidia.com>
Signed-off-by: Taylor Yeonbok Lee <249374542+taylor-yb-lee@users.noreply.github.com>
Signed-off-by: Alyosha-Swamy <raghav@arcee.ai>
Signed-off-by: TensorRT LLM <90828364+tensorrt-cicd@users.noreply.github.com>
…14684) Signed-off-by: junq <22017000+QiJune@users.noreply.github.com>
Signed-off-by: Yibin Li <109242046+yibinl-nvidia@users.noreply.github.com> Signed-off-by: Chang Liu <9713593+chang-l@users.noreply.github.com> Co-authored-by: Chang Liu <9713593+chang-l@users.noreply.github.com>
Signed-off-by: TensorRT LLM <90828364+tensorrt-cicd@users.noreply.github.com>
Signed-off-by: Taylor Yeonbok Lee <249374542+taylor-yb-lee@users.noreply.github.com>
Signed-off-by: nv-guomingz <137257613+nv-guomingz@users.noreply.github.com>
…/ --extra_llm_api_options YAML (#14812) Signed-off-by: marinayanov <256585945+marinayanov@users.noreply.github.com>
Signed-off-by: Bo Deng <deemod@nvidia.com>
Signed-off-by: Tyler Burt <195370667+tburt-nv@users.noreply.github.com>
Signed-off-by: TensorRT LLM <90828364+tensorrt-cicd@users.noreply.github.com>
…h Cutlass backend - Part 1 (#14923) Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>
…isagg (#14935) Signed-off-by: Dongfeng Yu <dongfengy@nvidia.com>
Signed-off-by: Dongfeng Yu <dongfengy@nvidia.com>
tensorrt-cicd
force-pushed
the
chenfeiz/bisect-6280721-culprit-head
branch
from
June 11, 2026 08:20
eedfdcd to
9f19890
Compare
Collaborator
|
/bot run --disable-fail-fast --stage-list "GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU4-Post-Merge-4" |
Collaborator
Author
|
PR_Github #53531 [ run ] triggered by Bot. Commit: |
Collaborator
Author
|
PR_Github #53531 [ run ] completed with state
|
tensorrt-cicd
force-pushed
the
chenfeiz/bisect-6280721-culprit-head
branch
from
June 15, 2026 06:05
9f19890 to
b4d44d3
Compare
tensorrt-cicd
merged commit Jun 15, 2026
b4d44d3
into
chenfeiz/bisect-6280721-culprit-base
162 of 163 checks passed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Pre-merge CI bisect probe for nvbugs/6280721.
This PR isolates the suspected culprit commit
8e5d9e2[TRTLLM-11508][refactor] Merge Eagle3 and MTP-eagle one-model workers (#12353) by basing the PR on its parent910826b. The PR diff is therefore exactly that one commit (plus an empty sign-off commit; tree is identical to8e5d9e2).Background
Post-merge job 2762 (
316430f) failed the test below while job 2761 (33b0a32) passed. The GEN server crashes during generation CUDA-graph warmup (cudaErrorLaunchFailurein_capture_generation_cuda_graphs), so the disagg endpoint never becomes healthy and the test times out after 5400s. The config usesdecoding_type: MTP, num_nextn_predict_layers: 1, and8e5d9e2is the only in-range commit that rewrites the MTP one-model / spec-metadata path inmodel_engine.py.Expected result: this test fails here, confirming the culprit.
Test under observation
Waiver note
This case is not present in
tests/integration/test_lists/waives.txtat this commit, so no waiver removal is needed — the test runs unwaived.Test Coverage
perf/test_perf_sanity.py::test_e2e[disagg_upload-gen_only-gb200_deepseek-v32-fp4_1k1k_con2048_ctx1_dep4_gen1_dep4_eplb0_mtp1_ccb-NIXL]