Add static libraries for batch manager#2
Merged
Conversation
Collaborator
|
LGTM, thanks for the quick fix. |
4 tasks
liuyhwangyh
pushed a commit
to liuyhwangyh/TensorRT-LLM
that referenced
this pull request
Mar 21, 2024
# This is the 1st commit message: add download models form www.modelscope.cn # This is the commit message NVIDIA#2: debug # This is the commit message NVIDIA#3: debug
4 tasks
yingcanw
added a commit
that referenced
this pull request
Jan 2, 2025
nv-guomingz
pushed a commit
that referenced
this pull request
Jan 24, 2025
* Add README * Add unified converter (#1) * init v3 lite feat * fix moe topk method * fix noaux_tc logic * fix deepseek v3 normal rope * refactor * wo conversion ok debugging build * add quantize for attn.dense * add unified converter support * testing unified converter * add convert checkpoint and update docs --------- Co-authored-by: Zeyu Wang <zeyuw@nvidia.com> * update README * add FP8 notes * Update run.py result * Update V3 README * Update usages of FP8 to BF16 instruction * fix model name mapping (#2) * Update HF ckpt BF16 conversion. * fix config of deepseek kv cache * Remove source code * Deepseek V3 FP8 Support --------- Co-authored-by: jershi425 <83951930+jershi425@users.noreply.github.com> Co-authored-by: Zeyu Wang <zeyuw@nvidia.com> Co-authored-by: Hanyue He <hanyueh@nvidia.com> Co-authored-by: root <root@h20-2.cm.cluster>
dongxuy04
added a commit
to dongxuy04/TensorRT-LLM
that referenced
this pull request
Apr 25, 2025
Signed-off-by: Dongxu Yang <78518666+dongxuy04@users.noreply.github.com>
dongxuy04
added a commit
that referenced
this pull request
Apr 25, 2025
* add MNNVL memory mapping support Signed-off-by: Dongxu Yang <78518666+dongxuy04@users.noreply.github.com> * add more MPI environment for trtllm-llmapi-launch Signed-off-by: Dongxu Yang <78518666+dongxuy04@users.noreply.github.com> * add MoE communication and prepare kernels Signed-off-by: Dongxu Yang <78518666+dongxuy04@users.noreply.github.com> * add MNNVL AlltoAll support for DeepSeekV3 Signed-off-by: Dongxu Yang <78518666+dongxuy04@users.noreply.github.com> * add output dump for throughput benchmark Signed-off-by: Dongxu Yang <78518666+dongxuy04@users.noreply.github.com> * support dynamic kernel launch grid Signed-off-by: Dongxu Yang <78518666+dongxuy04@users.noreply.github.com> * address review comments Signed-off-by: Dongxu Yang <78518666+dongxuy04@users.noreply.github.com> * address review comments #2 Signed-off-by: Dongxu Yang <78518666+dongxuy04@users.noreply.github.com> --------- Signed-off-by: Dongxu Yang <78518666+dongxuy04@users.noreply.github.com>
wu1du2
pushed a commit
to wu1du2/TensorRT-LLM
that referenced
this pull request
May 11, 2025
danielafrimi
added a commit
to danielafrimi/TensorRT-LLM
that referenced
this pull request
Jun 30, 2025
# This is the 1st commit message: kernel Signed-off-by: Ubuntu <dafrimi@nvidia.com> wip Signed-off-by: Ubuntu <dafrimi@nvidia.com> remove prints Signed-off-by: Ubuntu <dafrimi@nvidia.com> test pass Signed-off-by: Ubuntu <dafrimi@nvidia.com> test refactor with more use cases Signed-off-by: Ubuntu <dafrimi@nvidia.com> refacor Signed-off-by: Ubuntu <dafrimi@nvidia.com> refacor_2 Signed-off-by: Ubuntu <dafrimi@nvidia.com> add tuner wip Signed-off-by: Ubuntu <dafrimi@nvidia.com> autotuner works Signed-off-by: Ubuntu <dafrimi@nvidia.com> bfloat16 works. moer changes to the thop file Signed-off-by: Ubuntu <dafrimi@nvidia.com> is tune for autotuner is True --> gets real tactics configs Signed-off-by: Ubuntu <dafrimi@nvidia.com> wip Signed-off-by: Ubuntu <dafrimi@nvidia.com> wip Signed-off-by: Ubuntu <dafrimi@nvidia.com> zeros + quant mode is works Signed-off-by: Ubuntu <dafrimi@nvidia.com> act int8 Signed-off-by: Ubuntu <dafrimi@nvidia.com> removed fp8 for now Signed-off-by: Ubuntu <dafrimi@nvidia.com> wip Signed-off-by: Ubuntu <dafrimi@nvidia.com> w4a16 linear module Signed-off-by: Ubuntu <dafrimi@nvidia.com> wip Signed-off-by: Ubuntu <dafrimi@nvidia.com> changed cutalss for sm==89 Signed-off-by: Ubuntu <dafrimi@nvidia.com> wip Signed-off-by: Ubuntu <dafrimi@nvidia.com> test linear work Signed-off-by: Ubuntu <dafrimi@nvidia.com> add license Signed-off-by: Ubuntu <dafrimi@nvidia.com> works! Signed-off-by: Ubuntu <dafrimi@nvidia.com> refactor + linear test pass Signed-off-by: Ubuntu <dafrimi@nvidia.com> preprocess in load weights Signed-off-by: Ubuntu <dafrimi@nvidia.com> wip Signed-off-by: Ubuntu <dafrimi@nvidia.com> wip Signed-off-by: Ubuntu <dafrimi@nvidia.com> wip Signed-off-by: Ubuntu <dafrimi@nvidia.com> wip Signed-off-by: Ubuntu <dafrimi@nvidia.com> refactor + rebase Signed-off-by: Ubuntu <dafrimi@nvidia.com> wip Signed-off-by: Ubuntu <dafrimi@nvidia.com> wip Signed-off-by: Ubuntu <dafrimi@nvidia.com> Blackwell not supported Signed-off-by: Daniel Afrimi <dafrimi@nvidia.com> wip Signed-off-by: Daniel Afrimi <dafrimi@nvidia.com> skip blackwell Signed-off-by: Daniel Afrimi <dafrimi@nvidia.com> wip Signed-off-by: Daniel Afrimi <dafrimi@nvidia.com> works Signed-off-by: Ubuntu <dafrimi@nvidia.com> # This is the commit message NVIDIA#2: rebased Signed-off-by: Ubuntu <dafrimi@nvidia.com> # This is the commit message NVIDIA#3: align with my pld worked version of linear Signed-off-by: Ubuntu <dafrimi@nvidia.com> # This is the commit message NVIDIA#4: wip Signed-off-by: Ubuntu <dafrimi@nvidia.com> # This is the commit message NVIDIA#5: refactor Signed-off-by: Daniel Afrimi <danielafrimi8@gmail.com> # This is the commit message NVIDIA#6: refactor Signed-off-by: Daniel Afrimi <danielafrimi8@gmail.com> # This is the commit message NVIDIA#7: refactor Signed-off-by: Daniel Afrimi <danielafrimi8@gmail.com> # This is the commit message NVIDIA#8: refactor Signed-off-by: Daniel Afrimi <danielafrimi8@gmail.com> # This is the commit message NVIDIA#9: sys path Signed-off-by: Daniel Afrimi <danielafrimi8@gmail.com> # This is the commit message NVIDIA#10: sys path Signed-off-by: Daniel Afrimi <danielafrimi8@gmail.com>
yuxianq
added a commit
to yuxianq/TensorRT-LLM
that referenced
this pull request
Jul 17, 2025
Signed-off-by: Yuxian Qiu <142763828+yuxianq@users.noreply.github.com>
litaotju
pushed a commit
to litaotju/TensorRT-LLM
that referenced
this pull request
Jul 24, 2025
Signed-off-by: Yuxian Qiu <142763828+yuxianq@users.noreply.github.com> Signed-off-by: Fanrong Li <23290157+lfr-0531@users.noreply.github.com>
yuxianq
added a commit
to yuxianq/TensorRT-LLM
that referenced
this pull request
Jul 28, 2025
Signed-off-by: Yuxian Qiu <142763828+yuxianq@users.noreply.github.com> Signed-off-by: Fanrong Li <23290157+lfr-0531@users.noreply.github.com>
zongfeijing
pushed a commit
to zongfeijing/TensorRT-LLM
that referenced
this pull request
Jul 31, 2025
Signed-off-by: Yuxian Qiu <142763828+yuxianq@users.noreply.github.com> Signed-off-by: Fanrong Li <23290157+lfr-0531@users.noreply.github.com>
spease
added a commit
to spease/TensorRT-LLM
that referenced
this pull request
Dec 9, 2025
farazkh80
pushed a commit
to farazkh80/TensorRT-LLM
that referenced
this pull request
Dec 11, 2025
Fix: Add pynvml fallback for DGX Spark unified memory
5 tasks
5 tasks
zheyuf
added a commit
to zheyuf/TensorRT-LLM
that referenced
this pull request
Apr 27, 2026
… 1 PR #1) Moves 108 pre-merge + post-merge pytorch 4-GPU tests from l0_dgx_b200.yml to l0_gb200_multi_gpus.yml, deletes 4 dedup tests, and reshards two existing GB200 stages to keep wall-clock under the 4h cluster ceiling. Test moves (test-list YAML edits only): * pre-merge pytorch: 29 tests appended to l0_gb200_multi_gpus.yml * post-merge pytorch: 79 tests appended to l0_gb200_multi_gpus.yml * dedup deletes: 4 tests removed from l0_dgx_b200.yml only Sharding adjustments (jenkins/L0_Test.groovy): * GB200-4_GPUs-PyTorch-{1,2} 2-shard → 4-shard split (add -3, -4) * GB200-4_GPUs-PyTorch-Post-Merge-1 1-shard → 3-shard split (add -2, -3) Mirrors the established DGX_B200-4_GPUs-PyTorch-{1,2,3} + DGX_B200-4_GPUs-PyTorch-Post-Merge-{1,2} pattern, scaled for our larger added test load. Without resharding, projected wall-clocks hit ~4.75h pre-merge / ~7h post-merge, both over the 4h cluster ceiling. After resharding, projected wall-clock per shard: ~2.9h pre-merge / ~3.0h post-merge — within budget with margin. Going-forward capacity reclaimed on B200: ~2,200 GPU-hrs/wk. Out of scope for this PR (Wave 1 PR NVIDIA#2 will handle): * 1-GPU tests (156 tests) - needs new GB200-PyTorch-{1,Post-Merge-1,AutoDeploy-1} stages * 4-GPU autodeploy tests (11) - needs new GB200-4_GPUs-AutoDeploy-1 stage * Perf (41), legacy TRT/Triton (14), 8-GPU, Ray, fabric-dep, aarch64-LOW Verified with scripts/test_to_stage_mapping.py that moved samples route to GB200-4_GPUs-PyTorch-{1,2,3,4} (pre-merge) and GB200-4_GPUs-PyTorch-Post-Merge-{1,2,3} (post-merge), not DGX_B200-PyTorch-*. Ground-truth verification recipe (per Tyler's guidance): compare tests that ran in L0_PostMerge/<latest_job> vs L0_MergeRequest/<this_pr_job> via OpenSearch — every test removed from B200 stages should now appear in GB200 stages, no coverage loss. Signed-off-by: Zheyu Fu <zheyuf@NVIDIA.com>
zheyuf
added a commit
to zheyuf/TensorRT-LLM
that referenced
this pull request
Apr 28, 2026
… 1 PR #1) Moves 108 pre-merge + post-merge pytorch 4-GPU tests from l0_dgx_b200.yml to l0_gb200_multi_gpus.yml (skipping 9 that already exist there as duplicates), deletes 4 hard-coded dedup tests, and reshards the affected GB200 stages to keep wall-clock under the 4h cluster ceiling. Test moves and dedup deletes (test-list YAML edits only): * pre-merge pytorch: 29 tests appended to l0_gb200_multi_gpus.yml * post-merge pytorch: 70 tests appended to l0_gb200_multi_gpus.yml (9 entries already existed in target — dedup, not added) * dedup deletes: 4 TestDeepSeekV3Lite::test_bfloat16_4gpus_python_scheduler entries removed from l0_dgx_b200.yml only (already in l0_gb200_multi_gpus.yml) Sharding adjustments (jenkins/L0_Test.groovy): * GB200-4_GPUs-PyTorch-{1,2} 2-shard → 4-shard split (add -3, -4) * GB200-4_GPUs-PyTorch-Post-Merge-1 1-shard → 3-shard split (add -2, -3) Mirrors the established DGX_B200-4_GPUs-PyTorch-{1,2,3} + DGX_B200-4_GPUs-PyTorch-Post-Merge-{1,2} pattern, scaled for our larger added test load. Without resharding, projected wall-clocks hit ~4.75h pre-merge / ~7h post-merge, both over the 4h cluster ceiling. After resharding, projected wall-clock per shard: ~2.9h pre-merge / ~3.0h post-merge — within budget with margin. Going-forward capacity reclaimed on B200: ~2,200 GPU-hrs/wk (slightly more, since 9 of the moved tests were previously running on BOTH B200 and GB200 post-merge — this consolidation removes the B200 copy). Out of scope for this PR (Wave 1 PR NVIDIA#2 will handle): * 1-GPU tests (156 tests) - needs new GB200-PyTorch-{1,Post-Merge-1,AutoDeploy-1} stages * 4-GPU autodeploy tests (11) - needs new GB200-4_GPUs-AutoDeploy-1 stage * Perf (41), legacy TRT/Triton (14), 8-GPU, Ray, fabric-dep, aarch64-LOW Verified with scripts/test_to_stage_mapping.py that moved samples route to GB200-4_GPUs-PyTorch-{1,2,3,4} (pre-merge) and GB200-4_GPUs-PyTorch-Post-Merge-{1,2,3} (post-merge), not DGX_B200-PyTorch-*. Verified l0_gb200_multi_gpus.yml has zero duplicate test entries (trt-test-db rejects YAMLs with same test/conditions twice). Ground-truth verification recipe (per Tyler's guidance): compare tests that ran in L0_PostMerge/<latest_job> vs L0_MergeRequest/<this_pr_job> via OpenSearch — every test removed from B200 stages should now appear in GB200 stages, no coverage loss. Signed-off-by: Zheyu Fu <zheyuf@NVIDIA.com>
zheyuf
added a commit
to zheyuf/TensorRT-LLM
that referenced
this pull request
Apr 28, 2026
…1 PR #1) Moves 95 pre-merge + post-merge pytorch 4-GPU tests from l0_dgx_b200.yml to l0_gb200_multi_gpus.yml, and reshards the affected GB200 stages to keep wall-clock under the 4h cluster ceiling. Test moves (test-list YAML edits only): * pre-merge pytorch: 25 tests appended to l0_gb200_multi_gpus.yml * post-merge pytorch: 70 tests appended to l0_gb200_multi_gpus.yml 13 tests intentionally kept on BOTH B200 and GB200 (conservative cross-platform coverage; pre-existing duplication, ~90 GPU-hrs/wk): * 4 pre-merge: TestDeepSeekV3Lite::test_bfloat16_4gpus_python_scheduler[*] * 9 post-merge: TestLlama3_1_8BInstruct::test_bfloat16_4gpus[tp4-FLASHINFER-False], TestLlama3_1_8BInstruct::test_fp8_4gpus[pp4-...-TRTLLM-False], 7 × TestDeepSeekV3Lite::test_nvfp4_4gpus[CUTLASS|TRTLLM-mtp_nextn=*-...] These tests were originally added to both YAMLs (commits NVIDIA#11884, NVIDIA#7073/NVIDIA#7074, NVIDIA#11844, NVIDIA#13366) and have no platform-specific decorators. We could dedup, but keeping the duplication is safer if there's any latent host-CPU or NVLink topology-specific behavior. Sharding adjustments (jenkins/L0_Test.groovy): * GB200-4_GPUs-PyTorch-{1,2} 2-shard → 4-shard split (add -3, -4) * GB200-4_GPUs-PyTorch-Post-Merge-1 1-shard → 3-shard split (add -2, -3) Mirrors DGX_B200-4_GPUs-PyTorch-{1,2,3} + DGX_B200-4_GPUs-PyTorch-Post-Merge-{1,2}. Without resharding, projected wall-clocks hit ~4.75h pre-merge / ~7h post-merge, both over the 4h cluster ceiling. After resharding, projected wall-clock per shard: ~2.9h pre-merge / ~3.0h post-merge — within budget with margin. Going-forward capacity reclaimed on B200: ~2,030 GPU-hrs/wk (from the 95 moves; the 13 kept-duplicate tests continue running on B200). Out of scope for this PR (Wave 1 PR NVIDIA#2 will handle): * 1-GPU tests (156 tests) - needs new GB200-PyTorch-{1,Post-Merge-1,AutoDeploy-1} stages * 4-GPU autodeploy tests (11) - needs new GB200-4_GPUs-AutoDeploy-1 stage * Perf (41), legacy TRT/Triton (14), 8-GPU, Ray, fabric-dep, aarch64-LOW Verified with scripts/test_to_stage_mapping.py that moved samples route to GB200-4_GPUs-PyTorch-{1,2,3,4} (pre-merge) and GB200-4_GPUs-PyTorch-Post-Merge-{1,2,3} (post-merge), not DGX_B200-PyTorch-*. Verified l0_gb200_multi_gpus.yml has zero duplicate test entries (trt-test-db rejects same-test-same-conditions duplicates). Ground-truth verification recipe (per Tyler's guidance): compare tests run in L0_PostMerge/<latest_job> vs L0_MergeRequest/<this_pr_job> via OpenSearch. Every B200-removed test should appear in GB200 stages; the 13 intentionally kept tests will appear on both. Signed-off-by: Zheyu Fu <zheyuf@NVIDIA.com>
zheyuf
added a commit
to zheyuf/TensorRT-LLM
that referenced
this pull request
Apr 28, 2026
…1 PR #1) Moves 95 pre-merge + post-merge pytorch 4-GPU tests from l0_dgx_b200.yml to l0_gb200_multi_gpus.yml, and reshards the affected GB200 stages to keep wall-clock under the 4h cluster ceiling. Test moves (test-list YAML edits only): * pre-merge pytorch: 25 tests appended to l0_gb200_multi_gpus.yml * post-merge pytorch: 70 tests appended to l0_gb200_multi_gpus.yml 13 tests intentionally kept on BOTH B200 and GB200 (conservative cross-platform coverage; pre-existing duplication, ~90 GPU-hrs/wk): * 4 pre-merge: TestDeepSeekV3Lite::test_bfloat16_4gpus_python_scheduler[*] * 9 post-merge: TestLlama3_1_8BInstruct::test_bfloat16_4gpus[tp4-FLASHINFER-False], TestLlama3_1_8BInstruct::test_fp8_4gpus[pp4-...-TRTLLM-False], 7 × TestDeepSeekV3Lite::test_nvfp4_4gpus[CUTLASS|TRTLLM-mtp_nextn=*-...] These tests were originally added to both YAMLs (commits NVIDIA#11884, NVIDIA#7073/NVIDIA#7074, NVIDIA#11844, NVIDIA#13366) and have no platform-specific decorators. We could dedup, but keeping the duplication is safer if there's any latent host-CPU or NVLink topology-specific behavior. Sharding adjustments (jenkins/L0_Test.groovy): * GB200-4_GPUs-PyTorch-{1,2} 2-shard → 4-shard split (add -3, -4) * GB200-4_GPUs-PyTorch-Post-Merge-1 1-shard → 3-shard split (add -2, -3) Mirrors DGX_B200-4_GPUs-PyTorch-{1,2,3} + DGX_B200-4_GPUs-PyTorch-Post-Merge-{1,2}. Without resharding, projected wall-clocks hit ~4.75h pre-merge / ~7h post-merge, both over the 4h cluster ceiling. After resharding, projected wall-clock per shard: ~2.9h pre-merge / ~3.0h post-merge — within budget with margin. Going-forward capacity reclaimed on B200: ~2,030 GPU-hrs/wk (from the 95 moves; the 13 kept-duplicate tests continue running on B200). Out of scope for this PR (Wave 1 PR NVIDIA#2 will handle): * 1-GPU tests (156 tests) - needs new GB200-PyTorch-{1,Post-Merge-1,AutoDeploy-1} stages * 4-GPU autodeploy tests (11) - needs new GB200-4_GPUs-AutoDeploy-1 stage * Perf (41), legacy TRT/Triton (14), 8-GPU, Ray, fabric-dep, aarch64-LOW Verified with scripts/test_to_stage_mapping.py that moved samples route to GB200-4_GPUs-PyTorch-{1,2,3,4} (pre-merge) and GB200-4_GPUs-PyTorch-Post-Merge-{1,2,3} (post-merge), not DGX_B200-PyTorch-*. Verified l0_gb200_multi_gpus.yml has zero duplicate test entries (trt-test-db rejects same-test-same-conditions duplicates). Ground-truth verification recipe (per Tyler's guidance): compare tests run in L0_PostMerge/<latest_job> vs L0_MergeRequest/<this_pr_job> via OpenSearch. Every B200-removed test should appear in GB200 stages; the 13 intentionally kept tests will appear on both. Signed-off-by: Zheyu Fu <zheyuf@NVIDIA.com>
zheyuf
added a commit
to zheyuf/TensorRT-LLM
that referenced
this pull request
Apr 30, 2026
…1 PR #1) Moves 95 pre-merge + post-merge pytorch 4-GPU tests from l0_dgx_b200.yml to l0_gb200_multi_gpus.yml, and reshards the affected GB200 stages to keep wall-clock under the 4h cluster ceiling. Test moves (test-list YAML edits only): * pre-merge pytorch: 25 tests appended to l0_gb200_multi_gpus.yml * post-merge pytorch: 70 tests appended to l0_gb200_multi_gpus.yml 13 tests intentionally kept on BOTH B200 and GB200 (conservative cross-platform coverage; pre-existing duplication, ~90 GPU-hrs/wk): * 4 pre-merge: TestDeepSeekV3Lite::test_bfloat16_4gpus_python_scheduler[*] * 9 post-merge: TestLlama3_1_8BInstruct::test_bfloat16_4gpus[tp4-FLASHINFER-False], TestLlama3_1_8BInstruct::test_fp8_4gpus[pp4-...-TRTLLM-False], 7 × TestDeepSeekV3Lite::test_nvfp4_4gpus[CUTLASS|TRTLLM-mtp_nextn=*-...] These tests were originally added to both YAMLs (commits NVIDIA#11884, NVIDIA#7073/NVIDIA#7074, NVIDIA#11844, NVIDIA#13366) and have no platform-specific decorators. We could dedup, but keeping the duplication is safer if there's any latent host-CPU or NVLink topology-specific behavior. Sharding adjustments (jenkins/L0_Test.groovy): * GB200-4_GPUs-PyTorch-{1,2} 2-shard → 4-shard split (add -3, -4) * GB200-4_GPUs-PyTorch-Post-Merge-1 1-shard → 3-shard split (add -2, -3) Mirrors DGX_B200-4_GPUs-PyTorch-{1,2,3} + DGX_B200-4_GPUs-PyTorch-Post-Merge-{1,2}. Without resharding, projected wall-clocks hit ~4.75h pre-merge / ~7h post-merge, both over the 4h cluster ceiling. After resharding, projected wall-clock per shard: ~2.9h pre-merge / ~3.0h post-merge — within budget with margin. Going-forward capacity reclaimed on B200: ~2,030 GPU-hrs/wk (from the 95 moves; the 13 kept-duplicate tests continue running on B200). Out of scope for this PR (Wave 1 PR NVIDIA#2 will handle): * 1-GPU tests (156 tests) - needs new GB200-PyTorch-{1,Post-Merge-1,AutoDeploy-1} stages * 4-GPU autodeploy tests (11) - needs new GB200-4_GPUs-AutoDeploy-1 stage * Perf (41), legacy TRT/Triton (14), 8-GPU, Ray, fabric-dep, aarch64-LOW Verified with scripts/test_to_stage_mapping.py that moved samples route to GB200-4_GPUs-PyTorch-{1,2,3,4} (pre-merge) and GB200-4_GPUs-PyTorch-Post-Merge-{1,2,3} (post-merge), not DGX_B200-PyTorch-*. Verified l0_gb200_multi_gpus.yml has zero duplicate test entries (trt-test-db rejects same-test-same-conditions duplicates). Ground-truth verification recipe (per Tyler's guidance): compare tests run in L0_PostMerge/<latest_job> vs L0_MergeRequest/<this_pr_job> via OpenSearch. Every B200-removed test should appear in GB200 stages; the 13 intentionally kept tests will appear on both. Signed-off-by: Zheyu Fu <zheyuf@NVIDIA.com>
chienchunhung
referenced
this pull request
in chienchunhung/TensorRT-LLM
Apr 30, 2026
…rmanent wedge Document the multi-signature disaggregated-serving wedge surfaced by the rc11 deployment. The report covers the 1P1D reproducer harness, the six labelled failure signatures (sender-side broken-promise after ready, trie cascade-prune assertion, decode-side bad optional access, gen-side checkGenTransferStatus blocking on at_least_num=1, receiver-side queued cancel broken-promise, and the suspected control-path send stall), their mapping to chained test/fix PR pairs (NVIDIA#13571/NVIDIA#13572 for sig #2, NVIDIA#13639/NVIDIA#13640 for sig #1), the in-flight fixes for sig #4 and sig #5, and the relationship to the unrelated companion fixes NVIDIA#12718 and NVIDIA#13119 which are not in rc11. Includes an investigation timeline that explains why each signature surfaced only after the previous one was fixed, and a test-coverage analysis of why the existing unit and integration tests did not catch any of these bugs. Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com> Made-with: Cursor
chienchunhung
referenced
this pull request
in chienchunhung/TensorRT-LLM
Apr 30, 2026
… section Fold two related pieces of analysis into the report as a new section between the Investigation Timeline and "Why the Existing Tests Did Not Catch This": (1) Signature taxonomy refining the naive "burst -> timeout -> cancellation -> bug" framing. Four-of-seven signatures (#1, #3, #5, NVIDIA#6) are direct cancellation-handling bugs; #4 is a structural latent blocking bug that cancellations expose; #2 is an eviction-driven bug that burst traffic exposes via memory pressure; NVIDIA#7 is a NIXL-internal contention bug that the same load shape happens to trigger but which is not strictly a cancellation bug. Includes a refined trigger chain diagram and two precise corrections (burst alone is not the trigger; "cancellation" is one of several entry points to cleanup paths). (2) Cascade map distinguishing two kinds of inter-signature tangling: - Type 1 (a fix produces a new signature): only one case, the #1 fix produces NVIDIA#6 by making the receiver-side !isReady early-return path reachable in production where a latent recv-buffer leak existed. This is why NVIDIA#6 PR (NVIDIA#13673) is explicitly chained on #1 fix PR (NVIDIA#13640). - Type 2 (a fix exposes a pre-existing signature): three cases where #4 fix exposes #5 / NVIDIA#6 and NVIDIA#6 fix exposes NVIDIA#7 because the upstream fix removes the masking effect on the downstream bug. These are not regressions of the fixes; they were latent pre-existing issues. - Subtler third relationship: #4 is structurally a defensive catcher for any upstream bug that produces a never-resolving receiver future. The #4 fix is independently valuable as defence in depth, not just a symptomatic patch. Also includes a fix-to-file mapping showing that the fixes do not overlap in code; the only structural dependency is the NVIDIA#6 -> #1 chain enforced by the PR base. Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com> Made-with: Cursor
1 task
2 tasks
1 task
chienchunhung
referenced
this pull request
in chienchunhung/TensorRT-LLM
May 20, 2026
…view Apply 5 review-driven edits to §16 to address residual concerns: Walk ordering (concern NVIDIA#8): change GMS RO and MX-receiver alias walks from per-module to top-level model.setup_aliases(). Matches the §7 mitigation contract from ai-dynamo/dynamo PR NVIDIA#7053 ("Call model.post_load_weights() (top-level only) before materialize_module_from_gms()"). transform_weights() and cache_derived_state() walks remain per-module since those bodies live on submodules. New "Why setup_aliases() is top-level-only" callout documents the asymmetry. Lifecycle of _weights_transformed (concern #4): new subsection specifying explicit set/reset/orthogonality rules. Includes a 2x2 truth table showing _weights_removed and _weights_transformed track different lifecycles and can take any combination. Reset is the orchestrator's responsibility (e.g., ModelLoader.reload() resets the flag before re-binding tensors); subclasses do not manage reset. Hard preconditions (concern #5): promote MX source-identity completeness from "open question" to "hard precondition P1." Lists transform-affecting parameters that MX identity must cover (attn_backend, quant backend list, FP8/NVFP4 fusion strategy, TP/EP layout, model revision). Specifies an in-tree backend-fingerprint fail-safe as the fallback if upstream MX cannot guarantee completeness. P2 documents that orchestrator owns _weights_transformed reset. Removes redundant open question NVIDIA#6 from the table. Cosmetic fixes (concerns #2, NVIDIA#6): "four stages" -> "three per-module stages plus orchestrator-managed per-process finalization." cache_derived_state description softened to "reserved for data-dependent state where it exists; many existing modules will have empty bodies." Scope clarifications (concerns #1, #3, NVIDIA#7): - Tiny PR scope reframed as "duck-typed helpers, not inheritance" with citations to existing model_loader.py walker pattern. Lists 4 walker helpers (_setup_aliases, _walk_transform, _walk_cache_state, _walk_full_post_load). - Migration callout: when migrating a subclass, the old post_load_weights() override MUST be removed; otherwise the new staged calls silently no-op. - Family PR #2 (Linear/Attention) gains a "quant-method callback decision" note with default = keep quant_method.post_load_weights callback name (no rename). No code changes. Drives Tiny prep PR scope and family-PR migration sequence. References: - TRT-LLM PR NVIDIA#13926 (GMS-only) - TRT-LLM PR NVIDIA#14151 (MX shim refactor) - ai-dynamo/dynamo PR NVIDIA#7053 (upstream GMS prototype) Signed-off-by: Chien-Chun Hung <chienchunh@nvidia.com> Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
1 task
1 task
aranadive
pushed a commit
to aranadive/TensorRT-LLM
that referenced
this pull request
Jun 9, 2026
remote-g2 change remote g2 lookup api
oen123456
referenced
this pull request
in ddps-lab/disaggregation-TensorRT-LLM
Jun 10, 2026
9-에이전트 워크플로우(v1.2.1 소스 전수조사) + 핵심 3주장 적대적 검증 결과: - §6 재작성: per-side RPS는 orchestrator :8000 /prometheus/metrics의 ctx_/gen_completed_requests_total(단일 엔드포인트, role 접두)로 — 워커별 스크레이핑보다 깔끔. per-side TPS는 공식불가(토큰 카운터 0개 confirmed) → RPS×고정길이 파생. 함정 #2(orchestrator 404) 정정. - §7: benchmark_serving E2EL 기본 미출력(--percentile-metrics에 e2el 필수) 검증 반영 + 단일엔드포인트라 per-side 불가 명시. - §9 TODO: "소스확정 → 런타임 노출만 GPU 스모크로" 확인항목으로 갱신. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
chuangz0
added a commit
to chuangz0/TensorRT-LLM
that referenced
this pull request
Jun 12, 2026
#1 FD <-> chunk position misalignment in P2pHandleExporter Old code kept POSIX FDs in a flat mExportedFds and addressed them in removeHandles() by accumulating chunk counts across earlier pools. Pool-extension paths in exportHandles() append chunks to an existing pool but FDs to the END of mExportedFds, breaking that invariant. Removing a registered range could close another pool's FDs and leak the actual ones, also corrupting any in-flight UDS handshake. Fix: store FDs on P2pMemPool.fds in lockstep with chunks. Removal walks only the owning pool's fds - no cross-pool indexing possible. NVIDIA#2 UDS socket exposed POSIX FDs to other local users /tmp/trt_llm_p2p_fd_*.sock was created with default umask, no peer-credential check. Any same-host process that could connect could SCM_RIGHTS-receive the FDs and import the exporter's GPU memory. Fix: tight umask + explicit chmod 0600 around bind, fail closed if either fails; SO_PEERCRED gate at accept() to refuse non-same-uid clients (defense in depth). NVIDIA#3 P2pTransferContextPool never erased per-thread Contexts Long-lived processes whose callers are transient threads leaked ~64MB cubTempStorage + a worker pool per dead thread; OS thread::id reuse could also hand a fresh thread a stale Context built in a different CUDA context. Fix: pool is now held by shared_ptr (factory + private ctor enforce this); contextForCurrentThread installs a thread_local guard that, on thread exit, calls eraseForCurrentThread via weak_ptr (no-op if pool is already gone, so destruction order between agent and caller threads is safe in either direction). NVIDIA#4 Reads of P2pHandleExporter::mLocalInfo were unprotected getLocalAgentDesc did isSupported() then getLocalInfo().serialize() - a TOCTOU pair where a concurrent registerMemory could append a pool between the two reads, producing a half-written wire blob. Fix: single mInfoMutex protects mLocalInfo / mDetectedHandleType / mUdsPath end-to-end. New serializeIfSupported() does both checks under the lock and replaces the pair at the only caller. The previous mPoolsMutex (added for FDs) is the same mutex, renamed. NVIDIA#8 Document acquire-then-record contract on CudaEventPool Recycled events carry their prior recording; calling cudaEventQuery before record() returns stale state. All current callers already follow acquire->record->query, but the contract is now spelled out on the API doc so future callers don't diverge. Tests: p2pTransferAgentTest +2 regressions - SerializeIfSupportedEmptyOnFreshAgent: locked accessor returns empty on a freshly constructed agent. - TransientCallerThreadsReleaseContextOnExit: 32 transient caller threads each take a Context from the same agent; after join, pool->numContexts() must be exactly 0 (proves thread-exit guard fires). Local validation: p2pMemInfoTest 11/11 transferAgentTest 18/18 (AgentDesc + VmmDescSplitter) p2pTransferAgentTest 36/36 nixlP2pE2ETest 7/7 (+2 env-gated, validated separately) Signed-off-by: Chuang Zhu <111838961+chuangz0@users.noreply.github.com>
chuangz0
added a commit
to chuangz0/TensorRT-LLM
that referenced
this pull request
Jun 15, 2026
#1 FD <-> chunk position misalignment in P2pHandleExporter Old code kept POSIX FDs in a flat mExportedFds and addressed them in removeHandles() by accumulating chunk counts across earlier pools. Pool-extension paths in exportHandles() append chunks to an existing pool but FDs to the END of mExportedFds, breaking that invariant. Removing a registered range could close another pool's FDs and leak the actual ones, also corrupting any in-flight UDS handshake. Fix: store FDs on P2pMemPool.fds in lockstep with chunks. Removal walks only the owning pool's fds - no cross-pool indexing possible. NVIDIA#2 UDS socket exposed POSIX FDs to other local users /tmp/trt_llm_p2p_fd_*.sock was created with default umask, no peer-credential check. Any same-host process that could connect could SCM_RIGHTS-receive the FDs and import the exporter's GPU memory. Fix: tight umask + explicit chmod 0600 around bind, fail closed if either fails; SO_PEERCRED gate at accept() to refuse non-same-uid clients (defense in depth). NVIDIA#3 P2pTransferContextPool never erased per-thread Contexts Long-lived processes whose callers are transient threads leaked ~64MB cubTempStorage + a worker pool per dead thread; OS thread::id reuse could also hand a fresh thread a stale Context built in a different CUDA context. Fix: pool is now held by shared_ptr (factory + private ctor enforce this); contextForCurrentThread installs a thread_local guard that, on thread exit, calls eraseForCurrentThread via weak_ptr (no-op if pool is already gone, so destruction order between agent and caller threads is safe in either direction). NVIDIA#4 Reads of P2pHandleExporter::mLocalInfo were unprotected getLocalAgentDesc did isSupported() then getLocalInfo().serialize() - a TOCTOU pair where a concurrent registerMemory could append a pool between the two reads, producing a half-written wire blob. Fix: single mInfoMutex protects mLocalInfo / mDetectedHandleType / mUdsPath end-to-end. New serializeIfSupported() does both checks under the lock and replaces the pair at the only caller. The previous mPoolsMutex (added for FDs) is the same mutex, renamed. NVIDIA#8 Document acquire-then-record contract on CudaEventPool Recycled events carry their prior recording; calling cudaEventQuery before record() returns stale state. All current callers already follow acquire->record->query, but the contract is now spelled out on the API doc so future callers don't diverge. Tests: p2pTransferAgentTest +2 regressions - SerializeIfSupportedEmptyOnFreshAgent: locked accessor returns empty on a freshly constructed agent. - TransientCallerThreadsReleaseContextOnExit: 32 transient caller threads each take a Context from the same agent; after join, pool->numContexts() must be exactly 0 (proves thread-exit guard fires). Local validation: p2pMemInfoTest 11/11 transferAgentTest 18/18 (AgentDesc + VmmDescSplitter) p2pTransferAgentTest 36/36 nixlP2pE2ETest 7/7 (+2 env-gated, validated separately) Signed-off-by: Chuang Zhu <111838961+chuangz0@users.noreply.github.com>
chuangz0
added a commit
to chuangz0/TensorRT-LLM
that referenced
this pull request
Jun 16, 2026
#1 FD <-> chunk position misalignment in P2pHandleExporter Old code kept POSIX FDs in a flat mExportedFds and addressed them in removeHandles() by accumulating chunk counts across earlier pools. Pool-extension paths in exportHandles() append chunks to an existing pool but FDs to the END of mExportedFds, breaking that invariant. Removing a registered range could close another pool's FDs and leak the actual ones, also corrupting any in-flight UDS handshake. Fix: store FDs on P2pMemPool.fds in lockstep with chunks. Removal walks only the owning pool's fds - no cross-pool indexing possible. NVIDIA#2 UDS socket exposed POSIX FDs to other local users /tmp/trt_llm_p2p_fd_*.sock was created with default umask, no peer-credential check. Any same-host process that could connect could SCM_RIGHTS-receive the FDs and import the exporter's GPU memory. Fix: tight umask + explicit chmod 0600 around bind, fail closed if either fails; SO_PEERCRED gate at accept() to refuse non-same-uid clients (defense in depth). NVIDIA#3 P2pTransferContextPool never erased per-thread Contexts Long-lived processes whose callers are transient threads leaked ~64MB cubTempStorage + a worker pool per dead thread; OS thread::id reuse could also hand a fresh thread a stale Context built in a different CUDA context. Fix: pool is now held by shared_ptr (factory + private ctor enforce this); contextForCurrentThread installs a thread_local guard that, on thread exit, calls eraseForCurrentThread via weak_ptr (no-op if pool is already gone, so destruction order between agent and caller threads is safe in either direction). NVIDIA#4 Reads of P2pHandleExporter::mLocalInfo were unprotected getLocalAgentDesc did isSupported() then getLocalInfo().serialize() - a TOCTOU pair where a concurrent registerMemory could append a pool between the two reads, producing a half-written wire blob. Fix: single mInfoMutex protects mLocalInfo / mDetectedHandleType / mUdsPath end-to-end. New serializeIfSupported() does both checks under the lock and replaces the pair at the only caller. The previous mPoolsMutex (added for FDs) is the same mutex, renamed. NVIDIA#8 Document acquire-then-record contract on CudaEventPool Recycled events carry their prior recording; calling cudaEventQuery before record() returns stale state. All current callers already follow acquire->record->query, but the contract is now spelled out on the API doc so future callers don't diverge. Tests: p2pTransferAgentTest +2 regressions - SerializeIfSupportedEmptyOnFreshAgent: locked accessor returns empty on a freshly constructed agent. - TransientCallerThreadsReleaseContextOnExit: 32 transient caller threads each take a Context from the same agent; after join, pool->numContexts() must be exactly 0 (proves thread-exit guard fires). Local validation: p2pMemInfoTest 11/11 transferAgentTest 18/18 (AgentDesc + VmmDescSplitter) p2pTransferAgentTest 36/36 nixlP2pE2ETest 7/7 (+2 env-gated, validated separately) Signed-off-by: Chuang Zhu <111838961+chuangz0@users.noreply.github.com>
chuangz0
added a commit
to chuangz0/TensorRT-LLM
that referenced
this pull request
Jun 17, 2026
#1 FD <-> chunk position misalignment in P2pHandleExporter Old code kept POSIX FDs in a flat mExportedFds and addressed them in removeHandles() by accumulating chunk counts across earlier pools. Pool-extension paths in exportHandles() append chunks to an existing pool but FDs to the END of mExportedFds, breaking that invariant. Removing a registered range could close another pool's FDs and leak the actual ones, also corrupting any in-flight UDS handshake. Fix: store FDs on P2pMemPool.fds in lockstep with chunks. Removal walks only the owning pool's fds - no cross-pool indexing possible. NVIDIA#2 UDS socket exposed POSIX FDs to other local users /tmp/trt_llm_p2p_fd_*.sock was created with default umask, no peer-credential check. Any same-host process that could connect could SCM_RIGHTS-receive the FDs and import the exporter's GPU memory. Fix: tight umask + explicit chmod 0600 around bind, fail closed if either fails; SO_PEERCRED gate at accept() to refuse non-same-uid clients (defense in depth). NVIDIA#3 P2pTransferContextPool never erased per-thread Contexts Long-lived processes whose callers are transient threads leaked ~64MB cubTempStorage + a worker pool per dead thread; OS thread::id reuse could also hand a fresh thread a stale Context built in a different CUDA context. Fix: pool is now held by shared_ptr (factory + private ctor enforce this); contextForCurrentThread installs a thread_local guard that, on thread exit, calls eraseForCurrentThread via weak_ptr (no-op if pool is already gone, so destruction order between agent and caller threads is safe in either direction). NVIDIA#4 Reads of P2pHandleExporter::mLocalInfo were unprotected getLocalAgentDesc did isSupported() then getLocalInfo().serialize() - a TOCTOU pair where a concurrent registerMemory could append a pool between the two reads, producing a half-written wire blob. Fix: single mInfoMutex protects mLocalInfo / mDetectedHandleType / mUdsPath end-to-end. New serializeIfSupported() does both checks under the lock and replaces the pair at the only caller. The previous mPoolsMutex (added for FDs) is the same mutex, renamed. NVIDIA#8 Document acquire-then-record contract on CudaEventPool Recycled events carry their prior recording; calling cudaEventQuery before record() returns stale state. All current callers already follow acquire->record->query, but the contract is now spelled out on the API doc so future callers don't diverge. Tests: p2pTransferAgentTest +2 regressions - SerializeIfSupportedEmptyOnFreshAgent: locked accessor returns empty on a freshly constructed agent. - TransientCallerThreadsReleaseContextOnExit: 32 transient caller threads each take a Context from the same agent; after join, pool->numContexts() must be exactly 0 (proves thread-exit guard fires). Local validation: p2pMemInfoTest 11/11 transferAgentTest 18/18 (AgentDesc + VmmDescSplitter) p2pTransferAgentTest 36/36 nixlP2pE2ETest 7/7 (+2 env-gated, validated separately) Signed-off-by: Chuang Zhu <111838961+chuangz0@users.noreply.github.com>
chuangz0
added a commit
to chuangz0/TensorRT-LLM
that referenced
this pull request
Jun 24, 2026
#1 FD <-> chunk position misalignment in P2pHandleExporter Old code kept POSIX FDs in a flat mExportedFds and addressed them in removeHandles() by accumulating chunk counts across earlier pools. Pool-extension paths in exportHandles() append chunks to an existing pool but FDs to the END of mExportedFds, breaking that invariant. Removing a registered range could close another pool's FDs and leak the actual ones, also corrupting any in-flight UDS handshake. Fix: store FDs on P2pMemPool.fds in lockstep with chunks. Removal walks only the owning pool's fds - no cross-pool indexing possible. NVIDIA#2 UDS socket exposed POSIX FDs to other local users /tmp/trt_llm_p2p_fd_*.sock was created with default umask, no peer-credential check. Any same-host process that could connect could SCM_RIGHTS-receive the FDs and import the exporter's GPU memory. Fix: tight umask + explicit chmod 0600 around bind, fail closed if either fails; SO_PEERCRED gate at accept() to refuse non-same-uid clients (defense in depth). NVIDIA#3 P2pTransferContextPool never erased per-thread Contexts Long-lived processes whose callers are transient threads leaked ~64MB cubTempStorage + a worker pool per dead thread; OS thread::id reuse could also hand a fresh thread a stale Context built in a different CUDA context. Fix: pool is now held by shared_ptr (factory + private ctor enforce this); contextForCurrentThread installs a thread_local guard that, on thread exit, calls eraseForCurrentThread via weak_ptr (no-op if pool is already gone, so destruction order between agent and caller threads is safe in either direction). NVIDIA#4 Reads of P2pHandleExporter::mLocalInfo were unprotected getLocalAgentDesc did isSupported() then getLocalInfo().serialize() - a TOCTOU pair where a concurrent registerMemory could append a pool between the two reads, producing a half-written wire blob. Fix: single mInfoMutex protects mLocalInfo / mDetectedHandleType / mUdsPath end-to-end. New serializeIfSupported() does both checks under the lock and replaces the pair at the only caller. The previous mPoolsMutex (added for FDs) is the same mutex, renamed. NVIDIA#8 Document acquire-then-record contract on CudaEventPool Recycled events carry their prior recording; calling cudaEventQuery before record() returns stale state. All current callers already follow acquire->record->query, but the contract is now spelled out on the API doc so future callers don't diverge. Tests: p2pTransferAgentTest +2 regressions - SerializeIfSupportedEmptyOnFreshAgent: locked accessor returns empty on a freshly constructed agent. - TransientCallerThreadsReleaseContextOnExit: 32 transient caller threads each take a Context from the same agent; after join, pool->numContexts() must be exactly 0 (proves thread-exit guard fires). Local validation: p2pMemInfoTest 11/11 transferAgentTest 18/18 (AgentDesc + VmmDescSplitter) p2pTransferAgentTest 36/36 nixlP2pE2ETest 7/7 (+2 env-gated, validated separately) Signed-off-by: Chuang Zhu <111838961+chuangz0@users.noreply.github.com>
chuangz0
added a commit
to chuangz0/TensorRT-LLM
that referenced
this pull request
Jun 25, 2026
#1 FD <-> chunk position misalignment in P2pHandleExporter Old code kept POSIX FDs in a flat mExportedFds and addressed them in removeHandles() by accumulating chunk counts across earlier pools. Pool-extension paths in exportHandles() append chunks to an existing pool but FDs to the END of mExportedFds, breaking that invariant. Removing a registered range could close another pool's FDs and leak the actual ones, also corrupting any in-flight UDS handshake. Fix: store FDs on P2pMemPool.fds in lockstep with chunks. Removal walks only the owning pool's fds - no cross-pool indexing possible. NVIDIA#2 UDS socket exposed POSIX FDs to other local users /tmp/trt_llm_p2p_fd_*.sock was created with default umask, no peer-credential check. Any same-host process that could connect could SCM_RIGHTS-receive the FDs and import the exporter's GPU memory. Fix: tight umask + explicit chmod 0600 around bind, fail closed if either fails; SO_PEERCRED gate at accept() to refuse non-same-uid clients (defense in depth). NVIDIA#3 P2pTransferContextPool never erased per-thread Contexts Long-lived processes whose callers are transient threads leaked ~64MB cubTempStorage + a worker pool per dead thread; OS thread::id reuse could also hand a fresh thread a stale Context built in a different CUDA context. Fix: pool is now held by shared_ptr (factory + private ctor enforce this); contextForCurrentThread installs a thread_local guard that, on thread exit, calls eraseForCurrentThread via weak_ptr (no-op if pool is already gone, so destruction order between agent and caller threads is safe in either direction). NVIDIA#4 Reads of P2pHandleExporter::mLocalInfo were unprotected getLocalAgentDesc did isSupported() then getLocalInfo().serialize() - a TOCTOU pair where a concurrent registerMemory could append a pool between the two reads, producing a half-written wire blob. Fix: single mInfoMutex protects mLocalInfo / mDetectedHandleType / mUdsPath end-to-end. New serializeIfSupported() does both checks under the lock and replaces the pair at the only caller. The previous mPoolsMutex (added for FDs) is the same mutex, renamed. NVIDIA#8 Document acquire-then-record contract on CudaEventPool Recycled events carry their prior recording; calling cudaEventQuery before record() returns stale state. All current callers already follow acquire->record->query, but the contract is now spelled out on the API doc so future callers don't diverge. Tests: p2pTransferAgentTest +2 regressions - SerializeIfSupportedEmptyOnFreshAgent: locked accessor returns empty on a freshly constructed agent. - TransientCallerThreadsReleaseContextOnExit: 32 transient caller threads each take a Context from the same agent; after join, pool->numContexts() must be exactly 0 (proves thread-exit guard fires). Local validation: p2pMemInfoTest 11/11 transferAgentTest 18/18 (AgentDesc + VmmDescSplitter) p2pTransferAgentTest 36/36 nixlP2pE2ETest 7/7 (+2 env-gated, validated separately) Signed-off-by: Chuang Zhu <111838961+chuangz0@users.noreply.github.com>
3 tasks
Merged
3 tasks
KleinBlueC
pushed a commit
to KleinBlueC/TensorRT-LLM
that referenced
this pull request
Jul 7, 2026
…E. All 8 acceptance criteria hold at runtime against...
QA verdict: APPROVE (weighted_score 8.6)
Change vs main: 7 files changed, 1376 insertions(+), 8 deletions(-)
Build iterations run: 13
Code-quality verification: clean after 2 round(s)
- round 1 (REJECT):
ROUND MIS-SCOPED — patch not actually reviewed. This round's /simplify, checklist, and /code-review executed with cwd = the agent-flow framework repo (branch kleinc/code_quality_improvement_v1) and reviewed the code-quality-stage feature itself, NOT this task's PR patch. The ChatGLM3 code is on branch agent-team/chatglm3-6b-bringup-v1 in the TensorRT-LLM snapshot repo (a different directory), so `/simplify`/`/code-review` (which diff the cwd's HEAD) never touched it. Root cause: the Quality agent runs in the framework repo's cwd, not the snapshot repo; another round will misfire identically unless the Quality agent runs with the TensorRT-LLM task branch as its working directory.
I re-reviewed the ACTUAL patch (git diff main in TensorRT-LLM). Real, unaddressed findings:
- TESTS (duplication): `_patch_chatglm3_hf_tied_compat()` duplicated in all 4 new test files; `_find_cuda_graph_runner()` and `_assert_cuda_graph_hard_path()` duplicated across test_chatglm3_gsm8k.py and test_chatglm3_replay.py; CHATGLM3_CKPT const, skip decorators, and RuntimeCfg class repeated. Consolidate to a shared util/conftest (keep all coverage).
- modeling_chatglm.py (simplification): redundant post-normalization reads of attention_bias/mlp_bias/head_dim via `getattr(...) or ...` after normalize_chatglm_config already set them; normalize_chatglm_config returns config but both callers discard it; a few verbose comments.
- ALTITUDE (larger, behavior-sensitive; flagged not applied): custom imperative load_weights vs framework BaseWeightMapper; ChatGLM-specific normalization special-cased in shared config_utils.py; partial_rotary_factor/rope_fusion hardcoded.
No fixes applied: the round did not validly review this patch, and edits to modeling/test code cannot be validated against the GSM8K accuracy criterion in this environment (risking a criterion regression). Recommend fixing the Quality agent's working directory (must be the snapshot repo/task branch) before the next round, then applying the test-dedup and post-normalization simplifications above.
- round 2 (REJECT):
Round 2 finally assessed the ACTUAL task patch (git diff main...agent-team/chatglm3-6b-bringup-v1 in the TensorRT-LLM repo), unlike round 1 whose skill steps mis-scoped to the agent-flow framework repo (cwd). Note: the automated stage still mis-scopes — /simplify and /code-review diff the cwd (agent-flow), not this task's snapshot repo — so I targeted the correct repo manually.
/code-review (correctness): CLEAN — returned []. No correctness bugs survive verification. The high-risk items are validated: partial-RoPE, QKV/GQA split sizes (32 heads / 2 KV / head_dim 128), and transformer.* weight-name mapping are exercised by the GSM8K-parity acceptance gate QA approved; the fused-residual path is completed by the final self.norm(hidden_states, residual); embedding matches the Qwen sibling and load_weights handles attention-DP (tp_size = 1 if enable_attention_dp else mapping.tp_size). Speculative "other ChatGLM variant" concerns (_set_missing fill-if-None, head_dim None if kv_channels absent, missing dense/output biases) are not reachable for chatglm3-6b and are guarded by strict _take/consumed accounting.
/simplify (round 2): applied one behavior-preserving cleanup (dropped the unused `return config` from normalize_chatglm_config, annotation -> None); it was subsequently reverted in the working tree, so the patch is unchanged.
OUTSTANDING (why REJECT): the checklist's test item is not satisfied — heavy test-helper duplication across the 4 new test files (`_patch_chatglm3_hf_tied_compat` in all 4; `_find_cuda_graph_runner`/`_assert_cuda_graph_hard_path`/`RuntimeCfg` across gsm8k+replay; repeated CHATGLM3_CKPT/skip decorators), ~100+ duplicated lines to consolidate into a shared util; plus minor post-normalization config-read redundancy in modeling_chatglm.py. NOT applied here because tensorrt_llm is not importable in this environment and there is no GSM8K/test rig, so a test-file move cannot be validated for collection/pass (risking the acceptance criterion), and the working-tree edit I did make was reverted.
RECOMMENDATION: this cannot be resolved by another automated round as currently wired — fix the Quality agent's working directory so /simplify and /code-review target the snapshot repo/task branch, then apply the test-helper de-duplication with the test rig available (or the author applies it directly). Correctness and framework-conformance are clean; the residual is a non-behavioral maintainability cleanup.
- round 1 (REJECT):
Code-quality round 1 for the chatglm3-6b bringup patch (modeling_chatglm.py, config_utils.py, _torch/models/__init__.py + 4 new test files).
/simplify applied (safe, behavior-preserving, verified against repo convention):
- modeling_chatglm.py: consume the normalized config fields instead of re-deriving booleans inline — `bias=config.attention_bias`, `dense_bias=config.mlp_bias`, GatedMLP `bias=config.mlp_bias`, and `head_dim=config.head_dim` at both sites (attention + load_weights). Confirmed no base infra reads these fields and that the model already hard-depends on normalize_chatglm_config having run; matches Llama/Phi3/Nemotron/Parakeet.
- Both integration tests: removed the unreachable BFS-over-object-graph fallback in `_find_cuda_graph_runner`, keeping the deterministic `_executor.engine.model_engine.cuda_graph_runner` chain (verified in source; the runner is at a fixed location in single-process mode and in a subprocess otherwise, where BFS also fails).
Codex checklist pass: the working tree carried minor docstring/comment trims in the ChatGLM3 test files (removed helper docstrings on `_patch_chatglm3_hf_tied_compat`, `_pick_metric_key`, `greedy_generate`, and an inline comment); no behavioral change.
/code-review (high) found 10 items. Fixed this turn: (NVIDIA#4) `_pick_metric_key` in test_chatglm3_gsm8k.py now raises loudly instead of silently gating on an arbitrary non-exact_match metric — strict hardening, no change to any real gsm8k run.
OUTSTANDING (why not auto-fixed):
- CORRECTNESS, latent/out-of-target: (NVIDIA#1) load_weights only handles chatglm3-6b's bias layout — it unconditionally fetches the QKV bias and never loads dense/MLP biases, so ChatGLM2/3 variants with add_qkv_bias=False or add_bias_linear=True would KeyError or trip the strict unconsumed-weight ValueError. (NVIDIA#2) normalize_chatglm_config never maps ChatGLM `rope_ratio` to `rope_theta` (get_hf_rope_theta reads config.rope_theta), so chatglm3-6b-32k/128k silently get theta=10000 and degrade at long context. Both are no-ops for base chatglm3-6b (add_qkv_bias=True/add_bias_linear=False/rope_ratio=1) so QA passed; fixing them adds untestable variant paths (no GPU/checkpoint here) and is scope creep on a base bringup — hand to the coder to fix+test or narrow the module's "ChatGLM2/3" claim.
- TEST ROBUSTNESS: (NVIDIA#3) the CUDA-graph hard-path tests require TLLM_WORKER_USE_SINGLE_PROCESS=1 but only name it in an assert string and never set it; without it, ENABLED-cfg tests hard-fail on `assert runner is not None` and BASELINE silently skips its checks. Not auto-fixed because gating/skipping would weaken the author's intended hard-path assertion — needs a CI-env or fixture decision.
- ALTITUDE/REUSE (larger refactors, unsafe without test execution): (NVIDIA#5) hand-rolled load_weights + consumed-set + ignorable-suffix lists reimplement BaseWeightMapper/register_mapper (cf. Glm4WeightLoader in modeling_glm.py); (NVIDIA#9) normalize special-cased in shared load_pretrained_config + duplicated in __init__ vs a ChatGLMConfigLoader; (NVIDIA#6) `_patch_chatglm3_hf_tied_compat` (4 copies) + RuntimeCfg/_build_llm/_find_cuda_graph_runner/_assert_cuda_graph_hard_path (2 copies) duplicated across two separate, GPU-gated test trees; (NVIDIA#10) _GreedyHFLM._model_generate hand-rolls HF greedy decode.
- EFFICIENCY: (NVIDIA#7) test_chatglm3_source_activation_replay reloads the 6B HF model + rebuilds TRT model for both parametrized scenarios though only the trailing cuda_graph block differs (use the _HFRef cache pattern).
- SCOPE CREEP: (NVIDIA#8) the diff reflows unrelated bart/minimaxm3/qwen imports (+ a mistral import in config_utils.py); left as-is because formatter ownership is ambiguous (line 33 is 89 chars vs isort/yapf's 80) and reverting risks fighting the pre-commit hook.
Rejecting: a substantive edit was made this turn and several actionable findings remain (test env-var robustness is the most important; the two latent correctness gaps and the altitude/reuse refactors should be triaged by the coder).
- round 2 (APPROVE):
Code-quality round 2 for the chatglm3-6b bringup patch.
/simplify (round 2) applied 3 trivial, behavior-neutral cosmetic cleanups in modeling_chatglm.py: inlined the single-use `head_dim` local in ChatGLMAttention (`head_dim=config.head_dim` passed directly to super()), inlined the single-use `head_dim` temp in normalize_chatglm_config, and dropped the noise `: set` annotation on `consumed`. Reuse/efficiency/altitude agents found nothing new safely-actionable and confirmed the two "small fold" hypotheses don't hold (base `skip_modules` matches model module names, not the checkpoint keys `_IGNORABLE_*` filters; the second normalize_chatglm_config call is load-bearing for the direct-ModelConfig test path) — the only remaining altitude/reuse items are large refactors.
Codex checklist pass (round 2): no net changes — independent verification confirmed the only diffs since round 1 were the 4 quality edits above/below, all behavior-neutral.
/code-review (high) re-surfaced the same stable finding set (the substantive code was unchanged; no coder commits landed between rounds). I acted on the most actionable one this turn:
- FIXED (NVIDIA#3, test robustness): the CUDA-graph hard-path tests required TLLM_WORKER_USE_SINGLE_PROCESS=1 (their ENABLED config hard-asserts the in-process runner) but never set it, so they would spuriously fail/silently-skip when run with a checkpoint outside a harness that exports it. Added an autouse `monkeypatch.setenv` fixture to both integration files (test_chatglm3_gsm8k.py, test_chatglm3_replay.py). Verified the env is read at runtime (utils.py:380, os.environ.get inside a function) so the fixture takes effect before _build_llm; the fix aligns each test with its own documented requirement and cannot regress the already-validated path.
- Previously fixed (round 1): _pick_metric_key now fails loudly instead of gating on an arbitrary metric; consumed bias/head_dim inline derivation replaced with normalized fields; BFS runner-finder fallback removed.
RESIDUAL (documented follow-ups — not safely resolvable by the automated quality loop):
- Out-of-target latent correctness: (NVIDIA#1) load_weights only handles the chatglm3-6b/chatglm2-6b bias layout (add_qkv_bias=True/add_bias_linear=False) — it unconditionally fetches the QKV bias and never loads dense/MLP biases, so ChatGLM variants with add_qkv_bias=False (KeyError) or add_bias_linear=True (unloaded bias + unconsumed-weight ValueError) fail; (NVIDIA#2) rope_ratio is never mapped to rope_theta, so long-context chatglm3-6b-32k/128k (rope_ratio>1) silently use theta=10000. Both are no-ops for the bringup target (base chatglm3-6b/chatglm2-6b, rope_ratio=1); fixing needs variant checkpoints to test and expands scope beyond the bringup — for the coder/a follow-up.
- Larger altitude/reuse/efficiency refactors needing test-execution access: (NVIDIA#4/NVIDIA#7) adopt BaseWeightMapper/register_mapper for load_weights + a ChatGLMConfigLoader/config-subclass to remove the shared-runtime normalize special-case and its second call; (NVIDIA#5) consolidate the 4x/2x-duplicated test helpers across the two GPU-gated test trees; (NVIDIA#6) cache the 6B HF/TRT model in the parametrized attention test; (NVIDIA#8) the _GreedyHFLM hand-rolled reference decode.
APPROVE rationale: base chatglm3-6b is fully correct (QA-approved, acceptance criteria met); two thorough rounds applied all safe in-scope polish culminating in the NVIDIA#3 test-robustness fix; the quality loop has converged (round 2 reproduced round 1's findings with no coder changes); and every remaining finding is either out-of-target latent correctness for untested ChatGLM variants or a larger refactor requiring test-execution access — none of which another automated /simplify→/code-review round can safely resolve. Residuals are surfaced above for human/coder triage.
QA report:
DECISION: APPROVE. All 8 acceptance criteria hold at runtime against the current committed code (HEAD 2323424, iter11). Reference is the checkpoint's own trust_remote_code modeling_chatglm.py — an independent HF reference sharing zero code with the TRT-LLM path.
PER-CRITERION (runtime evidence):
1. Bootstrap chain — PASS. slurm/core_gpu_tests.4147161.log: "Built target build_wheel_targets", "Successfully installed tensorrt_llm-1.3.0rc21 transformers-5.5.4" (pip install -e .[devel]), "Claude Code successfully installed!", "claude-agent-sdk 0.2.93"; all inside srun pyxis container with --container-mounts=<repo>:<repo> --container-workdir=<repo> under set -eo pipefail (tests ran only because bootstrap rc=0).
2. config_registration_and_weight_accounting — PASS (TIER1 rc=0). Real ckpt loads unedited, arch resolves to ChatGLMForCausalLM, TRTLLM backend + KVCacheManagerV2, strict state-dict accounting (raises on missing/unexpected/shape-mismatch; only rotary_pos_emb.inv_freq tolerated).
3. source_activation_replay — PASS both graph modes. layer0 cosine=1.000000 mean_abs=0.00000, layer27 cosine=0.999999; decode graph-vs-eager max_abs=0.0 cosine=1.0. Companion partial_rope_boundary_and_theta PASS (dims [0:64] rotate, [64:128] pass-through, theta=10000^(-i/32)).
4. source_logit_replay — PASS both configs. trt_tok==hf_tok (5231/30910/4802), cosine~0.99999; enabled num_captured_graphs=8, baseline 0.
5. generation_parity — PASS both configs. 5 prompts x 32 steps, token_mismatches=0, min per-step cosine~0.99998.
6. llm_api_smoke — PASS both configs (nonempty deterministic; V2 + TRTLLM asserted; enabled enabled=True graphs=8, baseline graphs=0).
7. gsm8k accuracy_canary — PASS (slurm/gsm8k_gate.4145862.log): HF=40.00, baseline=40.00 gap 0.00, enabled=40.00 gap 0.00; enabled num_captured_graphs=32.
8. gsm8k full_trtllm_eval — PASS: HF=52.77, baseline=52.99 gap 0.23, enabled=53.07 gap 0.30 (both < 2.0 tol), enabled captured 32 CUDA graphs; RESULT: PASS. Absolute ~53% matches published ChatGLM3-6B GSM8K, ruling out a shared-bug artifact.
PROVENANCE: Core log 4147161 started 05:55, after the 05:33 HEAD commit -> definitively iter11 (criteria 1-6). git diff iter10->iter11 (2d58cc7..2323424) touched ONLY the two test files (additive CUDA-graph hard-path assertions); tensorrt_llm/ is byte-identical. GSM8K log 4145862 prints the iter11-only [cuda_graph_hard_path] output -> it ran the current test file against the unchanged runtime code (live bind-mount imported iter11 mid-bootstrap). Criteria 7-8 are therefore verified against current code.
INDEPENDENT RERUN THIS TURN: I resubmitted both jobs. Core (4149172) matched the prior post-commit run and was cancelled as redundant. My GSM8K resubmit (4149173) FAILED in ~1.5 min inside build_wheel.py (FileNotFoundError: inherit_graph_6.md5 in copy_resolving_symlink) — a concurrent-build race from my own simultaneous submission of two bootstrap-heavy jobs on one live repo mount, NOT a ChatGLM defect. Accepted evidence is the sequential post-commit runs + full code/provenance review.
RED-TEAM (all addressed): independent HF reference (no shared helpers); partial GPT-J RoPE covered analytically + behaviorally (layer0 exact); CUDA-graph silent fallback ruled out by in-process CUDAGraphRunner introspection (runner.enabled + >=1 torch.cuda.CUDAGraph; baseline 0); KV-cache/mask/scale drift ruled out by 32-step parity (0 mismatches) + GSM8K absolute-score match; transformers 5.x bridges (all_tied_weights_keys={}, max_length, num_hidden_layers) are loader-compat only and do not alter ChatGLM forward semantics.
EVALUATION SCORES:
- Functionality (2.0): 9 — every criterion passes with exact argmax parity, cosine >0.9999, GSM8K within 0.3 pts of HF on both configs.
- Code Quality (1.0): 8 — clean, well-documented 373-line model reusing existing modules; narrow config hook; strict weight accounting; minor incidental import reformatting in __init__.py; a pre-existing (non-diff) ruff E501 at __init__.py:130.
- Performance (1.5): 8 — parity-first bring-up (no perf gate in task.yaml); production TRTLLM backend + CUDA graph + overlap scheduler + fused QKV/gate-up + KVCacheManagerV2 all working.
- Completeness (1.0): 9 — complete, buildable, no placeholders, full state-dict accounting, both runtime configs.
- Technical Sophistication (1.0): 9 — correct unfused partial GPT-J RoPE with TRTLLM backend + CUDA-graph capture, compact MQA KV, config/transformers bridges.
Weighted = (9x2.0 + 8x1.0 + 8x1.5 + 9x1.0 + 9x1.0)/6.5 = 56/6.5 = 8.6.
STRENGTHS: genuine independent HF reference; real CUDA-graph hard-path proof (not just config flags); strict weight accounting; large GSM8K margin; minimal composable design (no new schemas). WEAKNESSES: multi-GPU/perf/chunked-prefill deferred (out of scope per task.yaml); harness must run the two jobs sequentially to avoid the concurrent-build race; pre-existing E501 worth a one-line fix. RECOMMENDATION: APPROVE — no code defect identified; task.yaml completion_criteria (GSM8K within 2 pts with and without CUDA graph + overlap scheduler) is met on both configs.</summary>
</invoke>
limin2021
added a commit
to limin2021/TensorRT-LLM
that referenced
this pull request
Jul 8, 2026
…e dead tiled_copy chain Finding NVIDIA#2 from code review: the GVR-style GMEM load path in filtered_topk_kernel_per_row was using CopyUniversalOp (no .CONSTANT cache hint) instead of CopyG2ROp(invariant=True) (.CONSTANT hint, better for invariant/read-only loads). The CopyG2ROp copy_atom from _get_tiled_copy() was passed all the way down the call chain but never reached the actual cute.copy call. Fix: - Change _copy_atom in filtered_topk_kernel_per_row to use CopyG2ROp(invariant=True) for .CONSTANT cache hint on GMEM loads. - Remove tiler_mn, copy_atom, tiled_copy dead parameters from filtered_topk_kernel_per_row, filtered_topk_kernel (prefill + decode), and run_kernel (decode). None of these were used inside the kernel bodies. - Delete _get_tiled_copy() helper (was only used to supply these dead params; tiled_copy.size replaced by self.num_threads_per_cta in launchers). - Delete run_kernel() (dead code: defined but never called). Signed-off-by: Mindy Li <11663212+limin2021@users.noreply.github.com>
1 task
chenfeiz0326
added a commit
to chenfeiz0326/TensorRT-LLM
that referenced
this pull request
Jul 17, 2026
Reduces the pre-merge disagg PerfSanity set from 8 tests to 4 by moving these back to post-merge: * NVIDIA#2 gpt-oss-120b-fp4 TEP4 gen_only 1k1k con64 (GB200) * NVIDIA#4 deepseek-r1-fp4 TEP8 gen_only 1k1k con1 (GB200) * NVIDIA#5 kimi-k25-thinking-fp4 DEP8 e2e 1k1k con4096 (GB200) * NVIDIA#6 glm-5-fp4 TEP4 gen_only 1k1k con1 (GB300) Remaining 4 pre-merge FUNCTIONAL-ONLY stages: * gpt-oss-120b-fp4 DEP2 e2e 1k1k con2048 (GB200, 2n/8g) * kimi-k25-thinking-fp4 TEP4 gen_only 1k1k con4 (GB200, 2n/8g) * glm-5-fp4 DEP8 e2e 1k1k con4096 (GB300, 3n/12g) * deepseek-r1-fp4 DEP8 gen_only 1k1k con2048 (B200, 2n/16g) Corresponding Jenkins adjustments: 2 GB200 pre-merge stages deleted, 2 GB200 Post-Merge testCounts bumped, 1 GB300 pre-merge stage renamed back to Post-Merge. Signed-off-by: Chenfei Zhang <chenfeiz@nvidia.com>
chenfeiz0326
added a commit
to chenfeiz0326/TensorRT-LLM
that referenced
this pull request
Jul 17, 2026
Follow-up: user requires each pre-merge pick to already be actively
running in L0 CI post-merge on origin/main — i.e. no QA-lane
borrowing, no uncommenting of previously-disabled lines, no new
config yamls, no first-time-in-L0 e2e/gen_only combinations.
Prior revision's two picks that failed that bar:
* gpt-oss 8k1k con512 e2e dep2 — active in QA lane only; the
e2e mode of this 8k1k dep2 config was not in L0 anywhere.
* kimi 8k1k con4 tep8 mtp3 gen_only GB200 — was commented out
in the GB200 12-GPU 3-node post_merge block since April.
Replaced with:
* gpt-oss 8k1k con1024 tp4 e2e MTP0 — was active in L0 post-merge
on GB200 ctx1_node1_gpu1_gen1_node1_gpu4 (8 GPUs, 2 nodes).
Gen-worker parallelism shifts dep→tp for this slot.
* kimi 8k1k con4 tep8 mtp3 gen_only — was active in L0 post-merge
on GB300 ctx1_node1_gpu4_gen1_node2_gpu8 (12 GPUs, 3 nodes).
GPU family shifts GB200→GB300 for this slot.
File changes:
- l0_gb200_multi_nodes_perf_sanity_ctx1_node1_gpu1_gen1_node1_gpu2.yml:
reverted to origin/main state (no pre_merge block).
- l0_gb200_multi_nodes_perf_sanity_ctx1_node1_gpu1_gen1_node1_gpu4.yml:
new pre_merge block with gpt-oss 8k1k con1024 e2e; test line
removed from post_merge to avoid double-run.
- l0_gb200_multi_nodes_perf_sanity_ctx1_node1_gpu4_gen1_node2_gpu8.yml:
reverted to origin/main state (kimi 8k1k stays commented out).
- l0_gb300_multi_nodes_perf_sanity_ctx1_node1_gpu4_gen1_node2_gpu8.yml:
new pre_merge block with kimi 8k1k con4 tep8 mtp3 gen_only; same
test removed from post_merge.
Jenkins:
- Dropped stages: FUNCTIONAL-ONLY-CTX1-NODE1-GPU1-GEN1-NODE1-GPU2
(GB200) and FUNCTIONAL-ONLY-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8
(GB200 kimi variant).
- New stages: FUNCTIONAL-ONLY-CTX1-NODE1-GPU1-GEN1-NODE1-GPU4
(GB200 gpt-oss) and FUNCTIONAL-ONLY-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8
(GB300 kimi variant).
- Post-merge decrement: GB200-8_GPUs-2_Nodes CTX1-NODE1-GPU1-GEN1-
NODE1-GPU4 7→6; GB300-12_GPUs-3_Nodes CTX1-NODE1-GPU4-GEN1-NODE2-
GPU8 5→4. Others revert to origin/main testCounts.
Shape matrix (all 4 pre-merge slots now at 8k1k):
NVIDIA#1 DSR1 B200 8k1k con1536 dep e2e=NO gen_only MTP✓
NVIDIA#2 gpt-oss GB200 8k1k con1024 tp e2e=YES MTP✗
NVIDIA#3 glm-5 GB300 8k1k con1024 dep e2e=YES MTP✓
NVIDIA#4 kimi GB300 8k1k con4 tep e2e=NO gen_only MTP✓
Signed-off-by: Chenfei Zhang <chenfeiz@nvidia.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.