[NVBUG-6448152][test] TEST ONLY; DO NOT REVIEW: current-runtime source midpoint - #16756
[NVBUG-6448152][test] TEST ONLY; DO NOT REVIEW: current-runtime source midpoint#16756chienchunhung wants to merge 2 commits into
Conversation
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
|
/bot run --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1" |
|
PR_Github #61099 [ run ] triggered by Bot. Commit: |
|
PR_Github #61099 [ run ] completed with state
|
|
/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1" |
|
PR_Github #61103 [ run ] triggered by Bot. Commit: |
|
PR_Github #61103 [ run ] completed with state
|
TEST ONLY; DO NOT REVIEW
This draft is a bounded midpoint-checkpoint discriminator for the remaining
current-tree throughput regression. It is not a proposed product change.
What this isolates
official C++ asynchronous-consensus change.
202607151440-16194for x86 and SBSA.tree is unchanged.
The checkpoint is exactly halfway through the 92-commit source interval that
already uses this current runtime image: 46 commits precede it and 46 follow it.
The current image tags, both exact workload YAML blobs, and every preimage used
by the asynchronous-consensus factor match the current official base. There is
no compatibility patch, timeout change, workload change, or semantic adaptation.
Interpretation
1402output tokens/s establishes a fast endpoint under the currentruntime and localizes the regression to the 46 later source commits.
800output tokens/s shows that the slowdown was alreadypresent at this checkpoint and moves the next discriminator into the earlier
half.
794.61output tokens/s result and the historical
1548.84output tokens/s result.as performance evidence.
Verification
25e394aac54636e88433755595813bd89704cfa2861bd9167303a7beb9694867.33cc3f639ff6b4f1415e919c4c3fd8d2b9a2c718.git diff --check, clang-format, CMake format, codespell, DCO, and theapplicable pre-commit checks passed.
GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1.