[NVBUG-6448152][test] TEST ONLY earlier native-source discriminator - #16843
[NVBUG-6448152][test] TEST ONLY earlier native-source discriminator#16843chienchunhung wants to merge 5 commits into
Conversation
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
|
/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1" |
|
PR_Github #61606 [ run ] triggered by Bot. Commit: |
|
PR_Github #61606 [ run ] completed with state
|
|
Terminal discriminator evidence for exact head
801.91 output tok/s is +0.92% versus the later-tree 794.61 result and 51.77% of the known-fast historical 1548.84 result. There is no meaningful recovery: the slowdown is present at or before native source |
|
Closing this terminal TEST ONLY discriminator after capturing the valid benchmark evidence above. The branch is preserved for audit. |
This is one fixed-current-runtime source discriminator for the separate later C++ CTX PP forward/device-loop throughput regression. The asynchronous-consensus product change is already merged and is not under review here.
One-factor construction
58d8964d131a163dc91c871afbf85785857594f33455e1542c4f5a0ee47e7439f811db059f3c5c1f602718366c40e7078cec1d520c531a4fd119ccc2274043a7checkpoint from the preceding discriminator.202607151440-16194; stable patch ID:7e73140158499de670b40490317c6f77601d06fc.549e0085a06212cde0c47ab47072ad6acdcd1f58.mainis present only as a no-tree second parent for CI mergeability.Interpretation
274043a7.Requested evidence
Run only
GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1, requiring the exact selector, three nodes, twelve tasks, four GPUs per node,--no-container-mount-home, 512/512 successful requests, coordinator activation on all four CTX PP ranks, clean shutdown, and a valid official throughput metric.The draft will be closed after terminal evidence is captured; the branch will be preserved.