[NVBUG-6448152][test] TEST ONLY; DO NOT REVIEW: earlier C++ source discriminator - #16837
Conversation
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
|
/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1" |
|
PR_Github #61549 [ run ] triggered by Bot. Commit: |
|
PR_Github #61549 [ run ] completed with state
|
|
Terminal TEST ONLY result for the single-stage run:
This is +0.70% versus the later current-source 794.61 tok/s result, +0.62% versus the preceding 795.22 tok/s discriminator, and 51.66% of the historical 1548.84 tok/s result. There is no meaningful recovery, so the slowdown is already present at or before native checkpoint Closing this TEST ONLY draft. The branch is preserved. |
Purpose
TEST ONLY; DO NOT REVIEW OR MERGE.
This diagnostic narrows the separate later-source C++ GB300 context-pipeline forward/device-loop throughput regression while holding the current user-space runtime and exact workload fixed.
The merged asynchronous-consensus implementation fixes the earlier blocking pipeline rendezvous. On the matched historical product tree it measured 1548.84 output tokens/s versus 1557.83 for its control. The most recent closed source discriminator completed 512/512 at 795.22 output tokens/s, placing the separate slowdown at or before native source
52ae70934f4357d4aa8e0b3c6ed87f83d53d3149.Exact construction
274043a75a53441b98a057321c75ae9215dbf44b, tree6d51ad3d7bea310ff3db79917ddee35be0e7ccc2.52ae7093.1e9400f65b8c81999f98b16e5ef77819a717a692and5ba029d0bcd4f1a1cba41f4711139a7e8d1ff5a5.228bc57299bab8910d2f03fb0ba9527e21947a9e.549e0085a06212cde0c47ab47072ad6acdcd1f58, exactly matching the merged implementation.d0bfb69a84356be6a24f215c0f5a76cca74ca8f7, treed650cc86d5cd44937d7545dc03c9a65677ec743e.75b39d4368204267e70fc3daadaf1820d1fc99edis included only as the second parent of the signed CI-only merge commit. The merge changes ancestry, not the diagnostic tree.The combined native-to-head diff changes exactly the eleven asynchronous-consensus files. It is byte-identical to a direct application of the factor.
Fixed runtime and workload
jenkins/current_image_tags.properties:504fd8f234a8d10076e82f000d35e41507566593.795a6192bf410b35d67535a1b85ce44e42398905.TIMEOUT (180).Target selector:
disagg_upload-e2e-gb300_deepseek-r1-fp4_128k8k_con256_ctx1_pp4_gen1_dep8_eplb0_mtp1_ccb-NIXLTarget stage:
GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1No runtime image, workload, scheduler setting, transport setting, traffic shape, timeout, or production consensus behavior is intentionally changed.
Interpretation
274043a7.Acceptance
Interpret the run only after verifying:
--no-container-mount-home;This draft will be closed promptly after terminal evidence is captured.