[NVBUG-6448152][test] Recover token throughput for DS PP4 tests (DO NOT REVIEW YET) - #16806
Conversation
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
|
/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1" |
|
PR_Github #61356 [ run ] triggered by Bot. Commit: |
|
PR_Github #61356 [ run ] completed with state
|
|
/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1" |
|
PR_Github #61372 [ run ] triggered by Bot. Commit: |
|
PR_Github #61372 [ run ] completed with state
|
Purpose
TEST ONLY; DO NOT REVIEW.
Narrow the source interval for the later C++ GB300 context-pipeline forward/device-loop throughput regression while holding the current user-space runtime and exact workload fixed.
The official asynchronous-consensus implementation fixes the earlier blocking pipeline rendezvous. It recovers the historical frozen tree to 1548.84 output tokens/s, but its current review head measures 794.61 output tokens/s. This diagnostic asks whether the separate later slowdown is already present at an earlier native source checkpoint.
Exact construction
52ae70934f4357d4aa8e0b3c6ed87f83d53d3149, tree3638dd83cc9c4ef9b7e79eb30eef98957a633622.ec93e4860ab32273a3003e1b8838b1f0f9daf86a.DSparkWorker.forwardto_forward_impland add the corresponding test stub.79aed56a23240a65ec02e9f95888712bdd8b5af1.68acb46c159989eb47bca824fffbd1f222085ba3.549e0085a06212cde0c47ab47072ad6acdcd1f58, identical to the official implementation.c7d38a45e30665639b450d9ff8b5428551b31024, tree1dd9e358f1b641a42091b2db2cfea17aec6b84bd.Fixed runtime and workload
The candidate preserves these verified blobs:
jenkins/current_image_tags.properties:504fd8f234a8d10076e82f000d35e41507566593.tests/scripts/perf-sanity/disaggregated/gb300_deepseek-r1-fp4_128k8k_con256_ctx1_pp4_gen1_dep8_eplb0_mtp1_ccb-NIXL.yaml:795a6192bf410b35d67535a1b85ce44e42398905.Target selector:
disagg_upload-e2e-gb300_deepseek-r1-fp4_128k8k_con256_ctx1_pp4_gen1_dep8_eplb0_mtp1_ccb-NIXLTarget stage:
GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1No runtime image, workload, scheduler setting, transport setting, timeout, traffic shape, or production consensus behavior is intentionally changed.
Interpretation
This checkpoint is eleven first-parent commits earlier than the verified-slow native checkpoint
3ab71753fe71d8ea4ae73427d1632fe2f7ecf201.52ae7093.Acceptance
Interpret the run only after verifying:
--no-container-mount-home;The draft will be closed after terminal evidence is captured. It is not intended for review or merge.