Skip to content

[NVBUG-6448152][test] TEST ONLY earlier native-source discriminator - #16843

Closed
chienchunhung wants to merge 5 commits into
NVIDIA:mainfrom
chienchunhung:codex/nvbug-6448152-native-midpoint-58d8964d
Closed

[NVBUG-6448152][test] TEST ONLY earlier native-source discriminator#16843
chienchunhung wants to merge 5 commits into
NVIDIA:mainfrom
chienchunhung:codex/nvbug-6448152-native-midpoint-58d8964d

Conversation

@chienchunhung

Copy link
Copy Markdown
Collaborator

TEST ONLY — do not merge.

This is one fixed-current-runtime source discriminator for the separate later C++ CTX PP forward/device-loop throughput regression. The asynchronous-consensus product change is already merged and is not under review here.

One-factor construction

  • Native source: 58d8964d131a163dc91c871afbf85785857594f3
  • Native source tree: 3455e1542c4f5a0ee47e7439f811db059f3c5c1f
  • Native source parent: 602718366c40e7078cec1d520c531a4fd119ccc2
  • This source is 53 first-parent commits earlier than the valid slow 274043a7 checkpoint from the preceding discriminator.
  • Current-runtime overlay is limited to the NIXL installer/development pins and all five image tags using build 202607151440-16194; stable patch ID: 7e73140158499de670b40490317c6f77601d06fc.
  • The official asynchronous-consensus factor is byte-identical by stable patch ID: 549e0085a06212cde0c47ab47072ad6acdcd1f58.
  • One CMake list-order normalization was restored after applying the factor and has no net tree effect beyond adding the factor source.
  • The native five-way stage mapping, workload YAML, selector, and timeout are unchanged.
  • Current main is present only as a no-tree second parent for CI mergeability.

Interpretation

  • At least 1402 output tokens/s places the regression in the following 53 first-parent commits through 274043a7.
  • Near 800 output tokens/s places the regression at or before this checkpoint and moves the source bound earlier.
  • Any partial recovery will be quantified.
  • A build/import failure, timeout, or failed request is compatibility/censored evidence and is not a throughput result.

Requested evidence

Run only GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1, requiring the exact selector, three nodes, twelve tasks, four GPUs per node, --no-container-mount-home, 512/512 successful requests, coordinator activation on all four CTX PP ranks, clean shutdown, and a valid official throughput metric.

The draft will be closed after terminal evidence is captured; the branch will be preserved.

Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61606 [ run ] triggered by Bot. Commit: c1c1671 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61606 [ run ] completed with state FAILURE. Commit: c1c1671
/LLM/main/L0_MergeRequest_PR pipeline #49816 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@chienchunhung

Copy link
Copy Markdown
Collaborator Author

Terminal discriminator evidence for exact head c1c1671a4512781fa4aedfa4c0dc7e47e965db3f / native source 58d8964d131a163dc91c871afbf85785857594f3:

  • Exact SBSA child / Slurm 2670797 ran the requested DeepSeek GB300 selector only, on three nodes / twelve tasks / four GPUs per node with --no-container-mount-home.
  • Result artifact records 512/512 successful requests, 0 failed requests, 0.10 req/s, 801.91 output tok/s, and 13632.41 total tok/s.
  • Logs show asynchronous context-transfer consensus enabled on CTX PP ranks 0, 1, 2, and 3. CTX, GEN, and disaggregated servers each reached complete application shutdown; no fatal, protocol, segmentation, or double-free errors were found.
  • Jenkins is red only because the intentionally historical source misses the expected performance gate; the benchmark itself is valid.

801.91 output tok/s is +0.92% versus the later-tree 794.61 result and 51.77% of the known-fast historical 1548.84 result. There is no meaningful recovery: the slowdown is present at or before native source 58d8964d, excluding the following 53 first-parent commits through 274043a7 from being its introduction interval. The next source discriminator must move earlier.

@chienchunhung

Copy link
Copy Markdown
Collaborator Author

Closing this terminal TEST ONLY discriminator after capturing the valid benchmark evidence above. The branch is preserved for audit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants