Skip to content

[NVBUG-6448152][test] TEST ONLY; DO NOT REVIEW: current-runtime source midpoint - #16756

Closed
chienchunhung wants to merge 2 commits into
NVIDIA:mainfrom
chienchunhung:codex/nvbug-6448152-current-runtime-midpoint
Closed

[NVBUG-6448152][test] TEST ONLY; DO NOT REVIEW: current-runtime source midpoint#16756
chienchunhung wants to merge 2 commits into
NVIDIA:mainfrom
chienchunhung:codex/nvbug-6448152-current-runtime-midpoint

Conversation

@chienchunhung

Copy link
Copy Markdown
Collaborator

TEST ONLY; DO NOT REVIEW

This draft is a bounded midpoint-checkpoint discriminator for the remaining
current-tree throughput regression. It is not a proposed product change.

What this isolates

  • Product checkpoint: legacy C++ module removal checkpoint.
  • Treatment: the byte-identical asynchronous-consensus factor from the
    official C++ asynchronous-consensus change.
  • Runtime: unchanged current image build 202607151440-16194 for x86 and SBSA.
  • Workload: unchanged exact GB300 PP4 perf-sanity YAML and selector.
  • CI mergeability: current main is recorded only as a no-tree parent; the factor
    tree is unchanged.

The checkpoint is exactly halfway through the 92-commit source interval that
already uses this current runtime image: 46 commits precede it and 46 follow it.
The current image tags, both exact workload YAML blobs, and every preimage used
by the asynchronous-consensus factor match the current official base. There is
no compatibility patch, timeout change, workload change, or semantic adaptation.

Interpretation

  • At least 1402 output tokens/s establishes a fast endpoint under the current
    runtime and localizes the regression to the 46 later source commits.
  • Approximately 800 output tokens/s shows that the slowdown was already
    present at this checkpoint and moves the next discriminator into the earlier
    half.
  • Any partial recovery will be reported relative to the current 794.61
    output tokens/s result and the historical 1548.84 output tokens/s result.
  • A build or infrastructure failure is inconclusive and will not be interpreted
    as performance evidence.

Verification

  • Exact factor diff SHA-256:
    25e394aac54636e88433755595813bd89704cfa2861bd9167303a7beb9694867.
  • The factor commit and no-tree merge commit have the same tree:
    33cc3f639ff6b4f1415e919c4c3fd8d2b9a2c718.
  • git diff --check, clang-format, CMake format, codespell, DCO, and the
    applicable pre-commit checks passed.
  • Targeted run pending for
    GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1.

Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
@chienchunhung

Copy link
Copy Markdown
Collaborator Author

/bot run --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61099 [ run ] triggered by Bot. Commit: dec24cf Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61099 [ run ] completed with state FAILURE. Commit: dec24cf
/LLM/main/L0_MergeRequest_PR pipeline #49353 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@chienchunhung

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61103 [ run ] triggered by Bot. Commit: dec24cf Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61103 [ run ] completed with state SUCCESS. Commit: dec24cf
/LLM/main/L0_MergeRequest_PR pipeline #49357 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants