Skip to content

[NVBUG-6448152][test] TEST ONLY; DO NOT REVIEW lower compatible C++ midpoint - #16769

Closed
chienchunhung wants to merge 1 commit into
NVIDIA:mainfrom
chienchunhung:codex/nvbug-6448152-lower-compatible-midpoint
Closed

[NVBUG-6448152][test] TEST ONLY; DO NOT REVIEW lower compatible C++ midpoint#16769
chienchunhung wants to merge 1 commit into
NVIDIA:mainfrom
chienchunhung:codex/nvbug-6448152-lower-compatible-midpoint

Conversation

@chienchunhung

Copy link
Copy Markdown
Collaborator

Purpose

TEST ONLY; DO NOT REVIEW. Do not merge.

Run one fixed-current-runtime source discriminator for the separate later-tree C++ PP forward/device-loop throughput regression. The production async-consensus change remains in the official C++ review.

The prior compatible midpoint diagnostic completed 512/512 requests at 804.06 output tokens/s, showing that the dominant regression was already present at that checkpoint. This draft moves to the next audited lower native checkpoint while preserving every other tested factor.

Exact factor

  • Native source parent: 699e2277a4b8b159e67d6daed3085e20c050b62d
  • Diagnostic head: 02ef2f06104494d5aac7859d29cd3343c046a726
  • Product tree: 8e4a24b7ed69636f4f511be918495b96e5010e86
  • Async-consensus stable patch ID: 549e0085a06212cde0c47ab47072ad6acdcd1f58

All eight modified-file preimages match the official factor parent, all three added paths are absent at the native checkpoint, and all eleven postimages are byte-identical to the official async-consensus implementation.

The current image tags, exact GB300 stage YAML, exact workload YAML, and performance harness remain byte-identical to the current official baseline. The checkpoint also contains the native DSpark interface expected by that runtime.

Requested validation

Run only:

GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1

Exact selector:

disagg_upload-e2e-gb300_deepseek-r1-fp4_128k8k_con256_ctx1_pp4_gen1_dep8_eplb0_mtp1_ccb-NIXL

A valid result requires every submitted request to succeed, asynchronous coordinator activation on all four context pipeline ranks, clean shutdown, and an official output-token-throughput record. Any timeout, failed request, import failure, or build failure censors throughput.

Local verification

  • Stable patch identity and all preimage/postimage blobs verified.
  • Changed-line clang-format, CMake formatting, codespell, whitespace, test-list validation, DCO, and all applicable pre-commit hooks passed.
  • Independent read-only review found no P0/P1/P2 or compatibility issue.

No production or merge-readiness claim is made by this test-only pull request.

Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61170 [ run ] triggered by Bot. Commit: 02ef2f0 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61170 [ run ] completed with state FAILURE. Commit: 02ef2f0
/LLM/main/L0_MergeRequest_PR pipeline #49419 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants