Skip to content

[NVBUG-6448152][test] TEST ONLY bisect pre-admission native 85665f5f - #16928

Closed
chienchunhung wants to merge 8 commits into
NVIDIA:mainfrom
chienchunhung:codex/nvbug-6448152-native-pre-admission-85665f5f
Closed

[NVBUG-6448152][test] TEST ONLY bisect pre-admission native 85665f5f#16928
chienchunhung wants to merge 8 commits into
NVIDIA:mainfrom
chienchunhung:codex/nvbug-6448152-native-pre-admission-85665f5f

Conversation

@chienchunhung

Copy link
Copy Markdown
Collaborator

TEST ONLY

Historical-source/current-runtime discriminator for the separate C++ token-throughput regression. This draft is diagnostic only and is not intended for merge.

  • Native source checkpoint: 85665f5fd3, with exact native tree 8ada9734721bcd127722b367152eeb5bb22fa4ca and parent eaf5693b3aee9acea0d322dc8be997c4c0975280.
  • This checkpoint is the exact first-parent immediately before the admission and bounded-polling change, and 100 first-parent commits before the known-slow native checkpoint 9ff00c869a36055fcf2ebf4de4950a0471a3039f.
  • Admission-specific controller, requester timestamping, bounded polling, scheduler/configuration, and serialization changes are absent from the native tree. Current main is joined only as a no-tree second parent for CI ancestry; the exact checkout tree remains the audited first-parent tree.
  • Runtime-only overlay is held byte-identical to the completed slow discriminators: NIXL 1.3.1, aiperf 0.8.0, and the same five CI image tags from build 202607151440-16194.
  • CI-only compatibility holds the Jenkins harness and no-home-mount launch behavior byte-identical to those discriminators. A compile-only upstream Marlin empty-stub repair is also carried; the selected GB300 CUTEDSL workload does not execute Marlin.
  • Because the native checkpoint predates cancellation APIs, the reviewed asynchronous-consensus factor is a semantic pre-cancellation port. It runs protocol mode 2 with cancellation fixed off, retains the native blocking/finite-wait behavior, and is valid for categorical fast-versus-near-800 localization rather than small-percent attribution.
  • CMake context normalization and restoration have zero net ordering effect.
  • The exact workload YAML, test list, selector, performance baseline, performance test, and Slurm workload files are byte-identical to both the native checkpoint and completed slow discriminators. No DSpark, workload, selector, timeout, scheduler, or product compatibility adaptation is included.
  • The target remains the exact GB300 3-node/12-GPU disaggregated performance selector. A result at or above 1402 output tokens/s places the regression at the admission change or within the following 99 commits; a result near 800 proves the slowdown predates admission and moves the source bound earlier. Any build/import failure, timeout, protocol mismatch, cancellation, or failed request is non-interpretable.

The merged asynchronous-consensus change is not under re-review here; this draft only localizes the separate source throughput regression.

chienchunhung and others added 7 commits July 27, 2026 20:24
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
(cherry picked from commit 6fc7f33)
…sensus factor

Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
(cherry picked from commit 9c8bf59)
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
(cherry picked from commit 4b182f1)
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #62089 [ run ] triggered by Bot. Commit: b6ba1cb Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #62089 [ run ] completed with state FAILURE. Commit: b6ba1cb
/LLM/main/L0_MergeRequest_PR pipeline #50274 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

Copy link
Copy Markdown
Collaborator Author

Terminal result: censored before the benchmark

The exact historical source and sole requested GB300 stage were selected, but both prerequisite build families failed before any test child was created:

  • the standard builds exposed a DeepEP NVSHMEM header/static-library compatibility mismatch;
  • the LLVM builds treated an existing virtualMemory aggregate-initialization warning as an error.

There was no SBSA test child, Slurm allocation, pytest collection, serving request, or performance record. This run therefore provides no fast/slow evidence and does not move the source bound.

Both build failures are addressed by the later DLFW dependency compatibility change. Its relevant three-file compile/toolchain delta applies cleanly to this exact native checkpoint and introduces no admission-control or benchmark-path behavior. The next diagnostic will keep native source 85665f5fd331d0154a78172954846d843085e83f, the frozen runtime/harness, and the asynchronous-consensus factor unchanged, adding only that separately audited compatibility port.

Closing this TEST ONLY draft unmerged and preserving its branch.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants