Skip to content

[NVBUG-6448152][test] TEST ONLY native midpoint async-consensus discriminator - #16796

Closed
chienchunhung wants to merge 2 commits into
NVIDIA:mainfrom
chienchunhung:codex/nvbug-6448152-native-midpoint-3ab71753
Closed

[NVBUG-6448152][test] TEST ONLY native midpoint async-consensus discriminator#16796
chienchunhung wants to merge 2 commits into
NVIDIA:mainfrom
chienchunhung:codex/nvbug-6448152-native-midpoint-3ab71753

Conversation

@chienchunhung

Copy link
Copy Markdown
Collaborator

TEST ONLY — DO NOT REVIEW OR MERGE.

This draft is the sole C++ source discriminator for the separate later-main throughput regression. It does not change the production proposal in the official C++ async-consensus review.

Factor under test

  • Native source parent: 3ab71753
  • Treatment: the official async-consensus factor applied byte-for-byte
  • Stable patch ID: 549e0085a06212cde0c47ab47072ad6acdcd1f58
  • Current CI image, exact workload YAML, selector, harness, and timeouts are unchanged
  • No diagnostic instrumentation or additional product changes

This checkpoint is the audited compatible midpoint between the first current-runtime-compatible boundary and the measured-slow native checkpoint.

Required evidence

Run only:

GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1

Accept only an exact-head run with the exact selector, coordinator active on all four CTX PP ranks, every request successful, clean shutdown, and a valid official output-token metric. Any timeout or failed request censors throughput.

Interpretation

  • At least 1402 output tok/s: regression is later than this checkpoint.
  • About 800 output tok/s: regression is already present at or before this checkpoint.
  • Build/import failure: compatibility-only and not performance evidence.

Close this diagnostic immediately after terminal evidence; preserve the branch for reproducibility.

…iminator

Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61314 [ run ] triggered by Bot. Commit: 42cf8ff Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61314 [ run ] completed with state FAILURE. Commit: 42cf8ff
/LLM/main/L0_MergeRequest_PR pipeline #49545 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants