Skip to content

[None][infra] Run GB200 functional-only disagg perf-sanity on both oci-hsg and aws-dfw - #16619

Closed
chenfeiz0326 wants to merge 2 commits into
NVIDIA:mainfrom
chenfeiz0326:user/gb200-func-disagg-flex-split
Closed

[None][infra] Run GB200 functional-only disagg perf-sanity on both oci-hsg and aws-dfw#16619
chenfeiz0326 wants to merge 2 commits into
NVIDIA:mainfrom
chenfeiz0326:user/gb200-func-disagg-flex-split

Conversation

@chenfeiz0326

@chenfeiz0326 chenfeiz0326 commented Jul 20, 2026

Copy link
Copy Markdown
Collaborator

Description

The pre-merge GB200 2-node disaggregated perf-sanity stage
GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-FUNCTIONAL-ONLY-CTX1-NODE1-GPU1-GEN1-NODE1-GPU4
is functional-only — as the stage name and the surrounding comment note,
"perf regressions do not fail CI." Its only purpose is to verify that the
disaggregated serving path runs end-to-end.

It currently uses auto:gb200-flex, which pins it to the oci-hsg cluster.
oci-hsg is heavily loaded, and since this stage does not judge perf numbers,
there is no reason to keep it on a single cluster.

This PR switches the platform label to auto:gb200-flex-split so the stage can
also be scheduled on aws-dfw. Because the stage only checks functionality
(not perf), cross-cluster perf variance is irrelevant, and spreading it across
both clusters relieves pressure on the crowded oci-hsg cluster.

auto:gb200-flex-split is already used by the neighboring GB200 2-node
multi-node stages in the same file, so this reuses an existing, validated
placement for the same 8-GPU / 2-node topology.

Test Coverage

No new tests. This is a CI placement-only change to an existing pre-merge
functional-only stage; the stage's test content is unchanged. The stage itself
(GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-FUNCTIONAL-ONLY-...) exercises
the change by running on either oci-hsg or aws-dfw.

PR Checklist

  • Please check this after reviewing the above items as appropriate for this PR.

Summary by CodeRabbit

  • Chores
    • Updated the SBSA multi-node performance sanity test configuration to use the revised GB200 platform selection.

…i-hsg and aws-dfw

The pre-merge GB200 2-node disaggregated perf-sanity stage
(GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-FUNCTIONAL-ONLY-CTX1-NODE1-GPU1-GEN1-NODE1-GPU4)
is functional-only: perf regressions do not fail CI, it only verifies
that the disaggregated path runs. It currently uses "auto:gb200-flex",
which pins it to the oci-hsg cluster.

Switch it to "auto:gb200-flex-split" so it can also land on aws-dfw.
Because this stage only checks functionality (not perf numbers),
cross-cluster perf variance is irrelevant, and spreading it across both
clusters relieves pressure on the heavily-loaded oci-hsg cluster.

The "auto:gb200-flex-split" label is already used by the neighboring
GB200 2-node multi-node stages, so this reuses an existing, validated
placement for the same 8-GPU / 2-node topology.

Signed-off-by: Chenfei Zhang <chenfeiz@nvidia.com>
@chenfeiz0326
chenfeiz0326 requested a review from a team as a code owner July 20, 2026 10:12
@coderabbitai

coderabbitai Bot commented Jul 20, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 97411f0d-c743-40f3-be4c-cdae07fcbd44

📥 Commits

Reviewing files that changed from the base of the PR and between b8604c4 and 523d009.

📒 Files selected for processing (1)
  • jenkins/L0_Test.groovy

📝 Walkthrough

Walkthrough

The SBSA multi-node GB200 disaggregated perf-sanity Jenkins stage now resolves its platform using auto:gb200-flex-split instead of auto:gb200-flex.

Changes

GB200 platform configuration

Layer / File(s) Summary
Update stage platform selection
jenkins/L0_Test.groovy
The stage configuration changes its platform identifier to auto:gb200-flex-split.

Estimated code review effort: 1 (Trivial) | ~2 minutes

Suggested reviewers: dpitman-nvda

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly matches the change: moving the GB200 functional-only disagg perf-sanity stage to both oci-hsg and aws-dfw.
Description check ✅ Passed The description includes the required Description, Test Coverage, and PR Checklist sections and explains the change and rationale.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@chenfeiz0326

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-FUNCTIONAL-ONLY-CTX1-NODE1-GPU1-GEN1-NODE1-GPU4-1"

@chenfeiz0326
chenfeiz0326 requested a review from chzblych July 20, 2026 10:22
@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60352 [ run ] triggered by Bot. Commit: 523d009 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60352 [ run ] completed with state FAILURE. Commit: 523d009
/LLM/main/L0_MergeRequest_PR pipeline #48695 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@chenfeiz0326

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-FUNCTIONAL-ONLY-CTX1-NODE1-GPU1-GEN1-NODE1-GPU4-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60362 [ run ] triggered by Bot. Commit: 523d009 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60362 [ run ] completed with state FAILURE. Commit: 523d009
/LLM/main/L0_MergeRequest_PR pipeline #48703 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@chenfeiz0326

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-FUNCTIONAL-ONLY-CTX1-NODE1-GPU1-GEN1-NODE1-GPU4-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60368 [ run ] triggered by Bot. Commit: 523d009 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60368 [ run ] completed with state FAILURE. Commit: 523d009
/LLM/main/L0_MergeRequest_PR pipeline #48706 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@chenfeiz0326

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-FUNCTIONAL-ONLY-CTX1-NODE1-GPU1-GEN1-NODE1-GPU4-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60392 [ run ] triggered by Bot. Commit: 523d009 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60392 [ run ] completed with state FAILURE. Commit: 523d009
/LLM/main/L0_MergeRequest_PR pipeline #48730 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@chenfeiz0326

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-FUNCTIONAL-ONLY-CTX1-NODE1-GPU1-GEN1-NODE1-GPU4-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60532 [ run ] triggered by Bot. Commit: 523d009 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60532 [ run ] completed with state FAILURE. Commit: 523d009
/LLM/main/L0_MergeRequest_PR pipeline #48853 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

…s-dfw EFA

The GB200 functional-only disaggregated perf-sanity stage now also runs on
aws-dfw (via auto:gb200-flex-split). aws-dfw nodes use an Amazon EFA fabric,
where UCX auto-selects the gdr_copy transport for the NIXL KV-cache transfer.
gdr_copy fails on EFA with `gdr_copy_from_mapping failed. ret:22`, so the ctx
worker cannot send KV, the gen worker starves, and the test times out (verified
by reproducing on aws-dfw GB200 and via the PR NVIDIA#14933 cache-transceiver harness).

Set UCX_TLS=^gdr_copy for GB200 FUNCTIONAL-ONLY disagg stages so UCX keeps its
other transports and the KV transfer completes on EFA. The scope is deliberately
limited to FUNCTIONAL-ONLY stages: they only verify the disaggregated path runs
and do not judge perf, so the slower EFA fallback is acceptable. GB200 post-merge
perf-sanity stages (which share the same config yaml but run on oci-hsg over
NVLink/IB) keep `unset UCX_TLS` so their perf numbers are unaffected.

Note: this makes the functional-only stage PASS on aws-dfw but the EFA KV path is
host-staged (<1 GB/s); getting the disagg transfer onto NVLink/GPUDirect-RDMA is a
separate upstream NIXL/cache-transceiver item.

Signed-off-by: Chenfei Zhang <chenfeiz@nvidia.com>
@chenfeiz0326

Copy link
Copy Markdown
Collaborator Author

Why this PR now also touches jenkins/scripts/perf/submit.py

When I first flipped this functional-only disagg stage to auto:gb200-flex-split, CI failed once it actually landed on aws-dfw. I reproduced it on aws-dfw GB200 hardware and root-caused it end-to-end.

Root cause

aws-dfw GB200 nodes use an Amazon EFA fabric. The NIXL KV-cache transceiver runs over its UCX backend, and on GB200 the launch generator currently does unset UCX_TLS, so UCX auto-selects the gdr_copy transport. On EFA gdr_copy fails:

gdr_copy_md.c:205  UCX ERROR gdr_unpin_buffer failed. ret:22
gdr_copy_ep.c:92   UCX ERROR gdr_copy_from_mapping failed. ret:22
Fatal: mem type pack failed to uct_ep_get_short() Input/output error

The ctx worker can't send KV, the gen worker starves (num_scheduled_requests=0, pegged at the 5000 ms poll interval), and the test times out. On oci-hsg (NVLink/IB) this never happens — which is why the stage passed there before.

Fix (scoped to functional-only)

Set UCX_TLS=^gdr_copy for GB200 FUNCTIONAL-ONLY disagg stages only. This keeps the UCX backend but excludes the transport that's broken on EFA, so the KV transfer completes.

The scope is deliberately narrow. The same perf-sanity config also feeds a GB200 post-merge perf-sanity stage that runs on oci-hsg and does judge perf; those stages keep unset UCX_TLS so their numbers are untouched. Only functional-only stages — which just verify the disaggregated path runs — get the override, where the slower EFA fallback is acceptable.

Verification

On aws-dfw GB200 (real gpt-oss-120b disagg case):

  • Baseline (unset UCX_TLS): 0 completions — gdr_copy crash / KV-transfer timeouts.
  • UCX_TLS=^gdr_copy: completes (1054 requests served, gen worker generating normally, no gdr_copy errors).

I also verified the branch only triggers for GB200 + FUNCTIONAL-ONLY stage names and that GB200 post-merge/aggregated perf stages, GB300, and B200 stages are unchanged.

Known limitation / upstream follow-up

^gdr_copy makes the functional-only stage pass, but on EFA the KV path is host-staged (<1 GB/s) — it doesn't reach NVLink/GPUDirect-RDMA bandwidth. Using the PR #14933 cache-transceiver harness I confirmed that neither a fabric-memory KV pool (TRTLLM_KVCACHE_POOL_USE_FABRIC_MEMORY=1) nor same-NVL72-domain placement gets the full NIXL disagg transfer onto NVLink here — the transceiver always routes over EFA. Getting real GPUDirect-RDMA-over-EFA (or MNNVL cuda_ipc for the transfer itself) is a separate NIXL / cache-transceiver item; I'll file it upstream. For a functional-only pre-merge stage this is not blocking.

@chenfeiz0326

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-FUNCTIONAL-ONLY-CTX1-NODE1-GPU1-GEN1-NODE1-GPU4-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60905 [ run ] triggered by Bot. Commit: b078509 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #60905 [ run ] completed with state FAILURE. Commit: b078509
/LLM/main/L0_MergeRequest_PR pipeline #49172 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@chenfeiz0326

Copy link
Copy Markdown
Collaborator Author

Closing this PR — aws-dfw cannot run this disaggregated case (correcting an earlier claim)

After validating on real aws-dfw GB200 hardware, I'm closing this PR: moving this GB200 functional-only disaggregated stage to aws-dfw does not work, and no environment tweak fixes it.

Correction to my earlier comment

My earlier comment said UCX_TLS=^gdr_copy made the case "complete (1054 requests served)." That was a misread and is wrong. The gen-worker metric currank_total_requests = X/Y is completed / received; X (completed) was 0 in every iteration, and the POST /v1/completions 200 OK lines are the disagg server accepting requests, not the KV transfer completing. passed_test_list.txt was empty and the result was errors="1" ("Test terminated unexpectedly"). Apologies for the confusion.

What actually happens on aws-dfw (all runs → 0 completed requests)

The real disagg_upload-e2e-gb200_gpt-oss-120b-fp4_8k1k_con1024_...NIXL case was run to the benchmark on 2 same-NVL72-domain aws-dfw nodes under every candidate:

Config Completed reqs Symptom
default UCX 0 gdr_copy_from_mapping failed. ret:22 crash + KV-transfer timeouts
UCX_TLS=^gdr_copy 0 KV-transfer timeouts (no crash, but KV never transfers)
UCX_TLS=all UCX_RNDV_SCHEME=put_zcopy 0 KV-transfer timeouts
use_kv_cache_manager_v2: true 0 KV-transfer timeouts (no crash; V2 active)

Every config ends at 0 completed requests via the 60 s KV-transfer timeout. The PR #14933 cache-transceiver harness showed cuda_ipc/NVLink ~200 GB/s in a microbenchmark, but that path is not taken by the full NIXL disagg serve transceiver on aws-dfw — it routes over EFA and times out. So the root issue is the NIXL cross-node KV transfer over EFA itself, not the UCX TLS/rndv env or the KV-cache-manager version.

Decision

This disaggregated stage stays on auto:gb200-flex (oci-hsg only), which works. Enabling aws-dfw for it would require fixing the NIXL/UCX disagg KV transfer over EFA (an upstream cache-transceiver item), not a CI placement/env change. Closing this PR; happy to reopen if the EFA transfer path is fixed.

@chenfeiz0326

Copy link
Copy Markdown
Collaborator Author

Closing per the analysis above — aws-dfw EFA cannot run this NIXL disagg case; keeping it on oci-hsg (auto:gb200-flex).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants