[TRTLLM-13639][perf] Migrate Kimi perf-sanity tests to Transceiver v2 - #15966
Conversation
96efe46 to
588936c
Compare
|
/bot run --disable-fail-fast --stage-list "GB200PerfSanity*,GB300PerfSanity*" |
📝 WalkthroughWalkthroughFourteen disaggregated perf-sanity YAML test configuration files (GB200 and GB300 variants) are updated to explicitly add ChangesPerf-sanity YAML config updates
Estimated code review effort: 1 (Trivial) | ~5 minutes 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
|
PR_Github #57724 [ run ] triggered by Bot. Commit: |
|
PR_Github #57724 [ run ] completed with state
|
|
/bot run --disable-fail-fast --stage-list "GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU4-Post-Merge" |
|
PR_Github #58216 [ run ] triggered by Bot. Commit: |
|
PR_Github #58216 [ run ] completed with state
|
|
/bot run --disable-fail-fast --stage-list "GB200-4_GPUs-PyTorch-PerfSanity-Post-Merge-7" |
|
PR_Github #58224 [ run ] triggered by Bot. Commit: |
|
PR_Github #58224 [ run ] completed with state |
0837be9 to
fc8ad13
Compare
|
/bot run --disable-fail-fast --stage-list "GB200PerfSanity*,GB300PerfSanity*" |
|
PR_Github #58228 [ run ] triggered by Bot. Commit: |
|
PR_Github #58228 [ run ] completed with state
|
fc8ad13 to
a76a77a
Compare
|
/bot run --disable-fail-fast --stage-list "GB200PerfSanity*,GB300PerfSanity*" |
|
PR_Github #58421 [ run ] triggered by Bot. Commit: |
|
PR_Github #58421 [ run ] completed with state
|
|
/bot run --disable-fail-fast --stage-list "GB200-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1,GB200-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-2,GB200-36_GPUs-9_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE8-GPU32-Post-Merge-11,GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU4-Post-Merge-5,GB200-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-4,GB300-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU4-Post-Merge-3,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-5" |
|
PR_Github #58570 [ run ] triggered by Bot. Commit: |
|
PR_Github #58570 [ run ] completed with state
|
|
/bot run --disable-fail-fast --stage-list "GB200-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1,GB200-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-2,GB200-36_GPUs-9_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE8-GPU32-Post-Merge-11,GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU4-Post-Merge-5,GB200-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-4,GB300-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU4-Post-Merge-3,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-5" |
|
PR_Github #59172 [ run ] triggered by Bot. Commit: |
|
PR_Github #59172 [ run ] completed with state |
ef44752 to
2a52cb6
Compare
…-bounce size The 12 gb200/gb300 Kimi-K2.5 disaggregated perf-sanity cases, flipped to the Python (v2) KV-cache transceiver with a 384 MiB KV-bounce arena, for measuring the v1(C++)-vs-v2(Python) transceiver gap. Bounce min_blocks is env-gated (TRTLLM_KV_CACHE_BOUNCE_MIN_BLOCKS); default keeps bounce off on the tiny con4 contexts, set =1 to force it on. Signed-off-by: Shixiaowei02 <39303645+Shixiaowei02@users.noreply.github.com>
2a52cb6 to
f0179c8
Compare
|
/bot run --disable-fail-fast --stage-list "GB200-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1,GB200-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-2,GB200-36_GPUs-9_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE8-GPU32-Post-Merge-11,GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU4-Post-Merge-5,GB200-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-4,GB300-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU4-Post-Merge-3,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-5" |
|
PR_Github #59634 [ run ] triggered by Bot. Commit: |
|
PR_Github #59634 [ run ] completed with state |
|
/bot skip --comment "Related tests have passed." |
|
PR_Github #59661 [ skip ] triggered by Bot. Commit: |
|
PR_Github #59661 [ skip ] completed with state |
….5 con4096 disagg perf case
The gen_only case timed out on the CTX side: 2939 "KV cache transfer timeout:
elapsed ~60244ms > kv_transfer_timeout_ms=60000ms" lines across CTX ranks 0-3 and
4096/4096 failed requests.
Root cause is the data plane, not the timeout budget. With bounce disabled, each
8k request submits its KV as ~15.6k scattered per-block fragments in a single
NIXL transfer request, which UCX cannot stripe across rails, so transfers queue
behind one another until the 60s budget expires. Coalescing the request's blocks
into one contiguous fabric-VMM buffer restores a single multi-rail write.
Reproduced and verified locally on GB300 (20 GPUs / 5 nodes, gen_only), both arms
on the same commit, wheel and container:
as-is : FAILED - 0/4096 successful, 2879 CTX transfer timeouts
+ bounce : PASSED - 4096/4096 successful, 0 timeouts, 11580 tok/s output
per-transfer latency p50 2.16 ms / p90 2.99 ms
"[kv-bounce] coalesced 256 blocks / 274MiB into one region across 1 writer(s)"
Bounce engages on the single-writer path here because attention DP is enabled on
both sides, so each request has exactly one peer rank (expected_transfers == 1).
This mirrors the settings NVIDIA#15966 already applied to the other Kimi disagg cases.
The timeout budget is left at the 60s default on purpose: coalesced transfers
complete in ~2 ms, four orders of magnitude inside it.
Un-waives only the gen_only case, which is the one covered by this bug and the
one verified above. The e2e variant stays waived pending its own validation.
Signed-off-by: Shixiaowei02 <39303645+Shixiaowei02@users.noreply.github.com>
….5 con4096 disagg perf case
The gen_only case timed out on the CTX side: 2939 "KV cache transfer timeout:
elapsed ~60244ms > kv_transfer_timeout_ms=60000ms" lines across CTX ranks 0-3 and
4096/4096 failed requests.
Root cause is the data plane, not the timeout budget. With bounce disabled, each
8k request submits its KV as ~15.6k scattered per-block fragments in a single
NIXL transfer request, which UCX cannot stripe across rails, so transfers queue
behind one another until the 60s budget expires. Coalescing the request's blocks
into one contiguous fabric-VMM buffer restores a single multi-rail write.
Reproduced and verified locally on GB300 (20 GPUs / 5 nodes, gen_only), both arms
on the same commit, wheel and container:
as-is : FAILED - 0/4096 successful, 2879 CTX transfer timeouts
+ bounce : PASSED - 4096/4096 successful, 0 timeouts, 11580 tok/s output
per-transfer latency p50 2.16 ms / p90 2.99 ms
"[kv-bounce] coalesced 256 blocks / 274MiB into one region across 1 writer(s)"
Bounce engages on the single-writer path here because attention DP is enabled on
both sides, so each request has exactly one peer rank (expected_transfers == 1).
This mirrors the settings NVIDIA#15966 already applied to the other Kimi disagg cases.
The timeout budget is left at the 60s default on purpose: coalesced transfers
complete in ~2 ms, four orders of magnitude inside it.
Un-waives only the gen_only case, which is the one covered by this bug and the
one verified above. The e2e variant stays waived pending its own validation.
Signed-off-by: Shixiaowei02 <39303645+Shixiaowei02@users.noreply.github.com>
….5 con4096 disagg perf case
The gen_only case timed out on the CTX side: 2939 "KV cache transfer timeout:
elapsed ~60244ms > kv_transfer_timeout_ms=60000ms" lines across CTX ranks 0-3 and
4096/4096 failed requests.
Root cause is the data plane, not the timeout budget. With bounce disabled, each
8k request submits its KV as ~15.6k scattered per-block fragments in a single
NIXL transfer request, which UCX cannot stripe across rails, so transfers queue
behind one another until the 60s budget expires. Coalescing the request's blocks
into one contiguous fabric-VMM buffer restores a single multi-rail write.
Reproduced and verified locally on GB300 (20 GPUs / 5 nodes, gen_only), both arms
on the same commit, wheel and container:
as-is : FAILED - 0/4096 successful, 2879 CTX transfer timeouts
+ bounce : PASSED - 4096/4096 successful, 0 timeouts, 11580 tok/s output
per-transfer latency p50 2.16 ms / p90 2.99 ms
"[kv-bounce] coalesced 256 blocks / 274MiB into one region across 1 writer(s)"
Bounce engages on the single-writer path here because attention DP is enabled on
both sides, so each request has exactly one peer rank (expected_transfers == 1).
This mirrors the settings NVIDIA#15966 already applied to the other Kimi disagg cases.
The timeout budget is left at the 60s default on purpose: coalesced transfers
complete in ~2 ms, four orders of magnitude inside it.
Un-waives only the gen_only case, which is the one covered by this bug and the
one verified above. The e2e variant stays waived pending its own validation.
Signed-off-by: Shixiaowei02 <39303645+Shixiaowei02@users.noreply.github.com>
Description
This pull request updates several performance sanity test YAML configurations for disaggregated worker setups, primarily to enhance the
cache_transceiver_configsection. The main changes introduce a new runtime and adjust memory settings for the cache transceiver, which will impact how cache data is managed and transferred during tests.Cache transceiver configuration updates:
transceiver_runtime: PYTHONto all affected YAML files, switching the cache transceiver to use the Python runtime for improved flexibility and compatibility. [1] [2] [3] [4] [5] [6] [7] [8] [9] [10] [11] [12] [13] [14] [15] [16] [17] [18] [19] [20] [21] [22] [23] [24]kv_cache_bounce_size_mb: 384to selected YAML files, increasing the key-value cache bounce buffer size to 384 MB for improved cache transfer performance in relevant test scenarios. [1] [2] [3] [4] [5] [6] [7] [8]No other functional or logic changes were made; these updates are limited to configuration improvements for the test scripts.
Test Coverage
PR Checklist
Please review the following before submitting your PR:
PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.
PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.
Test cases are provided for new code paths (see test instructions)
If PR introduces API changes, an appropriate PR label is added - either
api-compatibleorapi-breaking. Forapi-breaking, includeBREAKINGin the PR title.Any new dependencies have been scanned for license and vulnerabilities
CODEOWNERS updated if ownership changes
Documentation updated as needed
Update tava architecture diagram if there is a significant design change in PR.
The reviewers assigned automatically/manually are appropriate for the PR.
Please check this after reviewing the above items as appropriate for this PR.
GitHub Bot Help
To see a list of available CI bot commands, please comment
/bot help.Summary by CodeRabbit