[TRTLLM-13639][test] Migrate Kimi dis-agg tests to Transceiver v2 and trim tests - #16482
Conversation
8d16a99 to
853df7a
Compare
|
/bot run --disable-fail-fast |
📝 WalkthroughWalkthroughReplaces the Kimi-K2 NVFP4 accuracy test with Kimi-K2.5 coverage, updates disaggregated performance selections, removes obsolete benchmark variants, and explicitly configures the Python runtime for selected NIXL transceivers. ChangesKimi-K2.5 disaggregated serving
Estimated code review effort: 3 (Moderate) | ~20 minutes Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/integration/defs/accuracy/test_disaggregated_serving.py`:
- Around line 2233-2234: Update the test description/docstring associated with
the configurations near the NIXL backend and PYTHON transceiver_runtime entries
to state that both roles explicitly use NIXL with the Python runtime, rather
than the default transceiver. Apply the same wording update to the additional
location noted in the review.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: aeb230f1-3914-445e-a075-d9339380429a
📒 Files selected for processing (26)
tests/integration/defs/accuracy/test_disaggregated_serving.pytests/integration/test_lists/qa/llm_function_core.txttests/integration/test_lists/qa/llm_perf_disagg.ymltests/integration/test_lists/qa/llm_perf_multinode.txttests/integration/test_lists/test-db/l0_gb200_multi_gpus_perf_sanity.ymltests/integration/test_lists/test-db/l0_gb200_multi_nodes_perf_sanity_ctx1_node1_gpu4_gen1_node1_gpu4.ymltests/integration/test_lists/test-db/l0_gb200_multi_nodes_perf_sanity_ctx1_node1_gpu4_gen1_node2_gpu8.ymltests/integration/test_lists/test-db/l0_gb300_multi_gpus_perf_sanity.ymltests/integration/test_lists/test-db/l0_gb300_multi_nodes_perf_sanity_ctx1_node1_gpu4_gen1_node1_gpu4.ymltests/integration/test_lists/test-db/l0_gb300_multi_nodes_perf_sanity_ctx1_node1_gpu4_gen1_node2_gpu8.ymltests/scripts/perf/disaggregated/gb200_kimi-k25-thinking-fp4_1k1k_con2048_ctx1_dep4_gen1_dep32_eplb0_mtp0_ccb-NIXL.yamltests/scripts/perf/disaggregated/gb200_kimi-k25-thinking-fp4_1k1k_con4096_ctx1_dep4_gen1_dep8_eplb0_mtp0_ccb-NIXL.yamltests/scripts/perf/disaggregated/gb200_kimi-k25-thinking-fp4_1k1k_con4096_ctx1_dep4_gen1_dep8_eplb0_mtp0_ccb-UCX.yamltests/scripts/perf/disaggregated/gb200_kimi-k25-thinking-fp4_1k1k_con4_ctx1_dep4_gen1_tep4_eplb0_mtp0_ccb-NIXL.yamltests/scripts/perf/disaggregated/gb200_kimi-k25-thinking-fp4_8k1k_con1024_ctx1_dep4_gen1_dep32_eplb416_mtp3_ccb-NIXL.yamltests/scripts/perf/disaggregated/gb200_kimi-k25-thinking-fp4_8k1k_con4096_ctx1_dep4_gen1_dep16_eplb0_mtp0_ccb-NIXL.yamltests/scripts/perf/disaggregated/gb200_kimi-k25-thinking-fp4_8k1k_con4096_ctx1_dep4_gen1_dep16_eplb0_mtp0_ccb-UCX.yamltests/scripts/perf/disaggregated/gb200_kimi-k25-thinking-fp4_8k1k_con4_ctx1_dep4_gen1_tep8_eplb0_mtp3_ccb-NIXL.yamltests/scripts/perf/disaggregated/gb300_kimi-k25-thinking-fp4_1k1k_con2048_ctx1_dep4_gen1_dep32_eplb0_mtp0_ccb-NIXL.yamltests/scripts/perf/disaggregated/gb300_kimi-k25-thinking-fp4_1k1k_con4096_ctx1_dep4_gen1_dep8_eplb0_mtp0_ccb-NIXL.yamltests/scripts/perf/disaggregated/gb300_kimi-k25-thinking-fp4_1k1k_con4096_ctx1_dep4_gen1_dep8_eplb0_mtp0_ccb-UCX.yamltests/scripts/perf/disaggregated/gb300_kimi-k25-thinking-fp4_1k1k_con4_ctx1_dep4_gen1_tep4_eplb0_mtp0_ccb-NIXL.yamltests/scripts/perf/disaggregated/gb300_kimi-k25-thinking-fp4_8k1k_con1024_ctx1_dep4_gen1_dep32_eplb416_mtp3_ccb-NIXL.yamltests/scripts/perf/disaggregated/gb300_kimi-k25-thinking-fp4_8k1k_con4096_ctx1_dep4_gen1_dep16_eplb0_mtp0_ccb-NIXL.yamltests/scripts/perf/disaggregated/gb300_kimi-k25-thinking-fp4_8k1k_con4096_ctx1_dep4_gen1_dep16_eplb0_mtp0_ccb-UCX.yamltests/scripts/perf/disaggregated/gb300_kimi-k25-thinking-fp4_8k1k_con4_ctx1_dep4_gen1_tep8_eplb0_mtp3_ccb-NIXL.yaml
💤 Files with no reviewable changes (16)
- tests/integration/test_lists/test-db/l0_gb300_multi_gpus_perf_sanity.yml
- tests/scripts/perf/disaggregated/gb200_kimi-k25-thinking-fp4_1k1k_con4_ctx1_dep4_gen1_tep4_eplb0_mtp0_ccb-NIXL.yaml
- tests/scripts/perf/disaggregated/gb200_kimi-k25-thinking-fp4_8k1k_con4_ctx1_dep4_gen1_tep8_eplb0_mtp3_ccb-NIXL.yaml
- tests/scripts/perf/disaggregated/gb300_kimi-k25-thinking-fp4_1k1k_con4_ctx1_dep4_gen1_tep4_eplb0_mtp0_ccb-NIXL.yaml
- tests/scripts/perf/disaggregated/gb300_kimi-k25-thinking-fp4_1k1k_con4096_ctx1_dep4_gen1_dep8_eplb0_mtp0_ccb-UCX.yaml
- tests/scripts/perf/disaggregated/gb300_kimi-k25-thinking-fp4_8k1k_con4096_ctx1_dep4_gen1_dep16_eplb0_mtp0_ccb-UCX.yaml
- tests/scripts/perf/disaggregated/gb200_kimi-k25-thinking-fp4_1k1k_con4096_ctx1_dep4_gen1_dep8_eplb0_mtp0_ccb-UCX.yaml
- tests/scripts/perf/disaggregated/gb200_kimi-k25-thinking-fp4_8k1k_con4096_ctx1_dep4_gen1_dep16_eplb0_mtp0_ccb-UCX.yaml
- tests/integration/test_lists/qa/llm_function_core.txt
- tests/scripts/perf/disaggregated/gb300_kimi-k25-thinking-fp4_8k1k_con4_ctx1_dep4_gen1_tep8_eplb0_mtp3_ccb-NIXL.yaml
- tests/integration/test_lists/test-db/l0_gb200_multi_nodes_perf_sanity_ctx1_node1_gpu4_gen1_node2_gpu8.yml
- tests/integration/test_lists/test-db/l0_gb200_multi_gpus_perf_sanity.yml
- tests/integration/test_lists/test-db/l0_gb300_multi_nodes_perf_sanity_ctx1_node1_gpu4_gen1_node2_gpu8.yml
- tests/integration/test_lists/test-db/l0_gb300_multi_nodes_perf_sanity_ctx1_node1_gpu4_gen1_node1_gpu4.yml
- tests/integration/test_lists/test-db/l0_gb200_multi_nodes_perf_sanity_ctx1_node1_gpu4_gen1_node1_gpu4.yml
- tests/integration/test_lists/qa/llm_perf_multinode.txt
|
PR_Github #59657 [ run ] triggered by Bot. Commit: |
853df7a to
ef95e3a
Compare
|
/bot run --stage-list "DGX_B200-8_GPUs-PyTorch-1,DGX_B200-8_GPUs-PyTorch-2,DGX_B200-8_GPUs-PyTorch-3,DGX_B200-8_GPUs-PyTorch-4" |
|
/bot kill |
|
PR_Github #59662 [ run ] triggered by Bot. Commit: |
|
PR_Github #59664 [ kill ] triggered by Bot. Commit: |
|
PR_Github #59657 [ run ] completed with state |
ef95e3a to
d1be051
Compare
da7dcf0 to
bc28173
Compare
|
/bot run --disable-fail-fast |
|
PR_Github #60221 [ run ] triggered by Bot. Commit: |
Mirror the GPT-OSS approach (NVIDIA#16479): KimiK25ForConditionalGeneration now declares "PYTHON" as its preferred KV-cache transceiver runtime, so disaggregated Kimi-K2.5 adopts the Python (v2) transceiver by default when the user leaves transceiver_runtime at 'auto' and the effective backend is NIXL. Explicit user settings and non-NIXL backends are unaffected. - modeling_kimi_k25.py: add get_preferred_transceiver_runtime() -> "PYTHON". - test_modeling_kimi_k25.py: unit test asserting the preference. - TestKimiK25::test_nvfp4: set cache_transceiver_config explicitly to backend=NIXL + transceiver_runtime=PYTHON (ctx and gen). The disagg test harness injects TRTLLM_USE_UCX_KVCACHE=1 for any non-NIXL backend, so a "DEFAULT" backend would resolve to UCX and force the C++ transceiver; the test sets NIXL explicitly so it actually exercises the Python transceiver rather than relying on the model-default resolution. - Remove the redundant TestKimiK2::test_nvfp4 disaggregated accuracy test; Kimi disaggregated accuracy is covered by TestKimiK25, and Kimi-K2-Thinking accuracy stays covered by the aggregated tests in test_llm_api_pytorch. Signed-off-by: Shixiaowei02 <39303645+Shixiaowei02@users.noreply.github.com>
bc28173 to
0089eeb
Compare
|
/bot run --disable-fail-fast |
|
PR_Github #60221 [ run ] completed with state
|
|
PR_Github #60233 [ run ] triggered by Bot. Commit: |
|
PR_Github #60233 [ run ] completed with state
|
|
/bot run --disable-fail-fast --stage-list "GH200-PackageSanityCheck-PY312-UB2404, A10-PyTorch-1, B300-PyTorch-1, H100_PCIe-AutoDeploy-1" |
|
PR_Github #60354 [ run ] triggered by Bot. Commit: |
|
PR_Github #60354 [ run ] completed with state |
|
/bot run --disable-fail-fast --only-multi-gpu-test |
|
PR_Github #60404 [ run ] triggered by Bot. Commit: |
|
PR_Github #60404 [ run ] completed with state
|
|
/bot run --disable-fail-fast --stage-list "DGX_B300-4_GPUs-PyTorch-1, GB200-4_GPUs-PyTorch-2, GB200-4_GPUs-PyTorch-3, GB200-4_GPUs-PyTorch-5, GB200-4_GPUs-PyTorch-PerfSanity-1, GB200-4_GPUs-PyTorch-PerfSanity-2, GB200-8_GPUs-2_Nodes-PyTorch-1" |
|
PR_Github #60536 [ run ] triggered by Bot. Commit: |
|
PR_Github #60536 [ run ] completed with state |
|
/bot skip --comment "Related tests have passed." |
|
PR_Github #60631 [ skip ] triggered by Bot. Commit: |
|
PR_Github #60631 [ skip ] completed with state |
Description
This pull request removes the
TestKimiK2disaggregated serving test and updates theTestKimiK25test to use the NIXL Python (v2) cache transceiver instead of the default backend. The test list is also updated to reflect the removal of the Kimi-K2 test. These changes streamline the test suite and ensure the Kimi-K2.5 test uses the latest cache transceiver configuration.Test removals:
TestKimiK2class and itstest_nvfp4method fromtest_disaggregated_serving.py, which tested the Kimi-K2 model in a disaggregated serving setup.TestKimiK2::test_nvfp4entry from the test list inqa/llm_function_core.txt.Test configuration updates:
TestKimiK25.test_nvfp4method to use theNIXLbackend withPYTHONtransceiver runtime for both context and generation servers, replacing the previousDEFAULTbackend. [1] [2]Test Coverage
PR Checklist
Please review the following before submitting your PR:
PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.
PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.
Test cases are provided for new code paths (see test instructions)
If PR introduces API changes, an appropriate PR label is added - either
api-compatibleorapi-breaking. Forapi-breaking, includeBREAKINGin the PR title.Any new dependencies have been scanned for license and vulnerabilities
CODEOWNERS updated if ownership changes
Documentation updated as needed
Update tava architecture diagram if there is a significant design change in PR.
The reviewers assigned automatically/manually are appropriate for the PR.
Please check this after reviewing the above items as appropriate for this PR.
GitHub Bot Help
To see a list of available CI bot commands, please comment
/bot help.Summary by CodeRabbit