[None][test] Add TLLM_SPEC_DECODE_FORCE_NUM_ACCEPTED_TOKENS in Spec Decoding Perf Test - #14438
Conversation
📝 WalkthroughWalkthroughThis PR extends a performance sanity test to measure and track speculative-decoding-only throughput metrics. It adds metric parsing for average decoded tokens per iteration, integrates it into categorization and regression logic, validates its presence for speculative-decoding clients, and conditionally uploads the metric to avoid OpenSearch baseline contamination. ChangesSpeculative-Decoding Metric Addition
🎯 2 (Simple) | ⏱️ ~10 minutes 🚥 Pre-merge checks | ✅ 3 | ❌ 2❌ Failed checks (2 warnings)
✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
tests/integration/defs/perf/test_perf_sanity.py (1)
1-1:⚠️ Potential issue | 🟡 Minor | ⚡ Quick winUpdate the SPDX copyright year.
This file is modified in this 2026 PR, but Line 1 still ends at 2025.
Proposed fix
-# SPDX-FileCopyrightText: Copyright (c) 2022-2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-FileCopyrightText: Copyright (c) 2022-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.As per coding guidelines,
Include NVIDIA copyright header on ALL new files; update year on modified files.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tests/integration/defs/perf/test_perf_sanity.py` at line 1, The SPDX copyright header at the top of test_perf_sanity.py still lists the end year as 2025; update the header comment to end with 2026 so the SPDX line matches the file modification year (change "2022-2025" to "2022-2026" in the top-of-file comment).
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Outside diff comments:
In `@tests/integration/defs/perf/test_perf_sanity.py`:
- Line 1: The SPDX copyright header at the top of test_perf_sanity.py still
lists the end year as 2025; update the header comment to end with 2026 so the
SPDX line matches the file modification year (change "2022-2025" to "2022-2026"
in the top-of-file comment).
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: e7df7015-259b-46f2-a83b-09270d749788
📒 Files selected for processing (1)
tests/integration/defs/perf/test_perf_sanity.py
c002d33 to
33ec282
Compare
|
/bot run --disable-fail-fast --stage-list "DGX_B200-16_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU8-Post-Merge-1,GB200-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-2" |
|
PR_Github #50517 [ run ] triggered by Bot. Commit: |
|
PR_Github #50517 [ run ] completed with state
|
561f140 to
4eb225f
Compare
|
/bot run --disable-fail-fast --stage-list "GB300-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU4-Post-Merge-4,GB300-8_GPUs-2_Nodes-PyTorch-PerfSanity-Node2-GPU8-Post-Merge-2,GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge-2" |
|
PR_Github #51038 [ run ] triggered by Bot. Commit: |
|
PR_Github #51038 [ run ] completed with state
|
|
/bot run --disable-fail-fast --stage-list "GB300-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU4-Post-Merge-4,GB300-8_GPUs-2_Nodes-PyTorch-PerfSanity-Node2-GPU8-Post-Merge-2" |
|
PR_Github #51057 [ run ] triggered by Bot. Commit: |
|
PR_Github #51057 [ run ] completed with state
|
feb2679 to
bd8fa9e
Compare
|
/bot run --disable-fail-fast --stage-list "GB300-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU4-Post-Merge-4,GB300-8_GPUs-2_Nodes-PyTorch-PerfSanity-Node2-GPU8-Post-Merge-2" |
|
PR_Github #51152 [ run ] triggered by Bot. Commit: |
|
PR_Github #51152 [ run ] completed with state
|
|
/bot run --disable-fail-fast --stage-list "GB300-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU4-Post-Merge-4,GB300-8_GPUs-2_Nodes-PyTorch-PerfSanity-Node2-GPU8-Post-Merge-2" |
|
PR_Github #51162 [ run ] triggered by Bot. Commit: |
bd8fa9e to
61f23f0
Compare
|
PR_Github #51162 [ run ] completed with state
|
|
/bot run --disable-fail-fast --stage-list "GB300-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU4-Post-Merge-4,GB300-8_GPUs-2_Nodes-PyTorch-PerfSanity-Node2-GPU8-Post-Merge-2" |
|
PR_Github #51174 [ run ] triggered by Bot. Commit: |
|
PR_Github #51174 [ run ] completed with state
|
|
/bot run --disable-fail-fast --stage-list "GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU4-Post-Merge-1,GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU4-Post-Merge-2,GB200-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-6,GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU1-GEN1-NODE1-GPU2-Post-Merge-3,GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU1-GEN1-NODE1-GPU2-Post-Merge-4,GB200-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-8,GB200-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU1-GEN1-NODE2-GPU8-Post-Merge-2" |
|
PR_Github #51186 [ run ] triggered by Bot. Commit: |
|
PR_Github #51186 [ run ] completed with state
|
|
/bot skip --comment "Only add new perf tests, no need to run the whole CI pipeline" |
|
PR_Github #51192 [ skip ] triggered by Bot. Commit: |
|
PR_Github #51192 [ skip ] completed with state |
Summary by CodeRabbit
Code Change Summary
▎ 1. Always pass --ignore-eos, and stabilize accepted-token count via TLLM_SPEC_DECODE_FORCE_NUM_ACCEPTED_TOKENS (set per server_config in yaml; never computed in code).
▎ 2. New jenkins/scripts/perf/aggregated/submit.py for multi-node aggregated PerfSanity, plus three-way routing in L0_Test.groovy. Needed because trtllm-llmapi-launch dispatches rank 0 to pytest and other
▎ ranks to mgmn_worker_node — env vars must be set in the slurm script prefix to reach all ranks.
▎ 3. New d_al metric (acceptance length) for spec decoding tests, with l_force_num_accepted_tokens added as a baseline match key so different forced values match separately.
▎ 4. New d_mean_gen_worker_per_iter_device_step_time metric for gen_only tests; gen_only regression gates on this instead of throughput.
▎ 5. Per-test log isolation — all slurm logs land in <output_dir>/<test_case_name>/.
▎ 6. Yaml schema cleanup — agg yamls move env to per-server-config server_env_var; disagg adds spec-decode env to worker_env_var only.
Description
Test Coverage
PR Checklist
Please review the following before submitting your PR:
PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.
PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.
Test cases are provided for new code paths (see test instructions)
If PR introduces API changes, an appropriate PR label is added - either
api-compatibleorapi-breaking. Forapi-breaking, includeBREAKINGin the PR title.Any new dependencies have been scanned for license and vulnerabilities
CODEOWNERS updated if ownership changes
Documentation updated as needed
Update tava architecture diagram if there is a significant design change in PR.
The reviewers assigned automatically/manually are appropriate for the PR.
Please check this after reviewing the above items as appropriate for this PR.
GitHub Bot Help
To see a list of available CI bot commands, please comment
/bot help.