[TRTLLM-14026][feat] BREAKING: Don't review Remove TensorRT from the C++ tree (t4)#16477
Conversation
|
/bot run --post-merge --disable-fail-fast |
|
PR_Github #59642 [ run ] triggered by Bot. Commit: |
|
/bot run --post-merge --disable-fail-fast |
|
PR_Github #59691 [ run ] triggered by Bot. Commit: |
|
PR_Github #59642 [ run ] completed with state |
|
PR_Github #59691 [ run ] completed with state
|
1b750aa to
9f0a563
Compare
|
/bot run --only-multi-gpu-test --post-merge --disable-fail-fast |
|
PR_Github #59905 [ run ] triggered by Bot. Commit: |
|
PR_Github #59905 [ run ] completed with state
|
Sever the C++ tree's compile- and link-time dependency on TensorRT so the shared core (runtime, batch manager, executor API, KV cache, sampling, kernels, nanobind bridge) builds, links, and runs without the TensorRT library. Follows the Python TensorRT-backend removal (TRTLLM-14022); PyTorch is the sole backend. Internal types (serialization-compatible): - add tensorrt_llm::DataType/Dims (common/tllmDataType.h); DataType enumerator values mirror nvinfer1::DataType for byte-compatible serialization, Dims layout mirrors nvinfer1::Dims - migrate nvinfer1::DataType/Dims -> tensorrt_llm:: across 206 surviving files; replace NvInfer*.h includes with the internal header; drop 12 dead NvInfer includes from files that never referenced nvinfer1 - add tests/unittest/bindings/test_datatype_parity.py guarding the enumerator values (auto-scheduled via the existing unittest/bindings test-db entries) Remove the TensorRT-engine execution path (unused by the PyTorch backend): - plugins/ (nvinfer_plugin_tensorrt_llm), engine runtime wrappers (tllmRuntime, tllmStreamReaders, layerProfiler, rawEngine, tllmLogger), TRT model adapters (trtGptModel*/trtEncoderModel/trtGptModelFactory), executor.cpp/executorImpl, the C++ Executor + TllmRuntime + LogitsPostProcessor nanobind bindings (0 Python users each), executor_worker, disaggServerUtil, engine I/O buffers and engine-only logits/decoder algos, executor/model.h (0 consumers), the ModelSpec test helper + binding (orphaned) - keep the retained KVCacheEvent ctor and executor::version() inline in executor.h (they are trivial; no separate .cpp needed) - decouple shared files (inflightBatchingUtils, medusaBuffers, lookaheadBuffers, dataTransceiver) from the removed engine runtime - remove classes orphaned by the engine-path removal: GuidedDecoder (the PyTorch backend has its own Python guided decoder; xgrammar stays for executor::GuidedDecodingConfig), IntervalSet and DynamicBatchTuner (both only consumed by the removed executorImpl), each with their tests; removing guidedDecoderTest empties cpp/tests/e2e_tests entirely - restore the engine-free half of test_executor_bindings.py (config/ request/result/response/stats construction and pickle tests) since the underlying bindings stay and serve the PyTorch backend; only the trtllm.Executor/engine-fixture tests are dropped; the second, duplicated test_speculative_decoding_config (shadowed, never ran) is renamed to test_decoding_config to match what it tests - delete the engine-driven C++ tests (e2e_tests engine tests, executorTestSmall*, tllmRuntimeTest, encDecBeamSearchTest, tests/utils engine builders), cudaGraphExecutorCacheTest (tests the removed CudaGraphExecutor), unit_tests/utils (tests the deleted tests/utils helpers), tests/unittest/others/test_leak.py (its body calls the removed python graph-building APIs; unscheduled in CI), C++ benchmarks (bertBenchmark, gptManagerBenchmark, disaggServerBenchmark; the prepare_dataset.py tooling used by trtllm-bench stays), examples/cpp executor examples, and the cpp/tests/resources engine-build scripts - fix kept tests: strip vestigial never-consumed TllmLogger members (7 files), repoint ropeTest.cu to kernels/gptKernels.h (+ cudaUtils.h for QuantMode/getSMVersion), inline the engine-free createDecoderBatchInputs helper into gptDecoderBatchedTest, drop the empty tensorrt_llm::runtime forward-declaration block left in batch_manager/utils/debugUtils.h, narrow requestTest's using-namespace-common to a TllmException using-declaration (common now also exports DataType, which made the unqualified executor::DataType references ambiguous) - make dependencies the deleted plugin target satisfied transitively explicit: MPI include dirs for cpp/tests, th_utils -> tensorrt_llm shared-lib link - remove the IS_BUILDING build-time env contract end to end (common/opUtils.h isBuilding + attentionOp gate + _common.py half) - remove the llm_args=None engine path from executor/base_worker.py (tllm.Executor no longer exists); llm_args is now required Build/packaging (no TensorRT): - drop find_package(TensorRT)/TRT_LIB/NvInfer include injection and the plugins/executor_worker/benchmarks subdirs from the cpp CMake; delete FindTensorRT.cmake - build_wheel.py: trt_root optional (default None), drop the tensorrt venv check and the nvinfer_plugin/executorWorker/benchmarks targets - setup.py no longer packages libnvinfer_plugin_tensorrt_llm.so / executorWorker; requirements.txt drops tensorrt - jenkins/Build.groovy: drop --benchmarks, the benchmark/libnvinfer-plugin tarball packaging, and the build_cpp_examples.py step - delist the removed tests from test-db/waives.txt/.test_durations and prune tests/integration/defs/cpp to the shared-core gtest wrappers - CI-run fixes: respell the nvinfer1 usages a post-rebase main commit (NVIDIA#16304) added to kvCacheManagerTest; retire the two CI stages emptied by the delisting (A10-CPP-1, H100_PCIe-CPP-Post-Merge-1) and drop the duplicate tensorrt-chunk scheduling of unittest/disaggregated/test_router.py; make test_to_stage_mapping's CLI check sample only stage-mapped tests - point the remaining benchmarks/cpp references (AutoDeploy bench/dist/ allreduce-strategy tests, llmc standalone packager + its test, trtllm-bench docs) at benchmarks/ and drop the orphaned get_cpp_benchmark() helper - fix the two OSS-mode gtest compile failures unmasked by dropping the nvinfer_plugin link (its PUBLIC USING_OSS_CUTLASS_*_GEMM defines had leaked into every gtest, keeping the internal-only branches dormant): inline the deleted tests/utils GpuTimer into gemmAllReduceTest.cu, and only add the mixtureOfExpertsInternalTest target when INTERNAL_CUTLASS_KERNELS_PATH provides the internal headers it includes Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>
9f0a563 to
e2e81d0
Compare
|
/bot help |
GitHub Bot Help
Provide a user friendly way for developers to interact with a Jenkins server. Run See details below for each supported subcommand. Details
Launch build/test pipelines. All previously running jobs will be killed.
kill
Kill all running builds associated with pull request. skip
Skip testing for latest commit on pull request. reuse-pipeline
Reuse a previous pipeline to validate current commit. This action will also kill all currently running builds associated with the pull request. IMPORTANT NOTE: This is dangerous since lack of user care and validation can cause top of tree to break. |
|
/bot run --only-multi-gpu-test --post-merge --disable-fail-fast --stage-list "DGX_B200-4_GPUs-PyTorch-Post-Merge-2,DGX_B200-8_GPUs-PyTorch-PerfSanity-Post-Merge-3,DGX_B300-4_GPUs-PyTorch-Post-Merge-1,DGX_B200-16_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU8-Post-Merge-4,DGX_H200-4_GPUs-PyTorch-Post-Merge-1,DGX_H200-8_GPUs-PyTorch-PerfSanity-Post-Merge-1,DGX_H200-8_GPUs-PyTorch-Post-Merge-1,GB200-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-6,GB200-8_GPUs-2_Nodes-PyTorch-PerfSanity-Node2-GPU8-Post-Merge-2,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU2-GEN1-NODE2-GPU8-Post-Merge-2,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-2,GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-2,GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-4,GB300-36_GPUs-9_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE8-GPU32-Post-Merge-2,GB300-8_GPUs-2_Nodes-PyTorch-PerfSanity-Node2-GPU8-Post-Merge-2,GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge-1" |
|
PR_Github #60142 [ run ] triggered by Bot. Commit: |
|
/bot run --only-multi-gpu-test --post-merge --disable-fail-fast --stage-list "DGX_B200-4_GPUs-PyTorch-Post-Merge-2,DGX_B200-8_GPUs-PyTorch-PerfSanity-Post-Merge-3,DGX_B300-4_GPUs-PyTorch-Post-Merge-1,DGX_B200-16_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU8-Post-Merge-1,DGX_B200-16_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE1-GPU8-Post-Merge-2,DGX_H200-4_GPUs-PyTorch-Post-Merge-1,DGX_H200-8_GPUs-PyTorch-PerfSanity-Post-Merge-1,DGX_H200-8_GPUs-PyTorch-Post-Merge-1,GB200-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1,GB200-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-2,GB200-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-3,GB200-8_GPUs-2_Nodes-PyTorch-PerfSanity-Node2-GPU8-Post-Merge-2,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU2-GEN1-NODE2-GPU8-Post-Merge-2,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-2,GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-2,GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-4,GB300-36_GPUs-9_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE8-GPU32-Post-Merge-1,GB300-8_GPUs-2_Nodes-PyTorch-PerfSanity-Node2-GPU8-Post-Merge-2,GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge-1" |
|
PR_Github #60143 [ run ] triggered by Bot. Commit: |
|
PR_Github #60142 [ run ] completed with state |
|
PR_Github #60143 [ run ] completed with state
|
@coderabbitai summary
Description
Test Coverage
PR Checklist
Please review the following before submitting your PR:
PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.
PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.
Test cases are provided for new code paths (see test instructions)
If PR introduces API changes, an appropriate PR label is added - either
api-compatibleorapi-breaking. Forapi-breaking, includeBREAKINGin the PR title.Any new dependencies have been scanned for license and vulnerabilities
CODEOWNERS updated if ownership changes
Documentation updated as needed
Update tava architecture diagram if there is a significant design change in PR.
The reviewers assigned automatically/manually are appropriate for the PR.
Please check this after reviewing the above items as appropriate for this PR.
GitHub Bot Help
To see a list of available CI bot commands, please comment
/bot help.