[TRTLLM-14026][feat] BREAKING: Don't review Remove TensorRT from the C++ tree#16476
Closed
Wanli-Jiang wants to merge 1 commit into
Closed
Conversation
Collaborator
Author
|
/bot run --disable-fail-fast |
Collaborator
|
PR_Github #59637 [ run ] triggered by Bot. Commit: |
1 task
Collaborator
|
PR_Github #59637 [ run ] completed with state
|
Sever the C++ tree's compile- and link-time dependency on TensorRT so the shared core (runtime, batch manager, executor API, KV cache, sampling, kernels, nanobind bridge) builds, links, and runs without the TensorRT library. Follows the Python TensorRT-backend removal (TRTLLM-14022); PyTorch is the sole backend. Internal types (serialization-compatible): - add tensorrt_llm::DataType/Dims (common/tllmDataType.h); DataType enumerator values mirror nvinfer1::DataType for byte-compatible serialization, Dims layout mirrors nvinfer1::Dims - migrate nvinfer1::DataType/Dims -> tensorrt_llm:: across 206 surviving files; replace NvInfer*.h includes with the internal header; drop 12 dead NvInfer includes from files that never referenced nvinfer1 - add tests/unittest/bindings/test_datatype_parity.py guarding the enumerator values (auto-scheduled via the existing unittest/bindings test-db entries) Remove the TensorRT-engine execution path (unused by the PyTorch backend): - plugins/ (nvinfer_plugin_tensorrt_llm), engine runtime wrappers (tllmRuntime, tllmStreamReaders, layerProfiler, rawEngine, tllmLogger), TRT model adapters (trtGptModel*/trtEncoderModel/trtGptModelFactory), executor.cpp/executorImpl, the C++ Executor + TllmRuntime + LogitsPostProcessor nanobind bindings (0 Python users each), executor_worker, disaggServerUtil, engine I/O buffers and engine-only logits/decoder algos, executor/model.h (0 consumers), the ModelSpec test helper + binding (orphaned) - keep the retained KVCacheEvent ctor and executor::version() inline in executor.h (they are trivial; no separate .cpp needed) - decouple shared files (inflightBatchingUtils, medusaBuffers, lookaheadBuffers, dataTransceiver) from the removed engine runtime - remove classes orphaned by the engine-path removal: GuidedDecoder (the PyTorch backend has its own Python guided decoder; xgrammar stays for executor::GuidedDecodingConfig), IntervalSet and DynamicBatchTuner (both only consumed by the removed executorImpl), each with their tests; removing guidedDecoderTest empties cpp/tests/e2e_tests entirely - restore the engine-free half of test_executor_bindings.py (config/ request/result/response/stats construction and pickle tests) since the underlying bindings stay and serve the PyTorch backend; only the trtllm.Executor/engine-fixture tests are dropped; the second, duplicated test_speculative_decoding_config (shadowed, never ran) is renamed to test_decoding_config to match what it tests - delete the engine-driven C++ tests (e2e_tests engine tests, executorTestSmall*, tllmRuntimeTest, encDecBeamSearchTest, tests/utils engine builders), cudaGraphExecutorCacheTest (tests the removed CudaGraphExecutor), unit_tests/utils (tests the deleted tests/utils helpers), tests/unittest/others/test_leak.py (its body calls the removed python graph-building APIs; unscheduled in CI), C++ benchmarks (bertBenchmark, gptManagerBenchmark, disaggServerBenchmark; the prepare_dataset.py tooling used by trtllm-bench stays), examples/cpp executor examples, and the cpp/tests/resources engine-build scripts - fix kept tests: strip vestigial never-consumed TllmLogger members (7 files), repoint ropeTest.cu to kernels/gptKernels.h (+ cudaUtils.h for QuantMode/getSMVersion), inline the engine-free createDecoderBatchInputs helper into gptDecoderBatchedTest, drop the empty tensorrt_llm::runtime forward-declaration block left in batch_manager/utils/debugUtils.h, narrow requestTest's using-namespace-common to a TllmException using-declaration (common now also exports DataType, which made the unqualified executor::DataType references ambiguous) - make dependencies the deleted plugin target satisfied transitively explicit: MPI include dirs for cpp/tests, th_utils -> tensorrt_llm shared-lib link - remove the IS_BUILDING build-time env contract end to end (common/opUtils.h isBuilding + attentionOp gate + _common.py half) - remove the llm_args=None engine path from executor/base_worker.py (tllm.Executor no longer exists); llm_args is now required Build/packaging (no TensorRT): - drop find_package(TensorRT)/TRT_LIB/NvInfer include injection and the plugins/executor_worker/benchmarks subdirs from the cpp CMake; delete FindTensorRT.cmake - build_wheel.py: trt_root optional (default None), drop the tensorrt venv check and the nvinfer_plugin/executorWorker/benchmarks targets - setup.py no longer packages libnvinfer_plugin_tensorrt_llm.so / executorWorker; requirements.txt drops tensorrt - jenkins/Build.groovy: drop --benchmarks, the benchmark/libnvinfer-plugin tarball packaging, and the build_cpp_examples.py step - delist the removed tests from test-db/waives.txt/.test_durations and prune tests/integration/defs/cpp to the shared-core gtest wrappers - CI-run fixes: respell the nvinfer1 usages a post-rebase main commit (NVIDIA#16304) added to kvCacheManagerTest; retire the two CI stages emptied by the delisting (A10-CPP-1, H100_PCIe-CPP-Post-Merge-1) and drop the duplicate tensorrt-chunk scheduling of unittest/disaggregated/test_router.py; make test_to_stage_mapping's CLI check sample only stage-mapped tests - point the remaining benchmarks/cpp references (AutoDeploy bench/dist tests, llmc standalone packager + its test, trtllm-bench docs) at benchmarks/ and drop the orphaned get_cpp_benchmark() helper - fix the two OSS-mode gtest compile failures unmasked by dropping the nvinfer_plugin link (its PUBLIC USING_OSS_CUTLASS_*_GEMM defines had leaked into every gtest, keeping the internal-only branches dormant): inline the deleted tests/utils GpuTimer into gemmAllReduceTest.cu, and only add the mixtureOfExpertsInternalTest target when INTERNAL_CUTLASS_KERNELS_PATH provides the internal headers it includes Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>
Wanli-Jiang
force-pushed
the
user/williamj/deprecated-trt-backend-removal-cpp-t3
branch
from
July 17, 2026 06:17
1b750aa to
9f0a563
Compare
Collaborator
Author
|
/bot run --post-merge --disable-fail-fast |
Collaborator
|
PR_Github #59904 [ run ] triggered by Bot. Commit: |
Collaborator
|
PR_Github #59904 [ run ] completed with state
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
@coderabbitai summary
Description
Test Coverage
PR Checklist
Please review the following before submitting your PR:
PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.
PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.
Test cases are provided for new code paths (see test instructions)
If PR introduces API changes, an appropriate PR label is added - either
api-compatibleorapi-breaking. Forapi-breaking, includeBREAKINGin the PR title.Any new dependencies have been scanned for license and vulnerabilities
CODEOWNERS updated if ownership changes
Documentation updated as needed
Update tava architecture diagram if there is a significant design change in PR.
The reviewers assigned automatically/manually are appropriate for the PR.
Please check this after reviewing the above items as appropriate for this PR.
GitHub Bot Help
To see a list of available CI bot commands, please comment
/bot help.