Skip to content

[None][fix] Fix Backend's CUDA graph pool handle lifetime - #16995

Closed
lori-ren wants to merge 1 commit into
NVIDIA:mainfrom
lori-ren:fix/torch-compile-graph-pool-scope
Closed

[None][fix] Fix Backend's CUDA graph pool handle lifetime#16995
lori-ren wants to merge 1 commit into
NVIDIA:mainfrom
lori-ren:fix/torch-compile-graph-pool-scope

Conversation

@lori-ren

Copy link
Copy Markdown
Contributor

@coderabbitai summary

Description

Backend._graph_pool_handle is a class attribute, created once per process and never reset. With torch.compile on, the decode CUDA graph runner uses it, but on engine teardown _release_cuda_graphs() resets the graphs, dropping the pool's use_count to 0, creating danling handle in Backend. If a second engine is built in the same process — e.g. pytest reusing the MPI worker pool across cases — it passes that dead pool id to graph capture and trips use_count > 0 INTERNAL ASSERT FAILED.

Fix: make the handle per-instance, so each engine starts from a fresh pool. Sharing within one engine is unchanged. Verified on 4×H100: the full test_fp8_4gpus sweep in one pytest process now passes, GSM8K accuracy unchanged.

Test Coverage

  • pytest TestLlama3_1_8BInstruct::test_fp8_4gpus should pass despite reusing same session.

PR Checklist

Please review the following before submitting your PR:

  • PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.

  • PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.

  • Test cases are provided for new code paths (see test instructions)

  • If PR introduces API changes, an appropriate PR label is added - either api-compatible or api-breaking. For api-breaking, include BREAKING in the PR title.

  • Any new dependencies have been scanned for license and vulnerabilities

  • CODEOWNERS updated if ownership changes

  • Documentation updated as needed

  • Update tava architecture diagram if there is a significant design change in PR.

  • The reviewers assigned automatically/manually are appropriate for the PR.

  • Please check this after reviewing the above items as appropriate for this PR.

GitHub Bot Help

To see a list of available CI bot commands, please comment /bot help.

Signed-off-by: Lori Ren <lorir@nvidia.com>
@lori-ren

Copy link
Copy Markdown
Contributor Author

/bot run

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #62451 [ run ] triggered by Bot. Commit: c71194b Link to invocation

@liji-nv

liji-nv commented Jul 29, 2026

Copy link
Copy Markdown
Collaborator

#16952 is fixing the same issue.

@lori-ren

Copy link
Copy Markdown
Contributor Author

/bot kill

@lori-ren

Copy link
Copy Markdown
Contributor Author

I see, closing this PR as duplicated

@lori-ren lori-ren closed this Jul 29, 2026
@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #62466 [ ] completed with state FAILURE. Commit: c71194b
Not allowed on merged PR

Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #62451 [ run ] completed with state SUCCESS. Commit: c71194b
/LLM/main/L0_MergeRequest_PR pipeline #50605 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants