Skip to content

[None][fix] disagg-ctx: detach kv_cache page-index buffers on early IndexMapper release - #14423

Merged
Tabrizian merged 1 commit into
NVIDIA:feat/deepseek_v4from
Tabrizian:user/itabrizian/fix-disagg-ctx-page-index-buf-aliasing
May 21, 2026
Merged

[None][fix] disagg-ctx: detach kv_cache page-index buffers on early IndexMapper release#14423
Tabrizian merged 1 commit into
NVIDIA:feat/deepseek_v4from
Tabrizian:user/itabrizian/fix-disagg-ctx-page-index-buf-aliasing

Conversation

@Tabrizian

@Tabrizian Tabrizian commented May 21, 2026

Copy link
Copy Markdown
Member

…ndexMapper release

KVCacheManagerV2.release_index_slot recycles the IndexMapper slot for a context-only request once prefill completes, while the kv_cache itself stays alive holding pages for the upcoming KV transfer. The kv_cache's _base_page_indices memoryviews still point into the host_kv_cache_block_offsets row identified by the freed slot. When a new request takes that slot, both kv_caches alias the same row; the old kv_cache's late page unlocks (SeqBlock dtors) then write BAD_PAGE_INDEX over the new request's live page indices, causing the DSv4 compressor scatter kernel to read -1 and IMA on the CTX worker.

Detach the kv_cache's base-page-index buffers before removing the IndexMapper sequence so any subsequent writes from the old kv_cache land on a private array.array copy instead of the shared host row.

@coderabbitai summary

Description

Test Coverage

PR Checklist

Please review the following before submitting your PR:

  • PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.

  • PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.

  • Test cases are provided for new code paths (see test instructions)

  • If PR introduces API changes, an appropriate PR label is added - either api-compatible or api-breaking. For api-breaking, include BREAKING in the PR title.

  • Any new dependencies have been scanned for license and vulnerabilities

  • CODEOWNERS updated if ownership changes

  • Documentation updated as needed

  • Update tava architecture diagram if there is a significant design change in PR.

  • The reviewers assigned automatically/manually are appropriate for the PR.

  • Please check this after reviewing the above items as appropriate for this PR.

GitHub Bot Help

To see a list of available CI bot commands, please comment /bot help.

…ndexMapper release

KVCacheManagerV2.release_index_slot recycles the IndexMapper slot for a
context-only request once prefill completes, while the kv_cache itself
stays alive holding pages for the upcoming KV transfer. The kv_cache's
_base_page_indices memoryviews still point into the host_kv_cache_block_offsets
row identified by the freed slot. When a new request takes that slot,
both kv_caches alias the same row; the old kv_cache's late page unlocks
(SeqBlock dtors) then write BAD_PAGE_INDEX over the new request's live
page indices, causing the DSv4 compressor scatter kernel to read -1
and IMA on the CTX worker.

Detach the kv_cache's base-page-index buffers before removing the
IndexMapper sequence so any subsequent writes from the old kv_cache
land on a private array.array copy instead of the shared host row.

Signed-off-by: Iman Tabrizian <10105175+tabrizian@users.noreply.github.com>
@Tabrizian
Tabrizian requested a review from a team as a code owner May 21, 2026 17:01
@Tabrizian

Copy link
Copy Markdown
Member Author

/bot run --disable-fail-fast

@Tabrizian
Tabrizian requested review from pcastonguay and peihu-nv May 21, 2026 17:04
@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #49741 [ run ] triggered by Bot. Commit: 2a92312 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #49741 [ run ] completed with state SUCCESS. Commit: 2a92312
/LLM/main/L0_MergeRequest_PR pipeline #39344 completed with status: 'SUCCESS'

CI Report

Link to invocation

@Tabrizian
Tabrizian merged commit 02ac906 into NVIDIA:feat/deepseek_v4 May 21, 2026
8 of 9 checks passed
lfr-0531 pushed a commit to lfr-0531/TensorRT-LLM that referenced this pull request Jun 10, 2026
…ndexMapper release (NVIDIA#14423)

Signed-off-by: Iman Tabrizian <10105175+tabrizian@users.noreply.github.com>
lfr-0531 pushed a commit to lfr-0531/TensorRT-LLM that referenced this pull request Jun 11, 2026
…ndexMapper release (NVIDIA#14423)

Signed-off-by: Iman Tabrizian <10105175+tabrizian@users.noreply.github.com>
lfr-0531 pushed a commit to lfr-0531/TensorRT-LLM that referenced this pull request Jun 12, 2026
…ndexMapper release (NVIDIA#14423)

Signed-off-by: Iman Tabrizian <10105175+tabrizian@users.noreply.github.com>
jiaganc added a commit to jiaganc/TensorRT-LLM that referenced this pull request Jun 26, 2026
…ndexMapper release (NVIDIA#14423)

Source-Commit: 02ac906
Signed-off-by: Jiagan Cheng <jiaganc@nvidia.com>
jiaganc added a commit to jiaganc/TensorRT-LLM that referenced this pull request Jun 29, 2026
…ndexMapper release (NVIDIA#14423)

Source-Commit: 02ac906
Signed-off-by: Jiagan Cheng <jiaganc@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants