Revert "[None] Add sink/sliding window support for Triton" - #92
Merged
lucaslie merged 1 commit intoJul 17, 2025
Merged
Conversation
This reverts commit a37797b.
lucaslie
enabled auto-merge (squash)
July 17, 2025 15:34
lucaslie
disabled auto-merge
July 17, 2025 15:34
There was a problem hiding this comment.
Pull Request Overview
This PR reverts the addition of sink tokens and sliding window support for Triton attention kernels due to CUDA device assertions encountered during unit tests. The revert removes the new eager implementation of torch attention that was causing numerical stability issues.
- Removes sink token support from Triton attention kernels and test infrastructure
- Removes sliding window attention functionality from all kernel implementations
- Simplifies function signatures by removing optional sink and sliding window parameters
Reviewed Changes
Copilot reviewed 4 out of 4 changed files in this pull request and generated 1 comment.
| File | Description |
|---|---|
| test_attention_with_kv_cache.py | Removes test parameterization for sliding window and sink tokens, simplifies test functions |
| attention_with_kv_cache.py | Removes sliding window logic and sink token handling from Triton kernels |
| triton_attention.py | Removes sink and sliding window parameters from public API functions |
| _triton_attention_internal.py | Removes sink and sliding window parameters from internal kernel calls |
lucaslie
enabled auto-merge (squash)
July 17, 2025 15:34
lucaslie
disabled auto-merge
July 17, 2025 15:37
lucaslie
added a commit
that referenced
this pull request
Jul 18, 2025
lucaslie
added a commit
that referenced
this pull request
Jul 21, 2025
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Reverts #77
I am getting a couple of cuda device assertions when running the unit tests - I am assuming it is related to the numerical stability of the updated triton kernel