-
Notifications
You must be signed in to change notification settings - Fork 104
Pull requests: ROCm/FlyDSL
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[Kernel]Remove recast iter from sliding window attention kernel.
#997
opened Aug 10, 2026 by
amd-nprotaso
Contributor
Loading…
[Feat] Add fx.num_warp_threads() as constant accessor
#996
opened Aug 10, 2026 by
sjfeng1999
Collaborator
Loading…
1 task
Add type-aware extrema and integer ceil-division APIs [wip]
#995
opened Aug 10, 2026 by
Phil-amd
Member
Loading…
6 tasks done
Refactor gfx950 A16W16 GEMM to use the universal kernel
#992
opened Aug 10, 2026 by
xytpai
Contributor
Loading…
[fix] Resolve normalization-related CI regressions
#991
opened Aug 9, 2026 by
cschenjunlin
Contributor
Loading…
1 task
CI: upgrade PyTorch images to ROCm 7.14
multi-gpu
performance
Performance related issues
#989
opened Aug 9, 2026 by
coderfeli
Collaborator
Loading…
1 task
[WIP][Fix] Fix
lld invocation failed when ROCm is not at the baked-in path
#987
opened Aug 8, 2026 by
jli-melchior
Collaborator
Loading…
1 task
[Kernel] Fix single-accumulator RDNA3 GEMM tiles
#986
opened Aug 8, 2026 by
skyguan92
Contributor
Loading…
fix(pa): select the arch-native fp8 dtype in the metadata grid tuner
#984
opened Aug 7, 2026 by
zjin-lcf
Loading…
docs: Update docs style to fit style guide
#982
opened Aug 7, 2026 by
peterjunpark
•
2/2
Loading…
1 task
docs: Update configs for documentation to host on Read the Docs
#981
opened Aug 7, 2026 by
peterjunpark
•
1/2
Loading…
1 task
[Perf] Choose the RDNA3 GEMM tile from the shape
#980
opened Aug 7, 2026 by
vlluvia
Contributor
Loading…
[Feat] Add an experimental cuda nvvm backend
#976
opened Aug 6, 2026 by
sjfeng1999
Collaborator
Loading…
1 task
add A4W4 and FP8 P2P transport support to MegaMoE v2
multi-gpu
#972
opened Aug 6, 2026 by
Yaowu-Xiong
Contributor
Loading…
1 task
[Kernel][MI350] Add bias, alibi bias and sink to flash attention
#960
opened Aug 3, 2026 by
amd-nprotaso
Contributor
Loading…
[MoE] Port moe_gemm_2stage (stage1+stage2) to the new fx.* pipeline — fp8 (gfx942 + gfx950)
#947
opened Aug 1, 2026 by
coderfeli
Collaborator
Loading…
[Kernel] Add submanifold sparse 3D convolution (bf16 implicit GEMM, gfx950)
#942
opened Jul 31, 2026 by
jiacao-amd
Contributor
•
Draft
Add flex_attention (score_mod / mask_mod) on the generic flash-attention kernel
#931
opened Jul 30, 2026 by
RichardChamberlain1
•
Draft
3 of 4 tasks
DSL-ify raw arith float ops in kernels (flash/mla/pa/moe), fastmath v…
multi-gpu
#930
opened Jul 30, 2026 by
xudoyuan
Collaborator
Loading…
1 task
[Tool] Add per-kernel ISA resource delta table
#928
opened Jul 30, 2026 by
Phil-amd
Member
Loading…
3 of 4 tasks
Unify benchmark timing contracts and add calibrated CI gates
#924
opened Jul 30, 2026 by
jhinpan
Collaborator
Loading…
6 tasks done
[DSL] Preserve logical signedness of unsigned integer dtypes
#920
opened Jul 28, 2026 by
Arist12
Contributor
Loading…
6 tasks done
Previous Next
ProTip!
no:milestone will show everything without a milestone.