vulkan backend ops: implemented GATED_LINEAR_ATTN - #25601
Conversation
|
Hi @PranavUttarkar, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
Not a Large PR. maybe the bot thinks so due to the vulkan.csv file updated with this: PR follows the template |
d512d9d to
aebd8ed
Compare
Regenerate docs/ops/Vulkan.csv and docs/ops.md so ops docs match current backend CSVs. Assisted-by: Cursor Grok
34dd27b to
5fc8382
Compare
|
@CISC Hey I just updated the PR |
* vulkan : add GATED_LINEAR_ATTN op * docs : update Vulkan ops * vulkan : remove unused GLA spec constant * Updated ops.md * ops.md update
* vulkan : add GATED_LINEAR_ATTN op * docs : update Vulkan ops * vulkan : remove unused GLA spec constant * Updated ops.md * ops.md update
No shared git history exists with upstream (our root commit is parentless, a fresh source import rather than a real fork/clone), so a normal merge/rebase isn't possible -- cherry-picked commits individually instead, each verified independently. Landed 7 of 8 identified commits; skipped upstream's own Q2_0 (788e07d) since it collides with our already-verified, benchmarked custom Q2_0 (type 42) implementation -- same feature name, two independently-developed and incompatible kernels. - e34e9c6 vulkan: Refactor vk_queue to use per-instance mutexes and unique handles (ggml-org#23570) -- required updating our custom paged-KV/radix Vulkan code (pkv_vulkan_init/pkv_dispatch_init call sites), which accessed the old plain vk_queue struct fields directly. - 27b7fcf vulkan: add iq4_nl support back to FA (ggml-org#24585) - 29f1212 vulkan: Support quantized concat (ggml-org#25684) - 3cb4d39 vulkan: add POOL_1D op (ggml-org#25431) - a329897 vulkan: extend topk_moe fusion to support sqrt(softplus) (ggml-org#26124) - a72c87f vulkan backend ops: implemented GATED_LINEAR_ATTN (ggml-org#25601) - 0756c85 vulkan: fix submission batching size, add debug tools for diagnosing causes of DeviceLost drivers errors (ggml-org#26371) -- AMD-specific driver-timeout workaround, directly relevant to this fork's RDNA4 focus. Verified: full Release build of build-vulkan-rdna4-fresh (0 errors), real model load + inference request through the Vulkan backend produced correct output. docs/ops.md left as the incoming (upstream) version for now since it's auto-generated (scripts/create_ops_docs.py) from real test-backend-ops CSV dumps per backend -- needs a proper regeneration pass, not hand-editing, to reflect our custom ops too. Backup of pre-sync main state: branch + tag pre-upstream-vulkan-sync (pushed to origin). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* vulkan : add GATED_LINEAR_ATTN op * docs : update Vulkan ops * vulkan : remove unused GLA spec constant * Updated ops.md * ops.md update
Overview
Added vulkan support for GGML_OP_GATED_LINEAR_ATTN
This backend op used to not be supported on vulkan and fell back to cpu. Now the kernel follows the existing wkv6.comp pattern w/ a GLA-specific update-before-read ordering and an output "scale" push constant.
supports_op is limited to F32 and head_size == 64 (shader hardcodes BLOCK_SIZE 64, same as WKV6).
Part of #14909
Additional information
Modeled on the vulkan WKV6 path. Checked against the CPU reference in ggml_compute_forward_gla_f32
Test results (AMD Radeon 780M, Windows)
test-backend-ops.exe test -b Vulkan0 -o GATED_LINEAR_ATTN 4/4 tests passed
adjacent recurrence ops after rebase:
Full Vulkan0 suite: 14400/14443. Remaining failures are preexisting DIV(type=f16, ...)`NMSE misses on this device/driver. reproduced on clean master with this change stashed and they remain.
Requirements
YES: AI was used in the beginning to understand the codebase and for easier search and navigation to different similar points. Code was all handwritten and then AI was used to review correctness and any gaps/oversights.