[SYCL] Use native subgroup size for K-quant DMMV kernels on Intel - #21700
Conversation
|
Oddly I saw a TG improvement but not so much a PP with the B60 🤔 Llama-2-7B Q2_K (dual GPU)
Qwen3.5-27B Q2_K_XL (dual GPU)
Qwen2.5-1.5B-Instruct Q2_K (single GPU)
|
|
It needs to be verified on more GPUs: iGPU, Arc7xx, BMG and Xe iGPU (meteor lake or newer). Thank you! |
|
Corrected the benchmark results; my original numbers compared builds with different GGML_SYCL_F16 settings, which inflated the pp numbers significantly. @maxious thanks for testing on the B60. Your clean A/B comparison is what made the mismatch obvious. The real effect is a tg improvement on compute-bound K-quants (primarily Q2_K), not the pp speedup I originally claimed. Updated title and description to reflect this. The change is still architecturally correct; these are the only DMMV kernels still using the non-native subgroup size, and the DPCT register pressure warnings confirm 32 is too wide for Intel (at least, our cards). |
|
We use 32 as warp_size in some kernel for better performance by test. |
|
Sorry, it's my mistake to close this PR. |
|
The Arc770,BMG580, iGPU (UHD) are not impacted obviously. I think it's acceptable. Thank you! |
arthw
left a comment
There was a problem hiding this comment.
It's good job.
The sub-group size 16 is more useful on Intel GPUs.
The legacy code use 32 is based on test result.
Based on the latest driver and compiler, change them to 16 will make the code be clear and easy maintained.
It also approves the value 16 is better value to all Intel existed GPUs.
There are increase impact on B70/B60/PVC.
There is no impact to most of Intel old dGPU and iGPU (Arc770, BMG580, iGPU).
Except PVC has -4% of TG on Q4_K, there are no bad impact of performance on other Intel GPUs.
Thank you!
839c7e2 to
6f28d8c
Compare
|
Rebased onto current master and expanded scope to also cover the Q4_K and Q6_K reorder DMMV kernels added in #21638. After #21638 merged, dense Q4_K_M models (e.g. EVA-Qwen2.5-72B) hang during warmup on Intel GPUs. The reorder DMMV kernels were written against the unfixed QK_WARP_SIZE=32 code because this PR hadn't merged yet. This rebase applies the same Testing on B70:
Updated the PR description with full details. |
|
This PR is approved. Thank you! |
6f28d8c to
4213f77
Compare
|
Hi @arthw, thanks for the approval. I rebased the PR onto current master and resolved the conflict cleanly. While rebasing, I also updated the newly added Q3_K reorder DMMV kernel, so the PR now covers all 8 affected K-quant DMMV kernels: Q2/Q3/Q4/Q5/Q6 plus Q3/Q4/Q6 reorder. I reran focused validation on Intel SYCL:
The branch is now mergeable and still only touches |
Summary
Use
WARP_SIZE(16) instead ofQK_WARP_SIZE(32) for the K-quant DMMV kernels on Intel SYCL, including the reorder variants now present onmaster.These kernels were migrated from CUDA via DPCT and kept a 32-wide subgroup size. On Intel targets, native subgroup size is 16. DPCT flagged the original K-quant kernels with register pressure warnings recommending a smaller subgroup size, and the non-K-quant DMMV path already uses
WARP_SIZE.Each thread now processes both halves of the
QK_K=256block via afor (int im = 0; im < 2; ++im)loop. The inner dot-product computation is unchanged.Updated scope
This rebased version covers 8 affected kernels in
ggml/src/ggml-sycl/dmmv.cpp:dequantize_mul_mat_vec_q2_kdequantize_mul_mat_vec_q3_kdequantize_mul_mat_vec_q3_k_reorderdequantize_mul_mat_vec_q4_kdequantize_mul_mat_vec_q4_k_reorderdequantize_mul_mat_vec_q5_kdequantize_mul_mat_vec_q6_kdequantize_mul_mat_vec_q6_k_reorderThe original PR covered the five non-reorder K-quant kernels. Later master changes added K-quant reorder DMMV coverage, including Q3_K, Q4_K, and Q6_K, using the same old
QK_WARP_SIZE=32pattern. This rebase applies the same native-subgroup fix to all currently affected kernels.Changes
QK_WARP_SIZEtoWARP_SIZE.QK_WARP_SIZE / 2toWARP_SIZE / 2.QK_K=256halves in an explicitimloop, preserving total work.Current rebase verification
Rebased cleanly onto
ggml-org/llama.cpp@8ed274ef46e2e6c073e5400af4286f53547c003e.Test system: Intel Graphics
[0xe223]via Level Zero, oneAPI 2026.0, AOTbmg_g21, Release builds.cmake --build build-kquant-rebase-icpx --target test-backend-ops llama-cli llama-bench -j 12: passedtest-backend-ops -o MUL_MAT -b SYCL0withGGML_SYCL_F16=OFF:920/920 tests passedGGML_SYCL_PRIORITIZE_DMMV=1 GGML_SYCL_DEBUG=1 test-backend-ops -o MUL_MAT -b SYCL0:920/920 tests passedcmake --build build-kquant-rebase-icpx-f16 --target test-backend-ops -j 12: passedtest-backend-ops -o MUL_MAT -b SYCL0withGGML_SYCL_F16=ON:920/920 tests passedgit diff --check: passedQK_WARP_SIZEor staleDPCT1110warning blocks indmmv.cppForced-DMMV real-model smoke tests using
GGML_SYCL_PRIORITIZE_DMMV=1,llama-bench -p 64 -n 16 -r 1 -ngl 999 -dev SYCL0:Previous PR testing
Intel Arc Pro B70 (Xe2/Battlemage, 32 GB), dual GPU, oneAPI 2025.3.3, Ubuntu 26.04.
Hang fix:
Correctness:
test-backend-ops -o MUL_MAT -b SYCL0:911/911 passed, 0 failuresPerformance (single GPU, sequential, Release, JIT,
GGML_SYCL_F16=ON):Q8_0 is unaffected by this PR and already uses
WARP_SIZE; pp variance is normal run-to-run variation. Previous testing by @arthw across Arc770, BMG580, iGPU (UHD), and PVC confirmed the original non-reorder changes were acceptable. The reorder kernels follow the same pattern.