forked from ggml-org/llama.cpp
-
Notifications
You must be signed in to change notification settings - Fork 0
Home
Maciej Wakuła edited this page Jun 28, 2026
·
6 revisions
This project is my own fork of llama.cpp
meant for AMD rx 7900 xtx with 24GB VRAM
to use models (especially those from Bartowski)
with turboquant KV
while implementing features not yet present in official llama.cpp (which supports many devices, tools and frameworks).
BEWARE: My fork is likely to be buggy/unstable (but leave me a comment/star/watch if you found it useful :) )
I am vibe coding often (while rarely also coding myself... nowadays manual edits are too expensive)
Tools used: my own coder, claude, hermes, gemini, my own brain.