Skip to content
Maciej Wakuła edited this page Jun 28, 2026 · 6 revisions

This project is my own fork of llama.cpp

meant for AMD rx 7900 xtx with 24GB VRAM

to use models (especially those from Bartowski)

with turboquant KV

while implementing features not yet present in official llama.cpp (which supports many devices, tools and frameworks).

BEWARE: My fork is likely to be buggy/unstable (but leave me a comment/star/watch if you found it useful :) )


I am vibe coding often (while rarely also coding myself... nowadays manual edits are too expensive)


Tools used: my own coder, claude, hermes, gemini, my own brain.

Clone this wiki locally