1 repository on SrcLog
Local AMD GPU LLM inference engine (HIP/ROCm) — quantized GGUF, Q4_0/K-quants, flash attention, KV cache. ~515 tok/s decode on Qwen2.5-0.5B Q4_0 (RX 7900 XT).