llama.cpp OpenAI-compatible server on Vulkan for AMD Strix Halo (gfx1151), GGUF weights pinned to GTT not VRAM. Serves poolside Laguna S 2.1, Gemma 4 and Qwen3.6 GGUFs on the stock Vulkan image, plus an opt-in ROCmFP4 + MTP stack (Ubuntu 26.04 + TheRock ROCm 7.13). Docker Compose, with real measured benchmarks.
What is the hec-ovi/llama-vulkan-strix GitHub project? Description: "llama.cpp OpenAI-compatible server on Vulkan for AMD Strix Halo (gfx1151), GGUF weights pinned to GTT not VRAM. Serves poolside Laguna S 2.1, Gemma 4 and Qwen3.6 GGUFs on the stock Vulkan image, plus an opt-in ROCmFP4 + MTP stack (Ubuntu 26.04 + TheRock ROCm 7.13). Docker Compose, with real measured benchmarks.". Written in Python. Explain what it does, its main use cases, key features, and who would benefit from using it.
Question is copied to clipboard — paste it after the AI opens.
Clone via HTTPS
Clone via SSH
Download ZIP
Download main.zipReport bugs or request features on the llama-vulkan-strix issue tracker:
Open GitHub Issues