llama-swap config for MTP speculative decoding on AMD RX 7900 XTX (ROCm). Qwen3.6-35B-A3B-MTP and Gemma 4 26B-A4B-MTP with VRAM-tuned context sizing.
What is the blockfeed/llama-swap_homelab GitHub project? Description: "llama-swap config for MTP speculative decoding on AMD RX 7900 XTX (ROCm). Qwen3.6-35B-A3B-MTP and Gemma 4 26B-A4B-MTP with VRAM-tuned context sizing.". Explain what it does, its main use cases, key features, and who would benefit from using it.
Question is copied to clipboard — paste it after the AI opens.
Clone via HTTPS
Clone via SSH
Download ZIP
Download main.zipReport bugs or request features on the llama-swap_homelab issue tracker:
Open GitHub Issues