Multi-slot LLM inference on AMD Strix Halo: recipes + honest benchmarks (236 tok/s @ 32 streams, llama.cpp Vulkan)
What is the botAGI/strix-halo-multislot GitHub project? Description: "Multi-slot LLM inference on AMD Strix Halo: recipes + honest benchmarks (236 tok/s @ 32 streams, llama.cpp Vulkan)". Written in Python. Explain what it does, its main use cases, key features, and who would benefit from using it.
Question is copied to clipboard — paste it after the AI opens.
Clone via HTTPS
Clone via SSH
Download ZIP
Download main.zipReport bugs or request features on the strix-halo-multislot issue tracker:
Open GitHub Issues