Run llama.cpp across two GPUs of different vendors at once (AMD ROCm/HIP or Vulkan + NVIDIA CUDA) in one llama-server process, no RPC server. Windows (PowerShell) and Linux (bash) implementations, shared docs on ROCm memory bugs and benchmarks.
What is the daimonionnn/multi-gpu-llm-toolkit GitHub project? Description: "Run llama.cpp across two GPUs of different vendors at once (AMD ROCm/HIP or Vulkan + NVIDIA CUDA) in one llama-server process, no RPC server. Windows (PowerShell) and Linux (bash) implementations, shared docs on ROCm memory bugs and benchmarks.". Written in Shell. Explain what it does, its main use cases, key features, and who would benefit from using it.
Question is copied to clipboard — paste it after the AI opens.
Clone via HTTPS
Clone via SSH
Download ZIP
Download main.zipReport bugs or request features on the multi-gpu-llm-toolkit issue tracker:
Open GitHub Issues