Reproducible Docker + LiteLLM stack for running Qwen 3.6 27B with MTP on AMD RDNA4 (gfx1201). Verified on Radeon AI PRO R9700 32 GB. Built from llama.cpp PR #22673.
What is the AlanHuang99/qwen3.6-mtp-stack GitHub project? Description: "Reproducible Docker + LiteLLM stack for running Qwen 3.6 27B with MTP on AMD RDNA4 (gfx1201). Verified on Radeon AI PRO R9700 32 GB. Built from llama.cpp PR #22673.". Written in Python. Explain what it does, its main use cases, key features, and who would benefit from using it.
Question is copied to clipboard — paste it after the AI opens.
Clone via HTTPS
Clone via SSH
Download ZIP
Download main.zipReport bugs or request features on the qwen3.6-mtp-stack issue tracker:
Open GitHub Issues