Self-hosted multi-modal AI server: chat/vision LLMs (Qwen, Llama, Gemma, DeepSeek), text-to-speech (OuteTTS, Qwen3-TTS) and image generation (SDXL, Flux, Z-Image, Qwen-Image) behind one OpenAI-compatible endpoint. Wraps llama.cpp + stable-diffusion.cpp, auto-fetches GGUF models from Hugging Face, VRAM hot-swap, speculative decoding, no compiling.
What is the amirrouh/inferhost GitHub project? Description: "Self-hosted multi-modal AI server: chat/vision LLMs (Qwen, Llama, Gemma, DeepSeek), text-to-speech (OuteTTS, Qwen3-TTS) and image generation (SDXL, Flux, Z-Image, Qwen-Image) behind one OpenAI-compatible endpoint. Wraps llama.cpp + stable-diffusion.cpp, auto-fetches GGUF models from Hugging Face, VRAM hot-swap, speculative decoding, no compiling.". Written in Python. Explain what it does, its main use cases, key features, and who would benefit from using it.
Question is copied to clipboard — paste it after the AI opens.
Clone via HTTPS
Clone via SSH
Download ZIP
Download master.zipReport bugs or request features on the inferhost issue tracker:
Open GitHub Issues