inferhost
amirrouh/inferhost
Python
Self-hosted multi-modal AI server: chat/vision LLMs (Qwen, Llama, Gemma, DeepSeek), text-to-speech (OuteTTS, Qwen3-TTS) and image generation (SDXL, Flux, Z-Image, Qwen-Image) behind one OpenAI-compatible endpoint. Wraps llama.cpp + stable-diffusion.cpp, auto-fetches GGUF models from Hugging Face, VRAM hot-swap, speculative decoding, no compiling.