2 repositories on SrcLog
A high-throughput and memory-efficient inference and serving engine for LLMs
Intelligent Mixture-of-Models Router for Efficient Heterogeneous LLMs Inference