1 repository on SrcLog
The flame graph for "why won't this model fit": live-trace local-LLM VRAM (weights vs KV cache) and predict max context before OOM. Zero-dependency Go, AMD/ROCm first-class, Ollama & llama.cpp.