Architected a distributed notebook execution environment simulating Kaggle Kernels. Utilized the Kubernetes K8s API for dynamic pod orchestration and WebSockets for real-time streaming of code execution.
Warmth-aware load balancer and reverse proxy for self-hosted LLM inference — routes each request to the Ollama or vLLM backend that already has the model loaded