GitHub

Andrey Krisanov is a Staff Software Engineer at Severstal, where he builds production LLM inference platforms and AI infrastructure on Kubernetes and NVIDIA GPUs.

He leads the architecture and technical evolution of a shared model-serving platform for enterprise AI products and coding agents. His work covers vLLM-based model serving, inference routing, performance and reliability, observability, GPU capacity planning, controlled model rollouts and rollbacks, and multi-data-center resilience.

His main engineering interests are inference control planes, serving performance, and distributed systems.

Before specialising in AI infrastructure, Andrey built and scaled backend and platform systems across SaaS, fintech, data privacy, and high-traffic consumer products.

Check out his article on Kubernetes Model Serving in 2026 and more engineering notes on his technical blog.

Read the original on github.com ↗