Uploaded August 2026 | Updated September 2026, 2 weeks ago
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io
Shared GPU Scheduling & Proactive Autoscaling: A Production Blueprint for 1000+ GPUs - Jeonghyun Kim, SNOW Corporation & Reza Jelveh, Dynamia.ai
SNOW Corp. operates 1,000+ A100 GPUs serving 200 million users across three top-ranked GenAI applications (Snow, Epik, B612), handling 1,200+ AI workflows subject to extreme traffic volatility from viral AI trends.
The core bottleneck was Kubernetes' native GPU scheduling, which treats GPUs as atomic resources — forcing a 2x over-provisioning penalty on Train-to-Inference pipelines with no reliable visibility into actual GPU saturation.
This talk covers integrating HAMi for vGPU virtualization and extending KEDA with a custom Consumer Saturation metric for proactive autoscaling, hiding warm-up latency by scaling before traffic arrives.
We’ll detail the implementation: scheduler config, Prometheus metrics, and multi-region scaling via Helm GitOps. Results: over 50% GPU waste cut and 91% faster recovery during surges. You’ll get a production blueprint for efficient, shared GPU platforms.
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io
Shared GPU Scheduling & Proactive Autoscaling: A Production Blueprint for 1000+ GPUs - Jeonghyun Kim, SNOW Corporation & Reza Jelveh, Dynamia.ai
SNOW Corp. operates 1,000+ A100 GPUs serving 200 million users across three top-ranked GenAI applications (Snow, Epik, B612), handling 1,200+ AI workflows subject to extreme traffic volatility from viral AI trends.
The core bottleneck was Kubernetes' native GPU scheduling, which treats GPUs as atomic resources — forcing a 2x over-provisioning penalty on Train-to-Inference pipelines with no reliable visibility into actual GPU saturation.
This talk covers integrating HAMi for vGPU virtualization and extending KEDA with a custom Consumer Saturation metric for proactive autoscaling, hiding warm-up latency by scaling before traffic arrives.
We’ll detail the implementation: scheduler config, Prometheus metrics, and multi-region scaling via Helm GitOps. Results: over 50% GPU waste cut and 91% faster recovery during surges. You’ll get a production blueprint for efficient, shared GPU platforms.










