Shared GPU Scheduling & Proactive Autoscaling: A Production Blueprint for 1000+… J. Kim & R. Jelveh @cncf
Shared GPU Scheduling & Proactive Autoscaling: A Production Blueprint for 1000+… J. Kim & R. Jelveh  @cncf
Uploaded August 2026 | Updated September 2026, 2 weeks ago
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io

Shared GPU Scheduling & Proactive Autoscaling: A Production Blueprint for 1000+ GPUs - Jeonghyun Kim, SNOW Corporation & Reza Jelveh, Dynamia.ai

SNOW Corp. operates 1,000+ A100 GPUs serving 200 million users across three top-ranked GenAI applications (Snow, Epik, B612), handling 1,200+ AI workflows subject to extreme traffic volatility from viral AI trends.

The core bottleneck was Kubernetes' native GPU scheduling, which treats GPUs as atomic resources — forcing a 2x over-provisioning penalty on Train-to-Inference pipelines with no reliable visibility into actual GPU saturation.

This talk covers integrating HAMi for vGPU virtualization and extending KEDA with a custom Consumer Saturation metric for proactive autoscaling, hiding warm-up latency by scaling before traffic arrives.

We’ll detail the implementation: scheduler config, Prometheus metrics, and multi-region scaling via Helm GitOps. Results: over 50% GPU waste cut and 91% faster recovery during surges. You’ll get a production blueprint for efficient, shared GPU platforms.
Shared GPU Scheduling & Proactive Autoscaling: A Production Blueprint for 1000+… J. Kim & R. JelvehKeynote: Closing Remarks - Yoshiyuki TabataPolicy & GitOps Unite! Kyverno and Flux Save Cluster City - Cortney Nickerson & Leigh CapiliRook: Intro and Deep Dive With Ceph - Satoru Takeuchi & Dan van der SterContributing to Your Community | CNCF AmbassadorCounting What You Care About in Your Security Data Pipeline - Attila SzakácsCNCF Ambassador Arsh Sharma on CNCF Community⚡Lightning Talk: From First Contribution To Maintainer: An OSS Journey Inside Youki - Yusuke SakuraiWhen GitOps Becomes the Bottleneck: Scaling Argo CD for High-Churn Platforms - S. Kanabar & V. JainWhat the CNCF Ambassador Program Gave Faeka AnsariOpenAI: Winner of the CNCF End User Case Study ContestWhat Being a CNCF Ambassador Means to Knative Maintainer Calum Murray
CNCF [Cloud Native Computing Foundation] |

Shared GPU Scheduling & Proactive Autoscaling: A Production Blueprint for 1000+… J. Kim & R. Jelveh

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER