Beyond Single-Cluster Limits: Scaling GPU Workloads Across Kubernetes With… K. Das & E. Bayramova @cncf
Beyond Single-Cluster Limits: Scaling GPU Workloads Across Kubernetes With… K. Das & E. Bayramova  @cncf
Uploaded August 2026 | Updated September 2026, 3 weeks ago
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io

Beyond Single-Cluster Limits: Scaling GPU Workloads Across Kubernetes With Virtual Nodes - Kunal Das, Cast AI & Esmira Bayramova, Kimchi

Our ML platform team hit a wall when GPU demand became unpredictable. Fine-tuning jobs queued for hours on busy days while expensive on-prem GPUs sat idle on others. Adding cloud GPU clusters solved capacity but fractured our workflow , developers juggled multiple kubeconfigs, Kueue couldn't see cross-cluster queues,workloads couldn't failover because "multi-cluster" was really just isolated clusters with shared dashboards.

In this talk, we walk through how we used Liqo's virtual node pattern to unify 5 heterogeneous clusters , mixing on-prem A100/ H100 nodes with cloud spot GPU instances from cloud providers , into a single schedulable topology. We'll share how the vanilla Kubernetes scheduler handles placement using standard node selectors and affinities, how Cilium's cluster mesh compared to Liqo's WireGuard-based network fabric for cross-cluster pod traffic, and how we integrated Kueue for unified job queuing across the virtual cluster.
Beyond Single-Cluster Limits: Scaling GPU Workloads Across Kubernetes With… K. Das & E. BayramovaDetecting Compromised CI With eBPF and Cilium Tetragon - Liz Rice, Isovalent at CiscoShared Yet Isolated at Scale: Building Multi-Tenant Inference Platform on… Y. Hiraki & Y. TanakaNeurodiversity Meeting - December 2025The Death of the YAML-Engineer: Engineering Invisible Platf... Abhinav Sharma & Mumshad Mannambeth⚡Lightning Talk: Learning Kubernetes From Logs: Building a Foundation for Future… K. IshiiWhat Is a Virtual Power Plant (VPP) ? : Green Tech and the Modernization of the Grid - L. MalcomSIG API Machinery in the Era of AI: Updates - Federico Bongiovanni, Google CloudRunning Wasm Inside Your Storage Cluster With CSI and Gateway API - Ho Kim & SangJoon ParkCNCFs Jonathan Bryce on Why Software, Not Hardware, Drives AI EfficiencyHow Community Infrastructure Helped Grow Taiwan’s Open Source Contributors - C. Kuo & C. YangIoT Compute Unleashed: Transparent Offload From ESP32 To E... Anastassios Nanos & Charalampos Mainas
CNCF [Cloud Native Computing Foundation] |

Beyond Single-Cluster Limits: Scaling GPU Workloads Across Kubernetes With… K. Das & E. Bayramova

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER