Uploaded August 2026 | Updated September 2026, 2 weeks ago
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io
Topology-Aware Scheduling for AI Training & Inference With Kueue - Michał Woźniak, Google & Wei Huang, Meta
Kueue, the Kubernetes-native workload orchestrator, helps run AI workloads at scale with multi-tenant quota management and advanced scheduling.
In this session, we show through a case study at the Meta Superintelligence Lab how Kueue provided the right quota and scheduling layer for large-scale AI training and inference across complex, heterogeneous GPU topologies, enabling a shift from an in-house platform to an open-source stack on Kubernetes.
We then dive into Kueue’s multi-layer Topology-Aware Scheduling (TAS) for modern clusters with hierarchical network fabrics. We cover the main concepts and algorithms behind Hierarchical TAS, and show why multi-layer TAS was critical for this transition.
We close with other cutting-edge Kueue capabilities, including Workload-Aware Scheduler integration, elastic workloads, and multi-cluster job scheduling with MultiKueue.
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io
Topology-Aware Scheduling for AI Training & Inference With Kueue - Michał Woźniak, Google & Wei Huang, Meta
Kueue, the Kubernetes-native workload orchestrator, helps run AI workloads at scale with multi-tenant quota management and advanced scheduling.
In this session, we show through a case study at the Meta Superintelligence Lab how Kueue provided the right quota and scheduling layer for large-scale AI training and inference across complex, heterogeneous GPU topologies, enabling a shift from an in-house platform to an open-source stack on Kubernetes.
We then dive into Kueue’s multi-layer Topology-Aware Scheduling (TAS) for modern clusters with hierarchical network fabrics. We cover the main concepts and algorithms behind Hierarchical TAS, and show why multi-layer TAS was critical for this transition.
We close with other cutting-edge Kueue capabilities, including Workload-Aware Scheduler integration, elastic workloads, and multi-cluster job scheduling with MultiKueue.

