Uploaded August 2026 | Updated September 2026, 2 weeks ago
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io
Scaling in Kubernetes Safely on On-Prem KaaS Across 1,300+ Clusters and 40,000+ Nodes - Shota Yoshimura, LY Corporation
Many Kubernetes talks focus on scaling out to support growth, but large platforms also need safe ways to scale in. This is especially true for on-premises Kubernetes as a Service, where unused node capacity affects hardware procurement, rack capacity, and long-term operations.
This talk shares how we built a custom controller to consolidate underutilized worker nodes safely across 1,300+ clusters and 40,000+ nodes. At this scale, underutilized nodes were no longer isolated inefficiencies but a meaningful capacity problem. We needed a more conservative model than a general-purpose autoscaler could provide, so the controller uses observed node usage, seven-day peak metrics, minimum replica guarantees, and machine-group-level safety checks before taking action.
We will also cover safeguards including graceful shutdown expectations, DaemonSet termination ordering, and node deletion behavior designed to minimize user impact.
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io
Scaling in Kubernetes Safely on On-Prem KaaS Across 1,300+ Clusters and 40,000+ Nodes - Shota Yoshimura, LY Corporation
Many Kubernetes talks focus on scaling out to support growth, but large platforms also need safe ways to scale in. This is especially true for on-premises Kubernetes as a Service, where unused node capacity affects hardware procurement, rack capacity, and long-term operations.
This talk shares how we built a custom controller to consolidate underutilized worker nodes safely across 1,300+ clusters and 40,000+ nodes. At this scale, underutilized nodes were no longer isolated inefficiencies but a meaningful capacity problem. We needed a more conservative model than a general-purpose autoscaler could provide, so the controller uses observed node usage, seven-day peak metrics, minimum replica guarantees, and machine-group-level safety checks before taking action.
We will also cover safeguards including graceful shutdown expectations, DaemonSet termination ordering, and node deletion behavior designed to minimize user impact.










