Uploaded May 2026 | Updated September 2026, 3 weeks ago
Intelligent Load Balancing in Kubernetes
Gaurav Nanda and Vincent Cheng, Databricks
Kubernetes relies on kube-proxy and DNS for simple Layer 4 load balancing, which works for short-lived HTTP traffic but fails for persistent connections and high-throughput gRPC workloads. With thousands of requests multiplexed over a single TCP connection, clusters often see uneven load, pod hot-spotting, and rising tail latency.
This talk presents a client-side, control-plane-driven approach that removes kube-proxy and DNS from the data path. A lightweight control plane tracks Service and EndpointSlice updates, while client libraries receive live endpoint changes through xDS and make per-request routing decisions at Layer 7. We show how strategies like Power of Two Choices and zone-affinity routing improve load balance, stabilize tail latency, and reduce resource waste in production.
SREs and platform engineers will learn why default Kubernetes routing breaks down, how to design intelligent client-side load balancing, and what operational challenges emerge when deploying these systems at scale.
View the full SREcon26 Americas program at usenix.org/conference/srecon26americas/program
Intelligent Load Balancing in Kubernetes
Gaurav Nanda and Vincent Cheng, Databricks
Kubernetes relies on kube-proxy and DNS for simple Layer 4 load balancing, which works for short-lived HTTP traffic but fails for persistent connections and high-throughput gRPC workloads. With thousands of requests multiplexed over a single TCP connection, clusters often see uneven load, pod hot-spotting, and rising tail latency.
This talk presents a client-side, control-plane-driven approach that removes kube-proxy and DNS from the data path. A lightweight control plane tracks Service and EndpointSlice updates, while client libraries receive live endpoint changes through xDS and make per-request routing decisions at Layer 7. We show how strategies like Power of Two Choices and zone-affinity routing improve load balance, stabilize tail latency, and reduce resource waste in production.
SREs and platform engineers will learn why default Kubernetes routing breaks down, how to design intelligent client-side load balancing, and what operational challenges emerge when deploying these systems at scale.
View the full SREcon26 Americas program at usenix.org/conference/srecon26americas/program










