How the VAST AI Operating System Powers a Dynamic Data Plane for Ray | Ray Summit 2025 @anyscale
How the VAST AI Operating System Powers a Dynamic Data Plane for Ray | Ray Summit 2025  @anyscale
Uploaded December 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Glenn Lockwood from VAST Data shares how Ray’s new Label Selector API dramatically simplifies scheduling on heterogeneous GPU and accelerator clusters—eliminating workarounds and giving users fine-grained control over resource placement in distributed Ray applications.

He begins by outlining the challenge: acquiring the right accelerator resources (GPU tier, topology, interconnect, memory profile, CPU-to-GPU ratio, etc.) in a heterogeneous cluster is often difficult. Historically, users had to rely on fragile hacks such as custom resource tags or accelerator_type annotations to target specific hardware—solutions that don’t scale, break across environments, or cause scheduling failures.

Glenn introduces Ray’s Label Selector API as a robust, intuitive solution. This new API allows developers to schedule tasks, actors, and placement groups based on Ray node labels—labels that can be defined at RayCluster creation time or discovered automatically by Ray.

Key capabilities highlighted in the session include:

Support for both static and auto-scaling RayClusters

Per-bundle label selectors for precise placement inside multi-node workloads

Fallback strategies to maintain reliability under resource scarcity

Full integration with the Anyscale platform, Ray Dashboard, and KubeRay

Identical behavior across all deployment environments—from local clusters to production-grade autoscaling fleets

Glenn walks through common real-world use cases—such as selecting GPU types, restricting workloads to nodes with local NVMe, matching topology-specific requirements, or isolating latency-sensitive inference—and shows how the Label Selector API makes these constraints trivial to express.

The talk concludes with a live demo demonstrating how this new API enhances flexibility, reliability, and developer experience when scheduling Ray workloads on heterogeneous infrastructure.

Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
How the VAST AI Operating System Powers a Dynamic Data Plane for Ray | Ray Summit 2025Scaling Ray Train to 10K Kubernetes Nodes on GKE | Ray Summit 2024Motional’s Blueprint for High-Performance ML Systems in Autonomous Driving | Ray Summit 2025How Roblox Scaled Machine Learning by Leveraging Ray for Efficient Batch Inference | Ray Summit 2024Ray Train: Distributed Solutions for Removing Training Bottlenecks | Ray Summit 2025Pricing and Packaging Your AI Products for Scale | Ray Summit 2024Building a Multimodal Video Processing Pipeline with RayHow The AI Institute is Revolutionizing Robotics ML Training | Ray Summit 2024Why context engineering is going to play a big role in AI in the future #aiinfrastructureKubeRay + vLLM at DatalogyAI: Engineering Trillion-Scale Synthetic Data Systems | Ray Summit 2025The emerging OpenSource AI Stack for modern AI workloads ⚡  #aiinfrastructure #aiopsHow DigitalOcean Builds Next-Gen Inference with Ray, vLLM & More | Ray Summit 2025
Anyscale |

How the VAST AI Operating System Powers a Dynamic Data Plane for Ray | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER