How Runhouse Orchestrates Multi-Cluster Ray Workloads | Ray Summit 2025 @anyscale
How Runhouse Orchestrates Multi-Cluster Ray Workloads | Ray Summit 2025  @anyscale
Uploaded December 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Donny Greenberg from Runhouse shares how Kubetorch introduces a next-generation, Ray-inspired distributed programming paradigm for Kubernetes—enabling teams to build complex AI workloads “serverlessly” in Python, without drowning in YAML or manually managing Ray clusters.

He begins by outlining a trend across modern ML infrastructure: Ray has become a first-class distributed computing primitive, with teams composing multi-stage inference, training, and reinforcement learning workloads on Kubernetes alongside other compute backends. But this shift has surfaced new challenges—rapid debugging, fluent programmatic orchestration, fault-tolerant workflows, and the growing expectation that Ray clusters should be ephemeral and created per task, à la serverless computing.

Kubetorch aims to fill this gap.

Donny introduces Kubetorch as a programming model that extends Ray’s familiar Task and Actor abstractions to Kubernetes-native resources. In Kubetorch:

An Actor represents not just a process, but a full Kubernetes resource—including a KubeRay RayCluster.

Entire Ray programs, services, and pipelines can be composed as higher-order workflows directly in Python.

Teams get a dramatically improved developer experience with fast iteration, fault tolerance, and minimal operational overhead.

Workloads scale elastically and portably across Kubernetes environments.

Incremental adoption is natural—Kubetorch can wrap existing Ray workloads while offering a smoother serverless-like experience.

The session highlights how Kubetorch brings serverless Ray to life, offering instant cluster provisioning, ephemeral execution, and scalable workflow composition—all while staying grounded in the Ray programming model that ML practitioners already know.

Attendees will walk away with a clear picture of how Kubetorch simplifies distributed ML workflows, closes critical gaps in the Ray ecosystem, and makes Kubernetes-native AI infrastructure more developer-friendly than ever.

Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
How Runhouse Orchestrates Multi-Cluster Ray Workloads | Ray Summit 2025How vLLM and Ray Work TogetherCoinbases ML Training Evolution: From Sagemaker to Ray | Ray Summit 2024Secure & Scalable AI on Ray + Kubernetes: Google’s Decoupled Agent Pattern | Ray Summit 2025vLLM TPU: A new unified-backend supporting Pytorch and JAX natively on TPU | Ray Summit 2025Hugging Face + vLLM: One Model Definition to Rule Them All | Ray Summit 2025Matrix: Reliable Framework for Data-Centric Experimentation at Scale  | Ray Summit 2025From Spark to Ray: CSSs Data Revolution with Daft | Ray Summit 2024How Rubrik Unlocked AI at Scale with Ray Serve | Ray Summit 2024How Workday Achieved 50x Cheaper Model Serving with Ray Serve | Ray Summit 2025ByteDances Platform for Reinforcement Learning from Human Feedback | Ray Summit 2024Dynamic Scheduling for Large Language Model Serving | Ray Summit 2024
Anyscale |

How Runhouse Orchestrates Multi-Cluster Ray Workloads | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER