Secure & Scalable AI on Ray + Kubernetes: Google’s Decoupled Agent Pattern | Ray Summit 2025 @anyscale
Secure & Scalable AI on Ray + Kubernetes: Google’s Decoupled Agent Pattern | Ray Summit 2025  @anyscale
Uploaded November 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Alex Bulankou and Brandon Royal from Google share how to bring agentic AI systems out of the lab and into production through the Decoupled Agent Pattern—a scalable, resilient, and secure architecture built on Ray and Kubernetes.

They begin by outlining the core production challenge of agentic systems: integrating LLMs, tools, and long-lived stateful agents while ensuring security, elasticity, and high-throughput execution. Traditional architectures struggle to balance these constraints. The Decoupled Agent Pattern solves this by cleanly separating the stateful agent logic from the stateless, scalable tools it invokes.

At the heart of this pattern:

The agent’s core logic runs as a durable Ray Actor, with lifecycle and placement managed by Ray’s Global Control Store (GCS) for high availability.

Tools are executed as thousands of stateless Ray Tasks, enabling massive parallelism and elasticity.

Untrusted or dynamically generated code runs in gVisor sandboxes, providing kernel-level isolation without compromising throughput—made possible through Kubernetes’ secure runtime capabilities.

Alex and Brandon demonstrate the architecture with a series of live scenarios, including a financial analysis agent running on a Ray cluster on Google Kubernetes Engine (GKE).

They then show how the architecture leverages deep Kubernetes-native integrations:

KubeRay’s topology-aware placement allows Ray to understand node-level characteristics, enabling optimal scheduling.

This unlocks intelligent capacity management with tools like Kueue for cost-efficient batch scheduling.

And it provides a clear pathway to mission-critical resilience, supporting zero-downtime upgrades and fault-tolerant agent execution.

Attendees will leave with a practical blueprint for deploying agentic AI systems in production—combining Ray’s distributed computing strengths with Kubernetes’ security and orchestration capabilities to build scalable, resilient, and secure agentic runtimes.

Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
Secure & Scalable AI on Ray + Kubernetes: Google’s Decoupled Agent Pattern | Ray Summit 2025vLLM TPU: A new unified-backend supporting Pytorch and JAX natively on TPU | Ray Summit 2025Hugging Face + vLLM: One Model Definition to Rule Them All | Ray Summit 2025Matrix: Reliable Framework for Data-Centric Experimentation at Scale  | Ray Summit 2025From Spark to Ray: CSSs Data Revolution with Daft | Ray Summit 2024How Rubrik Unlocked AI at Scale with Ray Serve | Ray Summit 2024How Workday Achieved 50x Cheaper Model Serving with Ray Serve | Ray Summit 2025ByteDances Platform for Reinforcement Learning from Human Feedback | Ray Summit 2024Dynamic Scheduling for Large Language Model Serving | Ray Summit 2024How Prime Intellect Builds Scalable Infrastructure for Agentic RL | Ray Summit 2025BentoML or RayServe, You Can Choose Both with BentoRayHow KubeRay Is Evolving for Massive AI Workloads  | Ray Summit 2025
Anyscale |

Secure & Scalable AI on Ray + Kubernetes: Google’s Decoupled Agent Pattern | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER