The Emerging Stack for AI Compute: Kubernetes + Ray + PyTorch + vLLM @arizeai
The Emerging Stack for AI Compute: Kubernetes + Ray + PyTorch + vLLM  @arizeai
Uploaded July 2025 | Updated September 2026, 2 weeks ago
AI workloads require increasing scale for both compute and data, as well as significant heterogeneity across workloads, models, data types, and hardware accelerators. As a consequence, the software stack for running compute-intensive AI workloads is fragmented and rapidly evolving. Companies that productionize AI end up building large AI platform teams to manage these workloads. However, within the fragmented landscape, common patterns are beginning to emerge. This talk describes a popular software stack combining Kubernetes, Ray, PyTorch, and vLLM. It describes the role of each of these frameworks, how they operate together, and illustrates this combination with case studies from Pinterest, Uber, and Roblox as well as from today’s most popular post-training frameworks.

🧠 This talk covers:
Evolving workload types: multimodal data, agentic inference, post-training with RL
Real-world examples from Uber, Pinterest, Roblox, and DeepSeek
Why traditional stacks fail, and how Ray enables new patterns
What’s next for distributed AI systems

⏱️ Chapters
0:00 – Introduction & Background at UC Berkeley
0:51 – Why Ray: Scaling Was the Research Bottleneck
1:45 – Early Adoption: Uber, Pinterest, Ant Group
2:30 – GenAI Inflection Point for Distributed Compute
3:25 – Workload Shift #1: GPU + Multimodal Data Processing
5:10 – From SQL on CPUs to Inference on GPUs
6:35 – Workload Shift #2: Agentic AI & System Complexity
8:15 – Deployment Complexity: APIs, Models, Accelerators
9:10 – Workload Shift #3: Rise of Post-Training & RL
10:55 – A New Stack: Inference, Environment, and Training Loops
13:00 – Hardware Bottlenecks and Scaling Decisions
14:10 – Layer 1: Training & Inference Frameworks (PyTorch, vLLM)
15:30 – Transformer-Specific Optimizations (Speculative Decoding)
16:45 – Layer 2: Distributed Compute Engines (Ray, Spark)
18:00 – Layer 3: Container Orchestration (Kubernetes)
19:15 – Dynamic Interactions Between Layers (e.g. Autoscaling)
20:00 – Infra Case Studies: Amazon, Pinterest, Instacart
20:30 – Wrap-up
The Emerging Stack for AI Compute: Kubernetes + Ray + PyTorch + vLLMIntroducing Strands Agents: AWS on How To Improve the Customer Experience with AI AgentsAI Agent Mastery Certification Course: Module 4 – Tools & MCPBuilding Responsible AI in Highly Regulated Industries | BlackRock | Arize Observe 2026How to Catch AI Agent Failures Fast with Monitors and Alerts | Ep. 12Nous Research’s Hermes Agent: The Case for Open Models in Production | Arize Observe 2026AI Agent Mastery Certification Course: Module 6 – Agent EvaluationHow Handshake Builds a Future-Proofed, Adaptable LLM Orchestration Stack Amid Rapid Evolution In AIFrom 5 Minutes to 2 Days: The Scale of AI Agent AutonomyAI Agents Certification Course: Module 1 – Introduction to AgentsWhy Language Models Hallucinate OpenAI Paper, Explained by Author Santosh Vempala of GeorgiaTechVLAA Thinker: Dive Into Vision Language Models with Inherent Reasoning
Arize AI |

The Emerging Stack for AI Compute: Kubernetes + Ray + PyTorch + vLLM

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER