Uploaded July 2025 | Updated September 2026, 2 weeks ago
AI workloads require increasing scale for both compute and data, as well as significant heterogeneity across workloads, models, data types, and hardware accelerators. As a consequence, the software stack for running compute-intensive AI workloads is fragmented and rapidly evolving. Companies that productionize AI end up building large AI platform teams to manage these workloads. However, within the fragmented landscape, common patterns are beginning to emerge. This talk describes a popular software stack combining Kubernetes, Ray, PyTorch, and vLLM. It describes the role of each of these frameworks, how they operate together, and illustrates this combination with case studies from Pinterest, Uber, and Roblox as well as from today’s most popular post-training frameworks.
🧠 This talk covers:
Evolving workload types: multimodal data, agentic inference, post-training with RL
Real-world examples from Uber, Pinterest, Roblox, and DeepSeek
Why traditional stacks fail, and how Ray enables new patterns
What’s next for distributed AI systems
⏱️ Chapters
0:00 – Introduction & Background at UC Berkeley
0:51 – Why Ray: Scaling Was the Research Bottleneck
1:45 – Early Adoption: Uber, Pinterest, Ant Group
2:30 – GenAI Inflection Point for Distributed Compute
3:25 – Workload Shift #1: GPU + Multimodal Data Processing
5:10 – From SQL on CPUs to Inference on GPUs
6:35 – Workload Shift #2: Agentic AI & System Complexity
8:15 – Deployment Complexity: APIs, Models, Accelerators
9:10 – Workload Shift #3: Rise of Post-Training & RL
10:55 – A New Stack: Inference, Environment, and Training Loops
13:00 – Hardware Bottlenecks and Scaling Decisions
14:10 – Layer 1: Training & Inference Frameworks (PyTorch, vLLM)
15:30 – Transformer-Specific Optimizations (Speculative Decoding)
16:45 – Layer 2: Distributed Compute Engines (Ray, Spark)
18:00 – Layer 3: Container Orchestration (Kubernetes)
19:15 – Dynamic Interactions Between Layers (e.g. Autoscaling)
20:00 – Infra Case Studies: Amazon, Pinterest, Instacart
20:30 – Wrap-up
AI workloads require increasing scale for both compute and data, as well as significant heterogeneity across workloads, models, data types, and hardware accelerators. As a consequence, the software stack for running compute-intensive AI workloads is fragmented and rapidly evolving. Companies that productionize AI end up building large AI platform teams to manage these workloads. However, within the fragmented landscape, common patterns are beginning to emerge. This talk describes a popular software stack combining Kubernetes, Ray, PyTorch, and vLLM. It describes the role of each of these frameworks, how they operate together, and illustrates this combination with case studies from Pinterest, Uber, and Roblox as well as from today’s most popular post-training frameworks.
🧠 This talk covers:
Evolving workload types: multimodal data, agentic inference, post-training with RL
Real-world examples from Uber, Pinterest, Roblox, and DeepSeek
Why traditional stacks fail, and how Ray enables new patterns
What’s next for distributed AI systems
⏱️ Chapters
0:00 – Introduction & Background at UC Berkeley
0:51 – Why Ray: Scaling Was the Research Bottleneck
1:45 – Early Adoption: Uber, Pinterest, Ant Group
2:30 – GenAI Inflection Point for Distributed Compute
3:25 – Workload Shift #1: GPU + Multimodal Data Processing
5:10 – From SQL on CPUs to Inference on GPUs
6:35 – Workload Shift #2: Agentic AI & System Complexity
8:15 – Deployment Complexity: APIs, Models, Accelerators
9:10 – Workload Shift #3: Rise of Post-Training & RL
10:55 – A New Stack: Inference, Environment, and Training Loops
13:00 – Hardware Bottlenecks and Scaling Decisions
14:10 – Layer 1: Training & Inference Frameworks (PyTorch, vLLM)
15:30 – Transformer-Specific Optimizations (Speculative Decoding)
16:45 – Layer 2: Distributed Compute Engines (Ray, Spark)
18:00 – Layer 3: Container Orchestration (Kubernetes)
19:15 – Dynamic Interactions Between Layers (e.g. Autoscaling)
20:00 – Infra Case Studies: Amazon, Pinterest, Instacart
20:30 – Wrap-up










