The Evolution of Multi-GPU Inference in vLLM | Ray Summit 2024 @anyscale
The Evolution of Multi-GPU Inference in vLLM | Ray Summit 2024  @anyscale
Uploaded October 2024 | Updated September 2026, 1 week ago
At Ray Summit 2024, Sangbin Cho from Anyscale and Murali Andoorveedu from Centml explore the development and future of multi-GPU inference in vLLM. Their presentation focuses on the unique challenges posed by distributed inference for large language models, distinguishing it from distributed training.

Cho and Andoorveedu delve into various parallelism strategies, including tensor parallelism, pipeline parallelism, and expert parallelism, explaining how each works in detail. Using vLLM as a case study, they demonstrate how to construct an optimized architecture for efficient distributed inference. This talk provides valuable insights into the complexities of scaling LLM inference across multiple GPUs and offers a glimpse into the roadmap for future developments in this critical area of AI infrastructure.

--

Interested in more?
- Watch the full Day 1 Keynote: youtu.be/jwZHJthQvXo
- Watch the full Day 2 Keynote youtu.be/Lury2ad6KG8

--

đź”— Connect with us:
- Subscribe to our YouTube channel: youtube.com/@anyscale
- Twitter: https://x.com/anyscalecompute
- LinkedIn: linkedin.com/company/joinanyscale
- Website: anyscale.com
The Evolution of Multi-GPU Inference in vLLM | Ray Summit 2024The LLM-Cloud Synergy: NebiusAIs Insider Perspective | Ray Summit 2024How Coinbase Uses Ray, vLLM & LiteLLM to Power Secure LLM Services | Ray Summit 2025How Datadog is Transforming Time Series Forecasting with Toto | Ray Summit 2024MLOps with Ray on Anyscale | Ray Summit 2025Distributed Embeddings at Scale: Processing 10M+ Rows/ Day with Ray, GPUs & Qdrant | Ray Summit 2025Applied Intuition’s Blueprint for Scalable RL + Batch Inference | Ray Summit 2025Building Scalable Cross-Modal Search with Ray | Ray Summit 2024Ray at IBM: Transforming Large-Scale Data Processing for AI and Science | Ray Summit 2024Elastic Expert Parallelism for vLLM | Ray Summit 2025Apple’s Approach to Scalable Machine Learning Infrastructure on Ray | Ray Summit 2025How Uber Optimize Marketplaces with Ray | Ray Summit 2024
Anyscale |

The Evolution of Multi-GPU Inference in vLLM | Ray Summit 2024

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER