Uploaded October 2024 | Updated September 2026, 2 weeks ago
At Ray Summit 2024, Michael Goin and Robert Shaw from Neural Magic delve into the world of model quantization for vLLM deployments. Their presentation focuses on vLLM's support for various quantization methods, including FP8, INT8, and INT4, which are crucial for reducing memory usage and enhancing generation speed.
In the talk, Goin and Shaw explain the internal mechanisms of how vLLM leverages quantization to accelerate models. They also provide practical guidance on applying these quantization techniques to custom models using vLLM's llm-compressor framework. This talk offers valuable insights for developers and organizations looking to optimize their LLM deployments, balancing performance and resource efficiency in large-scale AI applications.
--
Interested in more?
- Watch the full Day 1 Keynote: youtu.be/jwZHJthQvXo
- Watch the full Day 2 Keynote youtu.be/Lury2ad6KG8
--
đź”— Connect with us:
- Subscribe to our YouTube channel: youtube.com/@anyscale
- Twitter: https://x.com/anyscalecompute
- LinkedIn: linkedin.com/company/joinanyscale
- Website: anyscale.com
At Ray Summit 2024, Michael Goin and Robert Shaw from Neural Magic delve into the world of model quantization for vLLM deployments. Their presentation focuses on vLLM's support for various quantization methods, including FP8, INT8, and INT4, which are crucial for reducing memory usage and enhancing generation speed.
In the talk, Goin and Shaw explain the internal mechanisms of how vLLM leverages quantization to accelerate models. They also provide practical guidance on applying these quantization techniques to custom models using vLLM's llm-compressor framework. This talk offers valuable insights for developers and organizations looking to optimize their LLM deployments, balancing performance and resource efficiency in large-scale AI applications.
--
Interested in more?
- Watch the full Day 1 Keynote: youtu.be/jwZHJthQvXo
- Watch the full Day 2 Keynote youtu.be/Lury2ad6KG8
--
đź”— Connect with us:
- Subscribe to our YouTube channel: youtube.com/@anyscale
- Twitter: https://x.com/anyscalecompute
- LinkedIn: linkedin.com/company/joinanyscale
- Website: anyscale.com



![[Ray Meetup] Ray + vLLM in Action: Lessons from Pinterest and Large Scale Distributed Inference
Listen in to our Ray Meetup where we explored batch inference at scale with Ray and vLLM! Learn how Pinterest scales batch inference using Ray, and get a first look at Anyscale’s latest tools—Ray Serve and Data LLM—for orchestrating large-scale LLM inference. We’ll cover topics like batch inference, prefill-decode disaggregation, DP/EP parallelism, and custom request routing.
Speakers:
Chia-Wei Chen, Software Engineer, ML Training Infra, Pinterest
Kourosh Hakhamaneshi, AI Lead, Anyscale [Ray Meetup] Ray + vLLM in Action: Lessons from Pinterest and Large Scale Distributed Inference](https://i.ytimg.com/vi/HDSy09hrm2I/mqdefault.jpg)






