Optimizing vLLM Performance through Quantization | Ray Summit 2024 @anyscale
Optimizing vLLM Performance through Quantization | Ray Summit 2024  @anyscale
Uploaded October 2024 | Updated September 2026, 2 weeks ago
At Ray Summit 2024, Michael Goin and Robert Shaw from Neural Magic delve into the world of model quantization for vLLM deployments. Their presentation focuses on vLLM's support for various quantization methods, including FP8, INT8, and INT4, which are crucial for reducing memory usage and enhancing generation speed.

In the talk, Goin and Shaw explain the internal mechanisms of how vLLM leverages quantization to accelerate models. They also provide practical guidance on applying these quantization techniques to custom models using vLLM's llm-compressor framework. This talk offers valuable insights for developers and organizations looking to optimize their LLM deployments, balancing performance and resource efficiency in large-scale AI applications.

--

Interested in more?
- Watch the full Day 1 Keynote: youtu.be/jwZHJthQvXo
- Watch the full Day 2 Keynote youtu.be/Lury2ad6KG8

--

đź”— Connect with us:
- Subscribe to our YouTube channel: youtube.com/@anyscale
- Twitter: https://x.com/anyscalecompute
- LinkedIn: linkedin.com/company/joinanyscale
- Website: anyscale.com
Optimizing vLLM Performance through Quantization | Ray Summit 2024How IBM Research Achieved vLLM Platform Portability with Triton Autotuning | Ray Summit 2024Greg Brockman on Founding OpenAI and Systems for AI | Ray Summit 2022Wisedocs’ Journey: Rebuilding & Accelerating ML with KubeRay | Ray Summit 2025[Ray Meetup] Ray + vLLM in Action: Lessons from Pinterest and Large Scale Distributed InferenceHow Zoox Built a Reliable, High-Velocity Model Serving Platform with Ray Serve | Ray Summit 2025Scaling Machine Learning at Tripadvisor: Our Journey with Ray and Anyscale | Ray Summit 2025AWS + vLLM: Building the Future of Open, Fast LLM Serving | Ray Summit 2025Ray + vLLM  Efficient Multi Node Orchestration for Sparse MoE Model Serving | Ray Summit 2025Hybrid RL + Imitation Learning for Robotics with Ray at RAI InstituteHow Runhouse Orchestrates Multi-Cluster Ray Workloads | Ray Summit 2025How vLLM and Ray Work Together
Anyscale |

Optimizing vLLM Performance through Quantization | Ray Summit 2024

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER