Uploaded March 2024 | Updated September 2026, 2 weeks ago
Quantization is an excellent technique to compress Large Language Models (LLM) and accelerate their inference.
Following up on part 1 youtu.be/kw7S-3s50uk, we look at and compare more advanced quantization techniques: SmoothQuant, GPTQ, AWQ, HQQ, and the Hugging Face Optimum Intel library based on Intel Neural Compressor and Intel OpenVINO.
Slides: fr.slideshare.net/slideshow/julien-simon-deep-dive-quantizing-llms/270921785
⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. Follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️
00:00 Introduction
00:55 SmoothQuant
07:00 Group-wise Precision Tuning Quantization (GPTQ)
12:35 Activation-aware Weight Quantization (AWQ)
18:10 Half-Quadratic Quantization (HQQ)
23:15 Optimum Intel
25:45 Accelerating Stable Diffusion with Intel OpenVINO
Quantization is an excellent technique to compress Large Language Models (LLM) and accelerate their inference.
Following up on part 1 youtu.be/kw7S-3s50uk, we look at and compare more advanced quantization techniques: SmoothQuant, GPTQ, AWQ, HQQ, and the Hugging Face Optimum Intel library based on Intel Neural Compressor and Intel OpenVINO.
Slides: fr.slideshare.net/slideshow/julien-simon-deep-dive-quantizing-llms/270921785
⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. Follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️
00:00 Introduction
00:55 SmoothQuant
07:00 Group-wise Precision Tuning Quantization (GPTQ)
12:35 Activation-aware Weight Quantization (AWQ)
18:10 Half-Quadratic Quantization (HQQ)
23:15 Optimum Intel
25:45 Accelerating Stable Diffusion with Intel OpenVINO










