Deep Dive: Quantizing Large Language Models, part 1 @juliensimonfr
Deep Dive: Quantizing Large Language Models, part 1  @juliensimonfr
Uploaded March 2024 | Updated September 2026, 2 weeks ago
Quantization is an excellent technique to compress Large Language Models (LLM) and accelerate their inference.

In this video, we discuss model quantization, first introducing what it is, and how to get an intuition of rescaling and the problems it creates. Then we introduce the different types of quantization: dynamic post-training quantization, static post-training quantization, and quantization-aware training. Finally, we start looking at and comparing actual quantization techniques: PyTorch, ZeroQuant, and bitsandbytes.

In part 2 youtu.be/fXBBwCIA0Ds, we look at and compare more advanced quantization techniques: SmoothQuant, GPTQ, AWQ, HQQ, and the Hugging Face Optimum Intel library based on Intel Neural Compressor and Intel OpenVINO.

Slides: fr.slideshare.net/slideshow/julien-simon-deep-dive-quantizing-llms/270921785

⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. Follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️

00:00 Introduction
02:05 What is quantization?
06:50 Rescaling weights and activations
08:17 The mapping function
12:38 Picking the input range
16:15 Getting rid of outliers
19:50 When can we apply quantization?
26:00 Dynamic post-training quantization with PyTorch
28:42 ZeroQuant
34:50 bitsandbytes
Deep Dive: Quantizing Large Language Models, part 1Building a Kanban Task Management Web App from Scratch with ReplitWhy Local Voices Matter in AI Development! 🌍🤖Azure ML: deploy Hugging Face models in minutes!Unlocking Innovation: Why Large Companies Cant Afford to Stall! 🚀Retail AI at the edge - Cisco Live 2025 livestreamArcee AI webinar: build agentic workflows with Arcee OrchestraHomunculus 12B and GLM-4-32B-Base-32K: 2 new Arcee AI research-oriented modelsDeploy Hugging Face models on Google Cloud: from the hub to Inference EndpointsSLM in Action: Local Inference with Arcee Nova 72B and OllamaPhi-2 on Intel Meteor Lake - Physics questionVirtuoso Lite and Virtuoso Medium v2: distilling DeepSeek-V3 to 10B & 32B
Julien Simon |

Deep Dive: Quantizing Large Language Models, part 1

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER