Deep Dive: Quantizing Large Language Models, part 2 @juliensimonfr
Deep Dive: Quantizing Large Language Models, part 2  @juliensimonfr
Uploaded March 2024 | Updated September 2026, 2 weeks ago
Quantization is an excellent technique to compress Large Language Models (LLM) and accelerate their inference.

Following up on part 1 youtu.be/kw7S-3s50uk, we look at and compare more advanced quantization techniques: SmoothQuant, GPTQ, AWQ, HQQ, and the Hugging Face Optimum Intel library based on Intel Neural Compressor and Intel OpenVINO.

Slides: fr.slideshare.net/slideshow/julien-simon-deep-dive-quantizing-llms/270921785

⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. Follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️

00:00 Introduction
00:55 SmoothQuant
07:00 Group-wise Precision Tuning Quantization (GPTQ)
12:35 Activation-aware Weight Quantization (AWQ)
18:10 Half-Quadratic Quantization (HQQ)
23:15 Optimum Intel
25:45 Accelerating Stable Diffusion with Intel OpenVINO
Deep Dive: Quantizing Large Language Models, part 2Arcee Llama Spark, a better Llama 3.1 #ai #largelanguagemodels #chatbot #opensourceUnderstanding AI Risk: Beyond the Surface in Enterprise!Unpacking the Complex World of Risk Management in AI – It’s Not What You Think!Unlocking the Secret to Impactful AI: Its All About Quality!Arcee AI live webinar - 18/09/2025Deep Dive: Optimizing LLM inferenceUncover the Truth Behind AI Model Bias - Its More Serious Than You Think!SLM in Action: Arcee Agent, A 7B model for function calls and tool usagePhi-2 on Intel Meteor Lake - Coding questionUnlock the Power of AI and Stock Market Data! 📈✨Arcee Orchestra - Build an Agentic Retrieval Workflow for the Energy Industry
Julien Simon |

Deep Dive: Quantizing Large Language Models, part 2

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER