Deep Dive: LLM Quantization, part 3 - FP8, FP4 @juliensimonfr
Deep Dive: LLM Quantization, part 3 - FP8, FP4  @juliensimonfr
Uploaded March 2026 | Updated September 2026, 2 weeks ago
Two years after parts 1 (youtu.be/kw7S-3s50uk) and 2 (youtu.be/fXBBwCIA0Ds), the quantization landscape has changed completely. FP8 is the new default; the serving stack has absorbed quantization, and MoE models have broken the old assumptions. This video shows you what practitioners actually use today.

⭐️⭐️⭐️ More content on Substack at airealist.ai ⭐️⭐️⭐️

In Part 3 of this series, I'm showing you exactly what to run, on which GPU, with which tool, and what breaks when you get it wrong. One model, start to finish: Arcee Trinity Mini (26B MoE, 3B active, 128 experts).
What's covered:
→ Why FP8 is the new BF16 — essentially lossless, one flag, works everywhere
→ The three things people call "4-bit" (INT4 vs MXFP4 vs NVFP4) and why they're completely different
→ Serving with vLLM: BF16, FP8 on-the-fly, FP8 checkpoint, with verified H100 numbers
→ Why MoE models break standard quantization and how to fix it
→ Creating your own checkpoints with llm-compressor and NVIDIA ModelOpt
→ NVFP4 on Blackwell: real 4-bit tensor core math, what works and what doesn't yet
→ The FP4 deep dive: E2M1 bit layout, MXFP4 vs NVFP4 scale factors, worked error example
→ Decision tree: which quantization path for your hardware and use case

*** Slides
fr.slideshare.net/slideshow/advanced-quantization-techniques-for-large-language-models-in-2026-a4c8/286754686

*** Arcee Trinity Mini:
26B total parameters, 3B active per token, 128 experts, 128K context, Apache 2.0.
huggingface.co/arcee-ai/Trinity-Mini
NVFP4: huggingface.co/arcee-ai/Trinity-Mini-NVFP4
FP8: huggingface.co/arcee-ai/Trinity-Mini-FP8-Block

#llm #quantization #fp8 #fp4 #nvfp4 #vllm #inference #nvidia #blackwell #moe #arcee
Deep Dive: LLM Quantization, part 3 - FP8, FP4Deep Dive: How Three MoE Reasoning Models Actually Work — Trinity, DeepSeek R1, Kimi K2 - Part 2Arcee.ai and Small Language Models - Mobile World Congress, Las Vegas (10/2024)LLMs from the trenches - LLMs are not intelligent, there is no reasoningUnlocking the Future of Finance with AI! 🧠💰Discover the Power of Arcee-Lite: Language Modeling Redefined!3 production-ready models released by Arcee AI on Hugging FaceDeploying Llama3 with Inference Endpoints and AWS Inferentia2The Ease of Running Small Language Models LocallyDeploy Hugging Face models on Google Cloud: from the hub to Vertex AI🚀 Unveiling the Future: Meet Arcee Nova, the Ultimate LLM! 🌟Open Source AI with Hugging Face - Dallas AI  meetup (05/2024)
Julien Simon |

Deep Dive: LLM Quantization, part 3 - FP8, FP4

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER