Deep Dive: Optimizing LLM inference @juliensimonfr
Deep Dive: Optimizing LLM inference  @juliensimonfr
Uploaded March 2024 | Updated September 2026, 2 weeks ago
Open-source LLMs are great for conversational applications, but they can be difficult to scale in production and deliver latency and throughput that are incompatible with your cost-performance objectives.

In this video, we zoom in on optimizing LLM inference, and study key mechanisms that help reduce latency and increase throughput: the KV cache, continuous batching, and speculative decoding, including the state-of-the-art Medusa approach.

Slides: fr.slideshare.net/slideshow/julien-simon-deep-dive-optimizing-llm-inference-69d3/270921961

⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. Follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️

00:00 Introduction
01:15 Decoder-only inference
06:05 The KV cache
11:15 Continuous batching
16:17 Speculative decoding
25:28 Speculative decoding: small off-the-shelf model
26:40 Speculative decoding: n-grams
30:25 Speculative decoding: Medusa
Deep Dive: Optimizing LLM inferenceUncover the Truth Behind AI Model Bias - Its More Serious Than You Think!SLM in Action: Arcee Agent, A 7B model for function calls and tool usagePhi-2 on Intel Meteor Lake - Coding questionUnlock the Power of AI and Stock Market Data! 📈✨Arcee Orchestra - Build an Agentic Retrieval Workflow for the Energy IndustryDiscover the Power of Arcee Lite: The Best 1.5B Model!Deep Dive: Quantizing Large Language Models, part 1Building a Kanban Task Management Web App from Scratch with ReplitWhy Local Voices Matter in AI Development! 🌍🤖Azure ML: deploy Hugging Face models in minutes!Unlocking Innovation: Why Large Companies Cant Afford to Stall! 🚀
Julien Simon |

Deep Dive: Optimizing LLM inference

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER