Lecture 22: Hackers Guide to Speculative Decoding in VLLM @GPUMODE
Lecture 22: Hackers Guide to Speculative Decoding in VLLM  @GPUMODE
Uploaded June 2024 | Updated September 2026, 10 hours ago
Abstract: We will discuss how vLLM combines continuous batching with speculative decoding with a focus on enabling external contributors. Topics include proposer/scorer/verifier framework, proposal methods, lookahead scheduling, dynamic speculative decoding, and future contribution ideas.

Speaker: Cade Daniel

Slides: docs.google.com/presentation/d/1p1xE-EbSAnXpTSiSI0gmy_wdwxN5XaULO3AnCWWoRe4/edit
Lecture 22: Hackers Guide to Speculative Decoding in VLLMLecture 90: Building resilient ML Engineering skillsLecture 76: BackendBench fixing the LLM kernel correctness problemLecture 43: int8 tensorcore matmul for TuringLecture 64: Multi-GPU programmingThe History of CUDA MODE (Now GPU MODE)Lecture 30: Quantized TrainingLivestream int8 tensorcore matmul for TuringLivestream The Ultra Scale PlaybookLecture 75 [ScaleML Series] GPU Programming Fundamentals + ThunderKittensLecture 56: Kernel Benchmarking TalesLecture 72: [ScaleML Series] Efficient & Effective Long-Context Modeling for Large Language Models
GPU MODE |

Lecture 22: Hacker's Guide to Speculative Decoding in VLLM

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER