Uploaded June 2024 | Updated September 2026, 10 hours ago
Abstract: We will discuss how vLLM combines continuous batching with speculative decoding with a focus on enabling external contributors. Topics include proposer/scorer/verifier framework, proposal methods, lookahead scheduling, dynamic speculative decoding, and future contribution ideas.
Speaker: Cade Daniel
Slides: docs.google.com/presentation/d/1p1xE-EbSAnXpTSiSI0gmy_wdwxN5XaULO3AnCWWoRe4/edit
Abstract: We will discuss how vLLM combines continuous batching with speculative decoding with a focus on enabling external contributors. Topics include proposer/scorer/verifier framework, proposal methods, lookahead scheduling, dynamic speculative decoding, and future contribution ideas.
Speaker: Cade Daniel
Slides: docs.google.com/presentation/d/1p1xE-EbSAnXpTSiSI0gmy_wdwxN5XaULO3AnCWWoRe4/edit








![Lecture 75 [ScaleML Series] GPU Programming Fundamentals + ThunderKittens
Speakers: William Brandon (Anthropic) and Simran Arora (ThunderKittens)
Full Schedule: https://scale-ml.org/bootcamp/
The GPU MODE x Scale ML speaker series is a 5-day, online event hosted on the GPU MODE YouTube channel where top researchers in AI will talk about various architectural and system-level advances that are integrated into OpenAI’s frontier open-source model, GPT-OSS.
Each day will consist of ~2 hours of talks and discussions (around noon PST, may start at slightly different times each day so please check frequently), covering a different component of the evolving transformer stack—from quirks in the attention mechanism and positional encodings to quantization, MoEs, and custom GPU kernels. Lecture 75 [ScaleML Series] GPU Programming Fundamentals + ThunderKittens](https://i.ytimg.com/vi/Cl2B_hmg4gA/mqdefault.jpg)

![Lecture 72: [ScaleML Series] Efficient & Effective Long-Context Modeling for Large Language Models
Speaker: Guangxuan Xiao (PhD at MIT)
Full Schedule: https://scale-ml.org/bootcamp/
The GPU MODE x Scale ML speaker series is a 5-day, online event hosted on the GPU MODE YouTube channel where top researchers in AI will talk about various architectural and system-level advances that are integrated into OpenAI’s frontier open-source model, GPT-OSS.
Each day will consist of ~2 hours of talks and discussions, covering a different component of the evolving transformer stack—from quirks in the attention mechanism and positional encodings to quantization, MoEs, and custom GPU kernels. Lecture 72: [ScaleML Series] Efficient & Effective Long-Context Modeling for Large Language Models](https://i.ytimg.com/vi/DFcKFDt0QEg/mqdefault.jpg)