Lecture 72: [ScaleML Series] Efficient & Effective Long-Context Modeling for Large Language Models @GPUMODE
Lecture 72: [ScaleML Series] Efficient & Effective Long-Context Modeling for Large Language Models  @GPUMODE
Uploaded August 2025 | Updated September 2026, 13 hours ago
Speaker: Guangxuan Xiao (PhD at MIT)

Full Schedule: scale-ml.org/bootcamp

The GPU MODE x Scale ML speaker series is a 5-day, online event hosted on the GPU MODE YouTube channel where top researchers in AI will talk about various architectural and system-level advances that are integrated into OpenAI’s frontier open-source model, GPT-OSS.

Each day will consist of ~2 hours of talks and discussions, covering a different component of the evolving transformer stack—from quirks in the attention mechanism and positional encodings to quantization, MoEs, and custom GPU kernels.
Lecture 72: [ScaleML Series] Efficient & Effective Long-Context Modeling for Large Language ModelsBonus Lecture: AMD Developer ChallengeLecture 14: Practitioners Guide to TritonLecture 102: quartet v2Lecture 54: Small RL Models at the Speed of Light with LeanRLLecture 45: Outperforming cuBLAS on H100One Layer Deeper competitionScan at the speed of lightGPU MODE IRL 2024 KeynotesLecture 10: Build a Prod Ready CUDA libraryLecture 19: Data Processing on GPUsLecture 44: NVIDIA Profiling
GPU MODE |

Lecture 72: [ScaleML Series] Efficient & Effective Long-Context Modeling for Large Language Models

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER