Lecture 36: CUTLASS and Flash Attention 3 @GPUMODE
Lecture 36: CUTLASS and Flash Attention 3  @GPUMODE
Uploaded November 2024 | Updated September 2026, 38 minutes ago
Speaker: Jay Shah
Slides: github.com/cuda-mode/lectures

Correction by Jay: "It turns out I inserted the wrong image for the intra-warpgroup overlapping (this was an older overlapping scheme we don't use now), while the algorithm was correct; I corrected this in the attached version."
Lecture 36: CUTLASS and Flash Attention 3Lecture 71: [ScaleML Series] FlexOlmo: Open Language Models for Flexible Data Use[Live] ScaleML Series Day 5 — GPU Programming for Foundation ModelsFactorio Learning EnvironmentLecture 27: gpu.cpp - Portable GPU compute using WebGPULecture 1 How to profile CUDA kernels in PyTorchLecture 82 Helion: A high-level DSL for ML kernelsLecture 21: Scan Algorithm Part 2Lecture 103: Fundamentals of CuTe Layout Algebra and Category-theoretic InterpretationConsumer GPU performanceLecture 46: Distributed GEMMLecture 2 Ch1-3 PMPP book
GPU MODE |

Lecture 36: CUTLASS and Flash Attention 3

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER