Lecture 80: How FlashAttention 4 Works @GPUMODE
Lecture 80: How FlashAttention 4 Works  @GPUMODE
Uploaded October 2025 | Updated September 2026, 2 hours ago
Speaker: Charles Frye

The source code (in CuTe) for FlashAttention4 on Blackwell GPUs has recently been released for the forward pass. The following blog: modal.com/blog/reverse-engineer-flash-attention-4 goes over their findings when reading through the source code, and changes between FA1,2,3 and now 4!
Lecture 80: How FlashAttention 4 WorksLecture 39: TorchtitanCornserve: Easy, Fast and Scalable Multimodal AIBonus Lecture: CUDA C++ llm.cppLecture 35: SGLangLecture 69: Quartet 4 bit trainingTIRxLecture 52: Scaling Laws for Low PrecisionLecture 66: Game ArenaProving Kernels Correct Instead of Testing ThemHow FlashAttention 4 WorksLecture 20: Scan Algorithm
GPU MODE |

Lecture 80: How FlashAttention 4 Works

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER