Lecture 79 Mirage (MPK): Compiling LLMs into Mega Kernels @GPUMODE
Lecture 79 Mirage (MPK): Compiling LLMs into Mega Kernels  @GPUMODE
Uploaded September 2025 | Updated September 2026, 9 hours ago
Talk by Mengdi Wu and Xinhao Cheng on Mirage. Mirage Persistent Kernel (MPK) is a compiler and runtime system that automatically transforms LLM inference into a single megakernel—a fused GPU kernel that performs all necessary computation and communication within a single kernel launch. This end-to-end GPU fusion approach reduces LLM inference latency by 1.2× to 6.7×, all while requiring minimal developer effort.

Repo: github.com/mirage-project/mirage
Lecture 79 Mirage (MPK): Compiling LLMs into Mega KernelsLecture 107: PithTrainLecture 58: Disaggregated LLM InferenceLecture 59: FastVideoLecture 93: Cornserve Easy, Fast and Scalable Multimodal AILecture 84: Numerics and AILive - Disaggregated LLM Inference: Past, Present and FutureCuTeLecture 108: One Layer Deeper competitionLivestream: Distributed GEMMLive: Scaling Laws for Low PrecisionLecture 57: CuTe
GPU MODE |

Lecture 79 Mirage (MPK): Compiling LLMs into Mega Kernels

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER