Mirage (MPK): Compiling LLMs into Mega Kernels @GPUMODE
Mirage (MPK): Compiling LLMs into Mega Kernels  @GPUMODE
Uploaded September 2025 | Updated September 2026, 45 minutes ago
Talk by Mengdi Wu and Xinhao Cheng on Mirage. Mirage Persistent Kernel (MPK) is a compiler and runtime system that automatically transforms LLM inference into a single megakernel—a fused GPU kernel that performs all necessary computation and communication within a single kernel launch. This end-to-end GPU fusion approach reduces LLM inference latency by 1.2× to 6.7×, all while requiring minimal developer effort.

Repo: github.com/mirage-project/mirage
Mirage (MPK): Compiling LLMs into Mega KernelsLecture 114: PyCuTeLearning CUTLASS the hard wayLecture 63: Search-Based Deep Learning CompilersLecture 89: cuTile (from friends at NVIDIA)PyCuTeSpectral Compute: Compile CUDA everywhereLecture 68: Landscape of GPU Centric communicationMulti-GPU programmingThe 4-bitter lesson: Balancing Stability and Performance in NVFP4 RLLecture 31: Beginners Guide to MetalLecture 81: High-performance purely functional data-parallel array programming
GPU MODE |

Mirage (MPK): Compiling LLMs into Mega Kernels

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER