Lecture 112: Production Megakernels for Real-World Inference @GPUMODE
Lecture 112: Production Megakernels for Real-World Inference  @GPUMODE
Uploaded August 2026 | Updated September 2026, 3 hours ago
Joe Fioti explains how the Luminal compiler brings megakernels into production inference alongside traditional kernel graphs, including persistent work queues, symbolic dependency tracking, sub-graph megakernels, and compiler search that chooses the best execution strategy.
Lecture 112: Production Megakernels for Real-World InferenceLecture 18: Fusing KernelsNeighborhood AttentionSmall RL Models at the Speed of Light with LeanRLLecture 53: torch.compile Q&ALecture 11: SparsityLecture 109: TIRxLecture 29: Triton InternalsLivestream: Mosaic GPULecture 104: Gluon and Linear LayoutsLecture 83: Formalized Kernel DerivationLecture 88: TinyTPU
GPU MODE |

Lecture 112: Production Megakernels for Real-World Inference

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER