Production Megakernels for Real-World Inference @GPUMODE
Production Megakernels for Real-World Inference  @GPUMODE
Uploaded August 2026 | Updated September 2026, 3 hours ago
Joe Fioti explains how the Luminal compiler brings megakernels into production inference alongside traditional kernel graphs, including persistent work queues, symbolic dependency tracking, sub-graph megakernels, and compiler search that chooses the best execution strategy.
Production Megakernels for Real-World InferenceLive Quartet 4 bit trainingLecture 24: Scan at the Speed of LightLecture 80: How FlashAttention 4 WorksLecture 39: TorchtitanCornserve: Easy, Fast and Scalable Multimodal AIBonus Lecture: CUDA C++ llm.cppLecture 35: SGLangLecture 69: Quartet 4 bit trainingTIRxLecture 52: Scaling Laws for Low PrecisionLecture 66: Game Arena
GPU MODE |

Production Megakernels for Real-World Inference

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER