Lecture 96: TLX @GPUMODE
Lecture 96: TLX  @GPUMODE
Uploaded February 2026 | Updated September 2026, 11 hours ago
Summary:

TLX provides a Triton-like programming model that removes much of the mechanical complexity required to reach peak GPU performance, while preserving full freedom for heuristic-driven, performance-critical decisions.

The talk will focus on concrete kernel designs on Blackwell GPUs, showing how simplicity enables aggressive fusion, explicit scheduling, and predictable performance through low-level optimization techniques such as warp specialization, async pipelines, and memory system control.
Lecture 96: TLXLecture 113: Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU CollectivesProduction Megakernels for Real-World InferenceLive Quartet 4 bit trainingLecture 24: Scan at the Speed of LightLecture 80: How FlashAttention 4 WorksLecture 39: TorchtitanCornserve: Easy, Fast and Scalable Multimodal AIBonus Lecture: CUDA C++ llm.cppLecture 35: SGLangLecture 69: Quartet 4 bit trainingTIRx
GPU MODE |

Lecture 96: TLX

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER