Lecture 28: Liger Kernel - Efficient Triton Kernels for LLM Training @GPUMODE
Lecture 28: Liger Kernel - Efficient Triton Kernels for LLM Training  @GPUMODE
Uploaded September 2024 | Updated September 2026, 16 hours ago
Byron Hsu presents LinkedIn's open-source collection of Triton kernels for efficient LLM training.

TIMESTAMPS

00:00 Host Opening
00:22 Main Focus
01:18 Outline
03:03 LLM Training Bottleneck
05:27 Live Demo - PyTorch Profiler
10:41 Why Triton
12:53 QA
13:49 Example - RMS Norm
18:00 QA
20:03 RMS Norm Tricks
21:20 Live Code - RMS Nrom
25:40 QA
28:14 Example - Fused Linear Cross Entropy
30:58 Gradient Checkpointing
31:51 Gradient-in-forward
32:53 QA
35:01 Chunking
36:23 QA
37:56 Live Code - Fused Linear Cross Entropy
39:59 QA
41:15 Convergence Test
42:39 Live Code - Convergence Test
44:12 Contiguity
45:15 Live Code - Contiguity
48:01 QA
49:45 Memory Address
50:38 Live Code - Memory Address
52:11 QA
1:00:28 QA - Liger Kernel
1:09:12 Acknowledgement

Slides: docs.google.com/presentation/d/1CGTV-uKw9crrBo13q1jAzAFCFzlpZFjeL4bnK67pTd8/edit?usp=sharing
Notebooks: github.com/cuda-mode/lectures/blob/main/README.md#lecture-28-liger-kernel
Lecture 28: Liger Kernel - Efficient Triton Kernels for LLM TrainingFormalized Deep Learning Architectures for Automated Low-Level Kernel OptimizationLecture 6 Optimizing OptimizersLecture 23: Tensor CoresMonarch applied to async RLGame ArenaLecture 32: UnslothcuTileLecture 78 Iris: Multi-GPU Programming in TritonLecture 33: BitblasLecture 41: FlashInferLecture 85: Factorio Learning Environment
GPU MODE |

Lecture 28: Liger Kernel - Efficient Triton Kernels for LLM Training

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER