Lecture 90: Building resilient ML Engineering skills @GPUMODE
Lecture 90: Building resilient ML Engineering skills  @GPUMODE
Uploaded January 2026 | Updated September 2026, 7 hours ago
Speaker: Stas Bekman

Abstract: With the aid of the ML Engineering open book we will explore foundational ML engineering skills that will hopefully help you to remain an expert in the environment that changes too fast. We will focus on things that don't change too much while discussing compute, networking and storage performance, and debugging and troubleshooting of the ML workloads.
Lecture 90: Building resilient ML Engineering skillsLecture 76: BackendBench fixing the LLM kernel correctness problemLecture 43: int8 tensorcore matmul for TuringLecture 64: Multi-GPU programmingThe History of CUDA MODE (Now GPU MODE)Lecture 30: Quantized TrainingLivestream int8 tensorcore matmul for TuringLivestream The Ultra Scale PlaybookLecture 75 [ScaleML Series] GPU Programming Fundamentals + ThunderKittensLecture 56: Kernel Benchmarking TalesLecture 72: [ScaleML Series] Efficient & Effective Long-Context Modeling for Large Language ModelsBonus Lecture: AMD Developer Challenge
GPU MODE |

Lecture 90: Building resilient ML Engineering skills

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER