Building resilient ML Engineering skills @GPUMODE
Building resilient ML Engineering skills  @GPUMODE
Uploaded January 2026 | Updated September 2026, 44 minutes ago
Speaker: Stas Bekman

Abstract: With the aid of the ML Engineering open book we will explore foundational ML engineering skills that will hopefully help you to remain an expert in the environment that changes too fast. We will focus on things that don't change too much while discussing compute, networking and storage performance, and debugging and troubleshooting of the ML workloads.
Building resilient ML Engineering skillsLandscape of GPU Centric communicationGPU Kernel Formal VerificationEXO 2Lecture 113: Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU CollectivesMega Lecture 91: Reinforcement Learning, Agents & OpenEnvLecture 36: CUTLASS and Flash Attention 3Lecture 71: [ScaleML Series] FlexOlmo: Open Language Models for Flexible Data Use[Live] ScaleML Series Day 5 — GPU Programming for Foundation ModelsFactorio Learning EnvironmentLecture 27: gpu.cpp - Portable GPU compute using WebGPULecture 1 How to profile CUDA kernels in PyTorch
GPU MODE |

Building resilient ML Engineering skills

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER