Lecture 58: Disaggregated LLM Inference @GPUMODE
Lecture 58: Disaggregated LLM Inference  @GPUMODE
Uploaded May 2025 | Updated September 2026, 2 hours ago
Speaker: Junda Chen
Lecture 58: Disaggregated LLM InferenceLecture 59: FastVideoLecture 93: Cornserve Easy, Fast and Scalable Multimodal AILecture 84: Numerics and AILive - Disaggregated LLM Inference: Past, Present and FutureCuTeLecture 108: One Layer Deeper competitionLivestream: Distributed GEMMLive: Scaling Laws for Low PrecisionLecture 57: CuTeLecture 42: Mosaic GPULecture 5: Going Further with CUDA for Python Programmers
GPU MODE |

Lecture 58: Disaggregated LLM Inference

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER