Lecture 87: Low Latency Communication Kernels with NVSHMEM @GPUMODE
Lecture 87: Low Latency Communication Kernels with NVSHMEM  @GPUMODE
Uploaded December 2025 | Updated September 2026, 40 minutes ago
Speaker: Prajwal Singhania

High-performance inference at scale is increasingly bottlenecked by communication, especially in decode-heavy LLM workloads where tensor parallelism dominates.

In this talk, we will introduce NVRAR - an NVSHMEM-based all-reduce tailored for inter-node settings.
Lecture 87: Low Latency Communication Kernels with NVSHMEMOutperforming cuBLAS on NVFP4Lecture 99: Distributed ML on consumer devicesLecture 79 Mirage (MPK): Compiling LLMs into Mega KernelsLecture 107: PithTrainLecture 58: Disaggregated LLM InferenceLecture 59: FastVideoLecture 93: Cornserve Easy, Fast and Scalable Multimodal AILecture 84: Numerics and AILive - Disaggregated LLM Inference: Past, Present and FutureCuTeLecture 108: One Layer Deeper competition
GPU MODE |

Lecture 87: Low Latency Communication Kernels with NVSHMEM

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER