Lecture 113: Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives @GPUMODE
Lecture 113: Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives  @GPUMODE
Uploaded August 2026 | Updated September 2026, 2 days ago
arxiv.org/abs/2607.16100
Lecture 113: Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU CollectivesProduction Megakernels for Real-World InferenceLive Quartet 4 bit trainingLecture 24: Scan at the Speed of LightLecture 80: How FlashAttention 4 WorksLecture 39: TorchtitanCornserve: Easy, Fast and Scalable Multimodal AIBonus Lecture: CUDA C++ llm.cppLecture 35: SGLangLecture 69: Quartet 4 bit trainingTIRxLecture 52: Scaling Laws for Low Precision
GPU MODE |

Lecture 113: Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER