Outperforming cuBLAS on NVFP4 @GPUMODE
Outperforming cuBLAS on NVFP4  @GPUMODE
Uploaded September 2026 | Updated September 2026, 3 hours ago
GPU MODE talk by Pranjal: Outperforming cuBLAS on NVFP4.

cudaforfun.substack.com/p/outperforming-cublas-on-nvfp4
Outperforming cuBLAS on NVFP4Lecture 99: Distributed ML on consumer devicesLecture 79 Mirage (MPK): Compiling LLMs into Mega KernelsLecture 107: PithTrainLecture 58: Disaggregated LLM InferenceLecture 59: FastVideoLecture 93: Cornserve Easy, Fast and Scalable Multimodal AILecture 84: Numerics and AILive - Disaggregated LLM Inference: Past, Present and FutureCuTeLecture 108: One Layer Deeper competitionLivestream: Distributed GEMM
GPU MODE |

Outperforming cuBLAS on NVFP4

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER