[Live] ScaleML Series Day 3 — Quantization in Large Models @GPUMODE
[Live] ScaleML Series Day 3 — Quantization in Large Models  @GPUMODE
Uploaded August 2025 | Updated September 2026, 5 hours ago
Day 3: Quantization in Large Models by Chris De Sa.

Full Schedule: scale-ml.org/bootcamp

The GPU MODE x Scale ML speaker series is a 5-day, online event hosted on the GPU MODE YouTube channel where top researchers in AI will talk about various architectural and system-level advances that are integrated into OpenAI’s frontier open-source model, GPT-OSS.

Each day will consist of ~2 hours of talks and discussions (around noon PST, may start at slightly different times each day so please check frequently), covering a different component of the evolving transformer stack—from quirks in the attention mechanism and positional encodings to quantization, MoEs, and custom GPU kernels.
[Live] ScaleML Series Day 3 — Quantization in Large ModelsLecture 100: InferenceX Continuous OSS Inference BenchmarkingLecture 88: TinyTPUcuDNN[Live] ScaleML Series Day 4 — Positional Encodings and PaTH AttentionLecture 4 Compute and Memory BasicsLecture 112: Production Megakernels for Real-World InferenceLecture 18: Fusing KernelsNeighborhood AttentionSmall RL Models at the Speed of Light with LeanRLLecture 53: torch.compile Q&ALecture 11: Sparsity
GPU MODE |

[Live] ScaleML Series Day 3 — Quantization in Large Models

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER