Uploaded November 2024 | Updated September 2026, 38 minutes ago
Speaker: Jay Shah
Slides: github.com/cuda-mode/lectures
Correction by Jay: "It turns out I inserted the wrong image for the intra-warpgroup overlapping (this was an older overlapping scheme we don't use now), while the algorithm was correct; I corrected this in the attached version."
Speaker: Jay Shah
Slides: github.com/cuda-mode/lectures
Correction by Jay: "It turns out I inserted the wrong image for the intra-warpgroup overlapping (this was an older overlapping scheme we don't use now), while the algorithm was correct; I corrected this in the attached version."
![Lecture 71: [ScaleML Series] FlexOlmo: Open Language Models for Flexible Data Use
Day 1: Overview of the series and a long talk on FlexOlmo by Professor Sewon Min.
Full Schedule: https://scale-ml.org/bootcamp/
The GPU MODE x Scale ML speaker series is a 5-day, online event hosted on the GPU MODE YouTube channel where top researchers in AI will talk about various architectural and system-level advances that are integrated into OpenAI’s frontier open-source model, GPT-OSS.
Each day will consist of ~2 hours of talks and discussions (around noon PST, may start at slightly different times each day so please check frequently), covering a different component of the evolving transformer stack—from quirks in the attention mechanism and positional encodings to quantization, MoEs, and custom GPU kernels. Lecture 71: [ScaleML Series] FlexOlmo: Open Language Models for Flexible Data Use](https://i.ytimg.com/vi/KorF7Xpozhg/mqdefault.jpg)
![[Live] ScaleML Series Day 5 — GPU Programming for Foundation Models
Day 5: GPU Programming for Foundation Models by William Brandon & Simran Arora.
Full Schedule: https://scale-ml.org/bootcamp/
The GPU MODE x Scale ML speaker series is a 5-day, online event hosted on the GPU MODE YouTube channel where top researchers in AI will talk about various architectural and system-level advances that are integrated into OpenAI’s frontier open-source model, GPT-OSS.
Each day will consist of ~2 hours of talks and discussions (around noon PST, may start at slightly different times each day so please check frequently), covering a different component of the evolving transformer stack—from quirks in the attention mechanism and positional encodings to quantization, MoEs, and custom GPU kernels. [Live] ScaleML Series Day 5 — GPU Programming for Foundation Models](https://i.ytimg.com/vi/LMk8nqIFXLo/mqdefault.jpg)








