Lecture 74: [ScaleML Series] Positional Encodings and PaTH Attention @GPUMODE
Lecture 74: [ScaleML Series] Positional Encodings and PaTH Attention  @GPUMODE
Uploaded August 2025 | Updated September 2026, 1 hour ago
Speaker: Songlin Yang (PhD at MIT).

Full Schedule: scale-ml.org/bootcamp

The GPU MODE x Scale ML speaker series is a 5-day, online event hosted on the GPU MODE YouTube channel where top researchers in AI will talk about various architectural and system-level advances that are integrated into OpenAI’s frontier open-source model, GPT-OSS.

Each day will consist of ~2 hours of talks and discussions (around noon PST, may start at slightly different times each day so please check frequently), covering a different component of the evolving transformer stack—from quirks in the attention mechanism and positional encodings to quantization, MoEs, and custom GPU kernels.
Lecture 74: [ScaleML Series] Positional Encodings and PaTH AttentionDistributed ML on consumer devicesLecture 60: Optimizing Linear AttentionLecture 8: CUDA Performance ChecklistLecture 16: On Hands ProfilingLecture 17: NCCLLecture 96: TLXLecture 113: Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU CollectivesProduction Megakernels for Real-World InferenceLive Quartet 4 bit trainingLecture 24: Scan at the Speed of LightLecture 80: How FlashAttention 4 Works
GPU MODE |

Lecture 74: [ScaleML Series] Positional Encodings and PaTH Attention

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER