Uploaded February 2026 | Updated September 2026, 12 hours ago
Summary:
TLX provides a Triton-like programming model that removes much of the mechanical complexity required to reach peak GPU performance, while preserving full freedom for heuristic-driven, performance-critical decisions.
The talk will focus on concrete kernel designs on Blackwell GPUs, showing how simplicity enables aggressive fusion, explicit scheduling, and predictable performance through low-level optimization techniques such as warp specialization, async pipelines, and memory system control.
Summary:
TLX provides a Triton-like programming model that removes much of the mechanical complexity required to reach peak GPU performance, while preserving full freedom for heuristic-driven, performance-critical decisions.
The talk will focus on concrete kernel designs on Blackwell GPUs, showing how simplicity enables aggressive fusion, explicit scheduling, and predictable performance through low-level optimization techniques such as warp specialization, async pipelines, and memory system control.
![[Live] ScaleML Series Day 3 — Quantization in Large Models
Day 3: Quantization in Large Models by Chris De Sa.
Full Schedule: https://scale-ml.org/bootcamp/
The GPU MODE x Scale ML speaker series is a 5-day, online event hosted on the GPU MODE YouTube channel where top researchers in AI will talk about various architectural and system-level advances that are integrated into OpenAI’s frontier open-source model, GPT-OSS.
Each day will consist of ~2 hours of talks and discussions (around noon PST, may start at slightly different times each day so please check frequently), covering a different component of the evolving transformer stack—from quirks in the attention mechanism and positional encodings to quantization, MoEs, and custom GPU kernels. [Live] ScaleML Series Day 3 — Quantization in Large Models](https://i.ytimg.com/vi/k8PcSGG249Y/mqdefault.jpg)



![[Live] ScaleML Series Day 4 — Positional Encodings and PaTH Attention
Day 4: Positional Encodings and PaTH Attention by Songlin Yang.
Full Schedule: https://scale-ml.org/bootcamp/
The GPU MODE x Scale ML speaker series is a 5-day, online event hosted on the GPU MODE YouTube channel where top researchers in AI will talk about various architectural and system-level advances that are integrated into OpenAI’s frontier open-source model, GPT-OSS.
Each day will consist of ~2 hours of talks and discussions (around noon PST, may start at slightly different times each day so please check frequently), covering a different component of the evolving transformer stack—from quirks in the attention mechanism and positional encodings to quantization, MoEs, and custom GPU kernels. [Live] ScaleML Series Day 4 — Positional Encodings and PaTH Attention](https://i.ytimg.com/vi/l6_fdwRvMPk/mqdefault.jpg)





