Uploaded February 2026 | Updated September 2026, 11 hours ago
Summary:
TLX provides a Triton-like programming model that removes much of the mechanical complexity required to reach peak GPU performance, while preserving full freedom for heuristic-driven, performance-critical decisions.
The talk will focus on concrete kernel designs on Blackwell GPUs, showing how simplicity enables aggressive fusion, explicit scheduling, and predictable performance through low-level optimization techniques such as warp specialization, async pipelines, and memory system control.
Summary:
TLX provides a Triton-like programming model that removes much of the mechanical complexity required to reach peak GPU performance, while preserving full freedom for heuristic-driven, performance-critical decisions.
The talk will focus on concrete kernel designs on Blackwell GPUs, showing how simplicity enables aggressive fusion, explicit scheduling, and predictable performance through low-level optimization techniques such as warp specialization, async pipelines, and memory system control.










