Formalized Deep Learning Architectures for Automated Low-Level Kernel Optimization @GPUMODE
Formalized Deep Learning Architectures for Automated Low-Level Kernel Optimization  @GPUMODE
Uploaded October 2025 | Updated September 2026, 1 hour ago
Abstract: Vincent Abbott is a PhD student at the Massachusetts Institute of Technology's (MIT) Zardini Lab who has developed a formal framework for describing the relationship between the mathematical function implemented by a deep learning model, its resource usage, and low-level implementation. These methods are based on category theoretic diagrams [1]. The Zardini Lab has developed these diagrams into a tool for rapidly deriving low-level algorithms, as presented in their recent work FlashAttention on a Napkin [2]. These methods have been put into practice, deriving a FlashAttention-like algorithm for an attention variant from first principles [3].

Recently, he has been working on encoding the underlying mathematics into an automated tool for diagram generation and algorithm optimization. In this talk, Vincent Abbott will cover formal diagrams for deep learning models, show how they can be used to derive low-level algorithms such as FlashAttention and corresponding performance models, and preview work related to automated tools for diagramming and analyzing algorithms.

[1] openreview.net/forum?id=RyZB4qXEgt
[2] openreview.net/forum?id=pF2ukh7HxA
[3] dl.acm.org/doi/10.1007/978-3-032-00686-8_1
Formalized Deep Learning Architectures for Automated Low-Level Kernel OptimizationLecture 6 Optimizing OptimizersLecture 23: Tensor CoresMonarch applied to async RLGame ArenaLecture 32: UnslothcuTileLecture 78 Iris: Multi-GPU Programming in TritonLecture 33: BitblasLecture 41: FlashInferLecture 85: Factorio Learning EnvironmentOptimizing Linear Attention in Triton
GPU MODE |

Formalized Deep Learning Architectures for Automated Low-Level Kernel Optimization

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER