Attention Approximates Sparse Distributed Memory @MITCBMM
Attention Approximates Sparse Distributed Memory  @MITCBMM
Uploaded October 2021 | Updated September 2026, 1 week ago
Trenton Bricken, Harvard University
Abstract: While Attention has come to be an important mechanism in deep learning, it emerged out of a heuristic process of trial and error, providing limited intuition for why it works so well. Here, we show that Transformer Attention closely approximates Sparse Distributed Memory (SDM), a biologically plausible associative memory model, under certain data conditions. We confirm that these conditions are satisfied in pre-trained GPT2 Transformer models. We discuss the implications of the Attention-SDM map and provide new computational and biological interpretations of Attention.
Attention Approximates Sparse Distributed MemoryRoadmap for the day¿Cómo ven nuestros ojos el cuadro completo?The Platonic Representation HypothesisCan we change someones emotions by showing them a picture?Selective responses to faces, scenes, and bodies in the ventral visual pathway of infantsWhat can computers tell us about biology?Sprouting: BienvenidaLearning to Reason, Insights from Language ModelingScaling Inference MissionQuest Engineering Team presentationParallel systems for social and spatial reasoning within the brains apex network
MITCBMM |

Attention Approximates Sparse Distributed Memory

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER