Titans: Learning to Memorize at Test Time (Paper Analysis) @YannicKilcher
Titans: Learning to Memorize at Test Time (Paper Analysis)  @YannicKilcher
Uploaded December 2025 | Updated September 2026, 2 weeks ago
Paper: arxiv.org/abs/2501.00663

Abstract:
Over more than a decade there has been an extensive research effort on how to effectively utilize recurrent models and attention. While recurrent models aim to compress the data into a fixed-size memory (called hidden state), attention allows attending to the entire context window, capturing the direct dependencies of all tokens. This more accurate modeling of dependencies, however, comes with a quadratic cost, limiting the model to a fixed-length context. We present a new neural long-term memory module that learns to memorize historical context and helps attention to attend to the current context while utilizing long past information. We show that this neural memory has the advantage of fast parallelizable training while maintaining a fast inference. From a memory perspective, we argue that attention due to its limited context but accurate dependency modeling performs as a short-term memory, while neural memory due to its ability to memorize the data, acts as a long-term, more persistent, memory. Based on these two modules, we introduce a new family of architectures, called Titans, and present three variants to address how one can effectively incorporate memory into this architecture. Our experimental results on language modeling, common-sense reasoning, genomics, and time series tasks show that Titans are more effective than Transformers and recent modern linear recurrent models. They further can effectively scale to larger than 2M context window size with higher accuracy in needle-in-haystack tasks compared to baselines.

Authors: Ali Behrouz, Peilin Zhong, Vahab Mirrokni

Links:
Homepage: ykilcher.com
Merch: ykilcher.com/merch
YouTube: youtube.com/c/yannickilcher
Twitter: twitter.com/ykilcher
Discord: ykilcher.com/discord
LinkedIn: linkedin.com/in/ykilcher

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: subscribestar.com/yannickilcher
Patreon: patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Titans: Learning to Memorize at Test Time (Paper Analysis)I created an AI-powered Social Network[ML News] GPT-3 learns to edit | Google Pathways | Make-A-Scene | CLIP meets GamePhysics | DouBlindLearning Rate Grafting: Transferability of Optimizer Tuning (Machine Learning Research Paper Review)Unsupervised Brain Models - How does Deep Learning inform Neuroscience? (w/ Patrick Mineault)Dynamic Inference with Neural Interpreters (w/ author interview)RWKV: Reinventing RNNs for the Transformer Era (Paper Explained)I BUILT A FULLY AUTOMATIC MANSPLAINER[ML News] Stable Diffusion Takes Over! (Open Source AI Art)[ML News] This AI completes Wikipedia! Meta AI Sphere | Google Minerva | GPT-3 writes a paper[ML News] LLaMA2 Released | LLMs for Robots | Multimodality on the RiseThis ChatGPT Skill will earn you $10B (also, AI reads your mind!) | ML News
Yannic Kilcher |

Titans: Learning to Memorize at Test Time (Paper Analysis)

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER