TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters (Paper Explained) @YannicKilcher
TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters (Paper Explained)  @YannicKilcher
Uploaded November 2024 | Updated September 2026, 2 weeks ago
A deep dive into the TokenFormer and an opinion about its impact, novelty, and relation to prior work.

Paper: arxiv.org/abs/2410.23168

Abstract:
Transformers have become the predominant architecture in foundation models due to their excellent performance across various domains. However, the substantial cost of scaling these models remains a significant concern. This problem arises primarily from their dependence on a fixed number of parameters within linear projections. When architectural modifications (e.g., channel dimensions) are introduced, the entire model typically requires retraining from scratch. As model sizes continue growing, this strategy results in increasingly high computational costs and becomes unsustainable. To overcome this problem, we introduce TokenFormer, a natively scalable architecture that leverages the attention mechanism not only for computations among input tokens but also for interactions between tokens and model parameters, thereby enhancing architectural flexibility. By treating model parameters as tokens, we replace all the linear projections in Transformers with our token-parameter attention layer, where input tokens act as queries and model parameters as keys and values. This reformulation allows for progressive and efficient scaling without necessitating retraining from scratch. Our model scales from 124M to 1.4B parameters by incrementally adding new key-value parameter pairs, achieving performance comparable to Transformers trained from scratch while greatly reducing training costs. Code and models are available at \url{this https URL}.

Authors: Haiyang Wang, Yue Fan, Muhammad Ferjad Naeem, Yongqin Xian, Jan Eric Lenssen, Liwei Wang, Federico Tombari, Bernt Schiele

Links:
Homepage: ykilcher.com
Merch: ykilcher.com/merch
YouTube: youtube.com/c/yannickilcher
Twitter: twitter.com/ykilcher
Discord: ykilcher.com/discord
LinkedIn: linkedin.com/in/ykilcher

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: subscribestar.com/yannickilcher
Patreon: patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters (Paper Explained)OpenAssistant is CompletedGLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsLLaMA Pro: Progressive LLaMA with Block Expansion (Paper Explained)Sparse is Enough in Scaling Transformers (aka Terraformer) | ML Research Paper ExplainedAGI is not coming!Context Rot: How Increasing Input Tokens Impacts LLM Performance (Paper Analysis)VOS: Learning What You Dont Know by Virtual Outlier Synthesis (Paper Explained)Is Stability turning into OpenAI?Were RNNs All We Needed? (Paper Explained)JEPA - A Path Towards Autonomous Machine Intelligence (Paper Explained)First Author Interview: AI & formal math (Formal Mathematics Statement Curriculum Learning)
Yannic Kilcher |

TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters (Paper Explained)

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER