ALiBi | Train Short, Test Long: Attention With Linear Biases Enables Input Length Extrapolation @TheAIEpiphany
ALiBi | Train Short, Test Long: Attention With Linear Biases Enables Input Length Extrapolation  @TheAIEpiphany
Uploaded August 2021 | Updated September 2026, 1 week ago
πŸ‘¨β€πŸ‘©β€πŸ‘§β€πŸ‘¦ JOIN OUR DISCORD COMMUNITY:
Discord β–Ί discord.gg/peBrCpheKE

πŸ“’ SUBSCRIBE TO MY MONTHLY AI NEWSLETTER:
Substack β–Ί aiepiphany.substack.com

❀️ Become The AI Epiphany Patreon ❀️ β–Ί patreon.com/theaiepiphany

In this video I cover ALiBi model from the "Train Short, Test Long: Attention With Linear Biases Enables Input Length Extrapolation" paper.

Instead of using positional embeddings (like e.g. the original transformer) they added non-learnable biases directly into the query-key matrix and achieved extraordinary extrapolation results (i.e. great perf on longer sequences).

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬
βœ… Paper: ofir.io/train_short_test_long.pdf
βœ… Code: github.com/ofirpress/attention_with_linear_biases
β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬

⌚️ Timetable:
00:00 Intro
02:00 Main results
03:30 Time and memory tradeoffs
05:00 ALiBi method explained
10:00 Results, low data regime
12:35 Generalization to different datasets
13:15 Results, big data regime
16:00 Why does it work and future work
21:10 Outro

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬
πŸ’° BECOME A PATREON OF THE AI EPIPHANY ❀️

If these videos, GitHub projects, and blogs help you,
consider helping me out by supporting me on Patreon!

The AI Epiphany β–Ί patreon.com/theaiepiphany
One-time donation:
paypal.com/paypalme/theaiepiphany

Much love! ❀️

Huge thank you to these AI Epiphany patreons:
Eli Mahler
Petar VeličkoviΔ‡
Zvonimir Sabljic

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬

πŸ’‘ The AI Epiphany is a channel dedicated to simplifying the field of AI using creative visualizations and in general, a stronger focus on geometrical and visual intuition, rather than the algebraic and numerical "intuition".

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬
πŸ‘‹ CONNECT WITH ME ON SOCIAL
LinkedIn β–Ί linkedin.com/in/aleksagordic
Twitter β–Ί twitter.com/gordic_aleksa

Instagram β–Ί instagram.com/aiepiphany
Facebook β–Ί facebook.com/aiepiphany

πŸ‘¨β€πŸ‘©β€πŸ‘§β€πŸ‘¦ JOIN OUR DISCORD COMMUNITY:
Discord β–Ί discord.gg/peBrCpheKE

πŸ“’ SUBSCRIBE TO MY MONTHLY AI NEWSLETTER:
Substack β–Ί aiepiphany.substack.com

πŸ’» FOLLOW ME ON GITHUB FOR ML PROJECTS:
GitHub β–Ί github.com/gordicaleksa

πŸ“š FOLLOW ME ON MEDIUM:
Medium β–Ί gordicaleksa.medium.com
β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬

#alibi #transformers #extrapolation
ALiBi | Train Short, Test Long: Attention With Linear Biases Enables Input Length ExtrapolationDay 23: Open NLLB - day before China trip! Refactoring & pushing  changes (Pt 1 cont.)Hyperbolic Graph Convolutional Networks | Geometric ML Paper ExplainedDiffusion Models Beat GANs on Image Synthesis | ML Coding Series | Part 2Facebook AIs DINO | PyTorch Code ExplainedUltimate Guide To Scaling ML Models - Megatron-LM | ZeRO | DeepSpeed | Mixed PrecisionT0: Multitask Prompted Training Enables Zero-Shot Task Generalization | Paper ExplainedDay 17: Open NLLB - analyzing batch iterators (Pt 2)State Space Models w/ Albert Gu & Karan Goel (Cartesia AI)ConvNeXt: A ConvNet for the 2020s | Paper ExplainedVQ-GAN: Taming Transformers for High-Resolution Image Synthesis | Paper ExplainedDay 14: Open NLLB - exploring BLEU, chrF++, logging (Pt 3. cont.)
Aleksa Gordić - The AI Epiphany |

ALiBi | Train Short, Test Long: Attention With Linear Biases Enables Input Length Extrapolation

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER