Attention for Neural Networks, Clearly Explained!!! @statquest
Attention for Neural Networks, Clearly Explained!!!  @statquest
Uploaded June 2023 | Updated September 2026, 1 week ago
Attention is one of the most important concepts behind Transformers and Large Language Models, like ChatGPT. However, it's not that complicated. In this StatQuest, we add Attention to a basic Sequence-to-Sequence (Seq2Seq or Encoder-Decoder) model and walk through how it works and is calculated, one step at a time. BAM!!!

NOTE: This StatQuest is based on two manuscripts. 1) The manuscript that originally introduced Attention to Encoder-Decoder Models: Neural Machine Translation by Jointly Learning to Align and Translate: arxiv.org/abs/1409.0473 and 2) The manuscript that first used the Dot-Product similarity for Attention in a similar context: Effective Approaches to Attention-based Neural Machine Translation arxiv.org/abs/1508.04025

NOTE: This StatQuest assumes that you are already familiar with basic Encoder-Decoder neural networks. If not, check out the 'Quest: youtu.be/L8HKweZIOmg

For a complete index of all the StatQuest videos, check out:
statquest.org/video-index

If you'd like to support StatQuest, please consider...

Patreon: patreon.com/statquest
...or...
YouTube Membership: youtube.com/channel/UCtYLUTtgS3k1Fg4y5tAhLbw/join

...buying one of my books, a study guide, a t-shirt or hoodie, or a song from the StatQuest store...
statquest.org/statquest-store

...or just donating to StatQuest!
paypal.me/statquest

Lastly, if you want to keep up with me as I research and create new StatQuests, follow me on twitter:
twitter.com/joshuastarmer

0:00 Awesome song and introduction
3:14 The Main Idea of Attention
5:34 A worked out example of Attention
10:18 The Dot Product Similarity
11:52 Using similarity scores to calculate Attention values
13:27 Using Attention values to predict an output word
14:22 Summary of Attention

#StatQuest #neuralnetwork #attention
Attention for Neural Networks, Clearly Explained!!!The Sensitivity, Specificity, Precision, Recall Sing-a-Long!!!Regularization Part 1: Ridge (L2) RegressionSupport Vector Machines Part 3: The Radial (RBF) Kernel (Part 3 of 3)Human Stories in AI: Rick MarksWord Embedding in PyTorch + LightningHuman Stories in AI: Xavier MoyáUsing Linear Models for t tests and ANOVA, Clearly Explained!!!Clustering with DBSCAN, Clearly Explained!!!Long Short-Term Memory with PyTorch + LightningStatistical Power, Clearly Explained!!!Christmas Morning
StatQuest with Josh Starmer |

Attention for Neural Networks, Clearly Explained!!!

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER