The matrix math behind transformer neural networks, one step at a time!!! @statquest
The matrix math behind transformer neural networks, one step at a time!!!  @statquest
Uploaded April 2024 | Updated September 2026, 1 week ago
Transformers, the neural network architecture behind ChatGPT, do a lot of math. However, this math can be done quickly using matrix math because GPUs are optimized for it. Matrix math is also used when we code neural networks, so learning how ChatGPT does it will help you code your own. Thus, in this video, we go through the math one step at a time and explain what each step does so that you can use it on your own with confidence.

NOTE: This StatQuest assumes that you are already familiar with:
Transformers: youtu.be/zxQyTK8quyY
The essential matrix algebra for neural networks: youtu.be/bQ5BoolX9Ag

For a complete index of all the StatQuest videos, check out:
statquest.org/video-index

If you'd like to support StatQuest, please consider...

Patreon: patreon.com/statquest
...or...
YouTube Membership: youtube.com/channel/UCtYLUTtgS3k1Fg4y5tAhLbw/join

...buying one of my books, a study guide, a t-shirt or hoodie, or a song from the StatQuest store...
statquest.org/statquest-store

...or just donating to StatQuest!
paypal.me/statquest

Lastly, if you want to keep up with me as I research and create new StatQuests, follow me on twitter:
twitter.com/joshuastarmer

0:00 Awesome song and introduction
1:43 Word Embedding
3:37 Position Encoding
4:28 Self Attention
12:09 Residual Connections
13:08 Decoder Word Embedding and Position Encoding
15:33 Masked Self Attention
20:18 Encoder-Decoder Attention
21:31 Fully Connected Layer
22:16 SoftMax

#StatQuest #Transformer #ChatGPT
The matrix math behind transformer neural networks, one step at a time!!!Tensors for Neural Networks, Clearly Explained!!!False Discovery Rates, FDR, clearly explainedSequence-to-Sequence (seq2seq) Encoder-Decoder Neural Networks, Clearly Explained!!!Matrix NotationSilly Songs, Clearly Explained!!!Design Matrices For Linear Models, Clearly Explained!!!AdaBoost, Clearly ExplainedThe SoftMax Derivative, Step-by-Step!!!Using Bootstrapping to Calculate p-values!!!Regularization Part 2: Lasso (L1) RegressionNaive Bayes, Clearly Explained!!!
StatQuest with Josh Starmer |

The matrix math behind transformer neural networks, one step at a time!!!

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER