Uploaded October 2022 | Updated September 2026, 2 weeks ago
This video tutorial explains how Bi-Directional Attention works in NLP. This attention mechanism is very similar to Self-attention method which was introduced in the previous video. The main difference is that Bi-Directional attention do not have Masking operation, which Self-attention has.
Also, while self-attention look to the previous words (or tokens) only, the Bi-Directional attention looks to both sides. For this reason this attention called as Bi-Directional.
This video do not cover math for this method. This lesson explain the logic in a high level how Bi-Directional attention works. To implement this method, there are developed Python packages to do it.
Bi-Directional attention is widely used in BERT, which means: Bidirectional Encoder Representation from Transformers.
This is the 3rd video in the mini course about Attention in NLP. Check it out the previous ones:
1. Encoder-Decoder attention and Dot-Product: youtu.be/z5upsjfVU9c
2. Self Attention: youtu.be/pfvYuX0Gjys
3. Bi-Directional Attention (this one).
4. Multi-Head attention (Upcoming).
You can read more about Bi-Directional Attention in the following sources:
- Standford University: Bidirectional Attention Flow with Self-Attention: https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1214/reports/final_reports/report170.pdf
- Medium.com article (BiDAF): towardsdatascience.com/the-definitive-guide-to-bi-directional-attention-flow-d0e96e9e666b
See you! - @DataScienceGarage
#attention #nlp #bidirectional #tokenizer #bert #selfattention #BiDAF #multihead #python #dotproduct #naturallanguageprocessing
This video tutorial explains how Bi-Directional Attention works in NLP. This attention mechanism is very similar to Self-attention method which was introduced in the previous video. The main difference is that Bi-Directional attention do not have Masking operation, which Self-attention has.
Also, while self-attention look to the previous words (or tokens) only, the Bi-Directional attention looks to both sides. For this reason this attention called as Bi-Directional.
This video do not cover math for this method. This lesson explain the logic in a high level how Bi-Directional attention works. To implement this method, there are developed Python packages to do it.
Bi-Directional attention is widely used in BERT, which means: Bidirectional Encoder Representation from Transformers.
This is the 3rd video in the mini course about Attention in NLP. Check it out the previous ones:
1. Encoder-Decoder attention and Dot-Product: youtu.be/z5upsjfVU9c
2. Self Attention: youtu.be/pfvYuX0Gjys
3. Bi-Directional Attention (this one).
4. Multi-Head attention (Upcoming).
You can read more about Bi-Directional Attention in the following sources:
- Standford University: Bidirectional Attention Flow with Self-Attention: https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1214/reports/final_reports/report170.pdf
- Medium.com article (BiDAF): towardsdatascience.com/the-definitive-guide-to-bi-directional-attention-flow-d0e96e9e666b
See you! - @DataScienceGarage
#attention #nlp #bidirectional #tokenizer #bert #selfattention #BiDAF #multihead #python #dotproduct #naturallanguageprocessing







![Perkūnkiemis - vakarinis aplinkelis 2016 (Žiema, gruodis) [TIMELAPSE]
Vilnius Timelapse - 30 sek.
Vieta: Pašilaičiai, Perkūnkiemis prie Vakarinio Vilniaus m. aplinkeliu (3 etapas) ties jungtimi su Ukmergės gatve.
Data: 2016 12 04
Vakarinio aplinkelio 3-iojo etapo atidarymas planuojamas jau 2016 metų gruodį.
BONUS: Perkūnkiemis iš oro skrendant lėktuvu: https://www.youtube.com/watch?v=Ypl73Iah5Kw#t=05m56s Perkūnkiemis - vakarinis aplinkelis 2016 (Žiema, gruodis) [TIMELAPSE]](https://i.ytimg.com/vi/ueAR32tB7qE/mqdefault.jpg)


