Uploaded August 2020 | Updated September 2026, 2 weeks ago
Should you pre-train your compressed transformer model before knowledge distillation from an off-the-shelf teacher? This paper says yes and explores a few details behind this pipeline. Thanks for watching! Please Subscribe!
Paper Links:
Well-Read Students Learn Better: arxiv.org/pdf/1908.08962.pdf
Patient Knowledge Distillation: arxiv.org/pdf/1908.09355.pdf
DistilBERT: arxiv.org/pdf/1910.01108.pdf
Don't Stop Pretraining: arxiv.org/pdf/2004.10964.pdf
SimCLRv2: arxiv.org/pdf/2006.10029.pdf
AllenNLP MLM Demo: demo.allennlp.org/masked-lm?text=The%20doctor%20ran%20to%20the%20emergency%20room%20to%20see%20%5BMASK%5D%20patient.
HuggingFace Transformers: huggingface.co/transformers
Chapters
0:00 Beginning
1:55 Pre-trained Distillation
4:05 Distillation Recap
6:03 Pre-training with MLM
6:25 Should we pre-train the student model?
6:52 Unlabeled data in knowledge distillation
7:33 Comparison with DistilBERT and Patient KD
12:05 Datasets used for Analysis
13:38 Pre-training over Truncation
15:08 Depth over Width for Compact Models
16:04 Results for each task
16:34 Robustness to Transfer Set Size
17:24 Robustness to Domain Shift
Should you pre-train your compressed transformer model before knowledge distillation from an off-the-shelf teacher? This paper says yes and explores a few details behind this pipeline. Thanks for watching! Please Subscribe!
Paper Links:
Well-Read Students Learn Better: arxiv.org/pdf/1908.08962.pdf
Patient Knowledge Distillation: arxiv.org/pdf/1908.09355.pdf
DistilBERT: arxiv.org/pdf/1910.01108.pdf
Don't Stop Pretraining: arxiv.org/pdf/2004.10964.pdf
SimCLRv2: arxiv.org/pdf/2006.10029.pdf
AllenNLP MLM Demo: demo.allennlp.org/masked-lm?text=The%20doctor%20ran%20to%20the%20emergency%20room%20to%20see%20%5BMASK%5D%20patient.
HuggingFace Transformers: huggingface.co/transformers
Chapters
0:00 Beginning
1:55 Pre-trained Distillation
4:05 Distillation Recap
6:03 Pre-training with MLM
6:25 Should we pre-train the student model?
6:52 Unlabeled data in knowledge distillation
7:33 Comparison with DistilBERT and Patient KD
12:05 Datasets used for Analysis
13:38 Pre-training over Truncation
15:08 Depth over Width for Compact Models
16:04 Results for each task
16:34 Robustness to Transfer Set Size
17:24 Robustness to Domain Shift








