Well-Read Students Learn Better @connor-shorten
Well-Read Students Learn Better  @connor-shorten
Uploaded August 2020 | Updated September 2026, 2 weeks ago
Should you pre-train your compressed transformer model before knowledge distillation from an off-the-shelf teacher? This paper says yes and explores a few details behind this pipeline. Thanks for watching! Please Subscribe!

Paper Links:
Well-Read Students Learn Better: arxiv.org/pdf/1908.08962.pdf
Patient Knowledge Distillation: arxiv.org/pdf/1908.09355.pdf
DistilBERT: arxiv.org/pdf/1910.01108.pdf
Don't Stop Pretraining: arxiv.org/pdf/2004.10964.pdf
SimCLRv2: arxiv.org/pdf/2006.10029.pdf
AllenNLP MLM Demo: demo.allennlp.org/masked-lm?text=The%20doctor%20ran%20to%20the%20emergency%20room%20to%20see%20%5BMASK%5D%20patient.
HuggingFace Transformers: huggingface.co/transformers

Chapters
0:00 Beginning
1:55 Pre-trained Distillation
4:05 Distillation Recap
6:03 Pre-training with MLM
6:25 Should we pre-train the student model?
6:52 Unlabeled data in knowledge distillation
7:33 Comparison with DistilBERT and Patient KD
12:05 Datasets used for Analysis
13:38 Pre-training over Truncation
15:08 Depth over Width for Compact Models
16:04 Results for each task
16:34 Robustness to Transfer Set Size
17:24 Robustness to Domain Shift
Well-Read Students Learn BetterAI Weekly Update - April 27th, 2020 (#19)MultiCite - New Research in Scientific Literature Mining!VidLanKDMeta Pseudo LabelsAI Weekly Update - June 9th, 2021 (#34!)Graph Embeddings and PyTorch-BigGraphDont Stop Pretraining!AI Weekly Update Preview - March 29th, 2021 (#30)MODALS: Modality-agnostic Automated Data Augmentation in the Latent Space
Connor Shorten |

Well-Read Students Learn Better

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER