L11.4 Why BatchNorm Works @SebastianRaschka
L11.4 Why BatchNorm Works  @SebastianRaschka
Uploaded March 2021 | Updated September 2026, 2 weeks ago
Sebastian's books: sebastianraschka.com/books

Slides: sebastianraschka.com/pdf/lecture-notes/stat453ss21/L11_norm-and-init__slides.pdf

BatchNorm papers:

Ioffe, S., & Szegedy, C. (2015). Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. In International Conference on Machine Learning (pp. 448-456). http://proceedings.mlr.press/v37/ioffe15.html

Santurkar, S., Tsipras, D., Ilyas, A., & Madry, A. (2018). How does batch normalization help optimization?. In Advances in Neural Information Processing Systems (pp. 2488-2498). arxiv.org/abs/1805.11604

Morcos, A. S., Barrett, D. G., Rabinowitz, N. C., & Botvinick, M. (2018). On the importance of single directions for generalization. arxiv.org/abs/1803.06959

Luo, P., Wang, X., Shao, W., & Peng, Z. (2018). Towards understanding regularization in batch normalization. arxiv.org/abs/1809.00846

Yang, G., Pennington, J., Rao, V., Sohl-Dickstein, J., & Schoenholz, S. S. (2019). A mean field theory of batch normalization. arxiv.org/abs/1902.08129

==============

Some Benchmarks: github.com/ducha-aiki/caffenet-benchmark/blob/master/batchnorm.md#bn----before-or-after-relu

-------

This video is part of my Introduction of Deep Learning course.

Next video: youtu.be/RsX01aYbQdI

The complete playlist: youtube.com/playlist?list=PLTKMiZHVd_2KJtIXOW0zFhFfBaJJilH51

A handy overview page with links to the materials: sebastianraschka.com/blog/2021/dl-course.html

-------

If you want to be notified about future videos, please consider subscribing to my channel: youtube.com/c/SebastianRaschka
L11.4 Why BatchNorm WorksL10.4 L2 Regularization for Neural NetsL13.9.2 Saving and Loading Models in PyTorchReinforcement Learning with Human Feedback (RLHF) in 4 minutesL19.5.2.5 GPT-v3: Language Models are Few-Shot LearnersL13.5 Whats The Difference Between Cross-Correlation And Convolution?L11.0 Input Normalization and Weight Initialization   Lecture OverviewBuild an LLM from Scratch 1: Set up your code environment13.3.2 Decision Trees & Random Forest Feature Importance (L13: Feature Selection)L13.9.1 LeNet-5 in PyTorchL17.4 Variational Autoencoder Loss FunctionL9.3.1 Multilayer Perceptron   Code Example Part 1/3 (Slide Overview)
Sebastian Raschka |

L11.4 Why BatchNorm Works

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER