Uploaded December 2020 | Updated September 2026, 3 weeks ago
Tutorial 1, Hot Chips 32 (2020), Sunday, August 16, 2020.
Organizer: Paulius Micikevicius, NVIDIA
NOTE: The first 15 minutes or so of this tutorial was not recorded due to some technical problems as we started up Hot Chips streaming for the first time. As a result, you may want to review the first slides covering basic concepts from the proceedings before watching the first talk!
After an introduction to the principles of large-scale machine learning, this tutorial includes a number of talks describing the experiences that several companies have had while scaling up ML workloads from small-scale analyses typically used for most research to the large-scale systems that can analyze vast quantities of data to train complex ML models.
Fundamentals of Scaling Out ML Training
Paulius Micikevicius, NVIDIA
— Part I: Scale Out Systems
DGX A100 SuperPOD
Michael Houston, NVIDIA
Google TPU Pod
Sameer Kumar and Dehao Chen, Google
Cerebras System
Natalia Vassilieva, Cerebras
— Part II: Scale Out Training Experiences
Megatron Language Model
Mohammad Shoeybi, NVIDIA
Distributed Parameter Server for Massive Recommender System
Weijie Zhao, Baidu
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Zhifeng Chen, Google
Tutorial 1, Hot Chips 32 (2020), Sunday, August 16, 2020.
Organizer: Paulius Micikevicius, NVIDIA
NOTE: The first 15 minutes or so of this tutorial was not recorded due to some technical problems as we started up Hot Chips streaming for the first time. As a result, you may want to review the first slides covering basic concepts from the proceedings before watching the first talk!
After an introduction to the principles of large-scale machine learning, this tutorial includes a number of talks describing the experiences that several companies have had while scaling up ML workloads from small-scale analyses typically used for most research to the large-scale systems that can analyze vast quantities of data to train complex ML models.
Fundamentals of Scaling Out ML Training
Paulius Micikevicius, NVIDIA
— Part I: Scale Out Systems
DGX A100 SuperPOD
Michael Houston, NVIDIA
Google TPU Pod
Sameer Kumar and Dehao Chen, Google
Cerebras System
Natalia Vassilieva, Cerebras
— Part II: Scale Out Training Experiences
Megatron Language Model
Mohammad Shoeybi, NVIDIA
Distributed Parameter Server for Massive Recommender System
Weijie Zhao, Baidu
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Zhifeng Chen, Google










