Uploaded December 2021 | Updated September 2026, 3 weeks ago
Tutorial 1, Part 3, Hot Chips 33 (2021), Sunday, August 22, 2021.
Organizers: Natalia Vassilieva, Cerebras, David Kanter, Real World Insights, and Vartika Singh, NVIDIA
Machine learning is a rich, varied, and rapidly evolving field. This tutorial explores the applications, performance characteristics, and key challenges of many different unique workloads across training and inference. In particular, it focuses on hardware/software co-optimization for the industry-standard MLPerf™ benchmarks and selected applications and considerations at prominent cloud players.
---------------------------
In Part 3 of the tutorial, ML researchers from Amazon, Google, and Microsoft describe how they addressed performance problems encountered in particular software scenarios: graph-based workloads and very large-scale training.
The Nature of Graph Neural Network Workloads
Da Zheng & George Karypis, Amazon
Challenges in large scale training of Giant models on large TPU machines
Sameer Kumar, Google
ZeRO-Infinity and DeepSpeed: Breaking the device Memory Wall for Extreme Scale Deep Learning
Yuxiong He & Samyam Rajbhandari, Microsoft
Tutorial 1, Part 3, Hot Chips 33 (2021), Sunday, August 22, 2021.
Organizers: Natalia Vassilieva, Cerebras, David Kanter, Real World Insights, and Vartika Singh, NVIDIA
Machine learning is a rich, varied, and rapidly evolving field. This tutorial explores the applications, performance characteristics, and key challenges of many different unique workloads across training and inference. In particular, it focuses on hardware/software co-optimization for the industry-standard MLPerf™ benchmarks and selected applications and considerations at prominent cloud players.
---------------------------
In Part 3 of the tutorial, ML researchers from Amazon, Google, and Microsoft describe how they addressed performance problems encountered in particular software scenarios: graph-based workloads and very large-scale training.
The Nature of Graph Neural Network Workloads
Da Zheng & George Karypis, Amazon
Challenges in large scale training of Giant models on large TPU machines
Sameer Kumar, Google
ZeRO-Infinity and DeepSpeed: Breaking the device Memory Wall for Extreme Scale Deep Learning
Yuxiong He & Samyam Rajbhandari, Microsoft










