Uploaded December 2021 | Updated September 2026, 3 weeks ago
Tutorial 1, Part 2, Hot Chips 33 (2021), Sunday, August 22, 2021.
Organizers: Natalia Vassilieva, Cerebras, David Kanter, Real World Insights, and Vartika Singh, NVIDIA
Machine learning is a rich, varied, and rapidly evolving field. This tutorial explores the applications, performance characteristics, and key challenges of many different unique workloads across training and inference. In particular, it focuses on hardware/software co-optimization for the industry-standard MLPerf™ benchmarks and selected applications and considerations at prominent cloud players.
---------------------------
In Part 2 of the tutorial, ML researchers from Graphcore, Intel, and Facebook explain some of their efforts at optimizing ML applications on radically different types of hardware.
NOTE: Approximately the first 10 minutes or so of the Graphcore talk was not recorded due to some technical problems with the recording. As a result, you may want to review the first slides from the proceedings before watching the rest of that talk!
Software/hardware co-optimization on the IPU: An MLPerf™ case study
Mario Michael Krell, Graphcore
Deep Learning Inference Optimizations on CPUs
Guokai Ma, Intel
AI at Scale for the Modern Era
Carole-Jean Wu & Niket Agarwal, Facebook
Tutorial 1, Part 2, Hot Chips 33 (2021), Sunday, August 22, 2021.
Organizers: Natalia Vassilieva, Cerebras, David Kanter, Real World Insights, and Vartika Singh, NVIDIA
Machine learning is a rich, varied, and rapidly evolving field. This tutorial explores the applications, performance characteristics, and key challenges of many different unique workloads across training and inference. In particular, it focuses on hardware/software co-optimization for the industry-standard MLPerf™ benchmarks and selected applications and considerations at prominent cloud players.
---------------------------
In Part 2 of the tutorial, ML researchers from Graphcore, Intel, and Facebook explain some of their efforts at optimizing ML applications on radically different types of hardware.
NOTE: Approximately the first 10 minutes or so of the Graphcore talk was not recorded due to some technical problems with the recording. As a result, you may want to review the first slides from the proceedings before watching the rest of that talk!
Software/hardware co-optimization on the IPU: An MLPerf™ case study
Mario Michael Krell, Graphcore
Deep Learning Inference Optimizations on CPUs
Guokai Ma, Intel
AI at Scale for the Modern Era
Carole-Jean Wu & Niket Agarwal, Facebook










