Ultimate Guide To Scaling ML Models - Megatron-LM | ZeRO | DeepSpeed | Mixed Precision @TheAIEpiphany
Ultimate Guide To Scaling ML Models - Megatron-LM | ZeRO | DeepSpeed | Mixed Precision  @TheAIEpiphany
Uploaded August 2022 | Updated September 2026, 1 week ago
πŸš€ Sign up for AssemblyAI's speech API using my link πŸš€
assemblyai.com/?utm_source=youtube&utm_medium=social&utm_campaign=theaiepiphany

πŸ‘¨β€πŸ‘©β€πŸ‘§β€πŸ‘¦ Join our Discord community πŸ‘¨β€πŸ‘©β€πŸ‘§β€πŸ‘¦
discord.gg/peBrCpheKE

In this video I show you what it takes to scale ML models up to trillions of parameters!

I cover the fundamental ideas behind all of the recent big ML models you must have heard of like Meta's OPT-175B, BigScience BLOOM 176B, EleutherAI's GPT-NeoX-20B, GPT-J, OpenAI's GPT-3, Google's PaLM, DeepMind's Chinchilla/Gopher models, etc.

I cover the ideas of data parallelism, model/pipeline parallelism (e.g. GPipe, PipeDream, etc.), model/tensor parallelism (Megatron-LM), activation checkpointing, mixed precision training, ZeRO (zero redundancy optimizer) from Microsoft's DeepSpeed library and many more.

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬
Papers:
βœ… Megatron-LM paper: arxiv.org/abs/1909.08053
βœ… ZeRO (DeepSpeed) paper: arxiv.org/abs/1910.02054v3
βœ… Mixed precision training paper: arxiv.org/abs/1710.03740
βœ… Gpipe (pipeline parallelism) paper: arxiv.org/abs/1811.06965

Articles:
βœ… Collective ops: en.wikipedia.org/wiki/Collective_operation
βœ… IEEE float16 format: en.wikipedia.org/wiki/Half-precision_floating-point_format
βœ… Google Brain's bfloat16 format: cloud.google.com/blog/products/ai-machine-learning/bfloat16-the-secret-to-high-performance-on-cloud-tpus
β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬

⌚️ Timetable:
00:00:00 Intro to training Large ML models (trillions of params!)
00:02:04 (sponsored) AssemblyAI's speech transcription API
00:03:31 Data parallelism
00:01:52 Pipeline/model parallelism
00:14:52 Megatron-LM paper (tensor/model parallelism)
00:18:22 Splitting the MLP block vertically
00:30:07 Splitting the attention block vertically
00:39:24 Activation checkpointing
00:42:12Combining data + model parallelism
00:45:42 Scaling is all you need and 3D parallelism
00:47:57 Mixed precision training paper
00:49:57 Single vs half vs bfloat number formats
00:51:32 Storing master weights in single precision
00:55:41 Loss scaling
00:58:13 Arithmetic precision matters
01:00:32 ZeRO optimizer paper (DeepSpeed library)
01:06:37 Partitioning is all you need?
01:11:02 Where did all the memory go?
01:21:42 Outro

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬
πŸ’° BECOME A PATREON OF THE AI EPIPHANY ❀️

If these videos, GitHub projects, and blogs help you,
consider helping me out by supporting me on Patreon!

The AI Epiphany - patreon.com/theaiepiphany
One-time donation - paypal.com/paypalme/theaiepiphany

Huge thank you to these AI Epiphany patreons:
Eli Mahler
Petar VeličkoviΔ‡

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬

πŸ’Ό LinkedIn - linkedin.com/in/aleksagordic
🐦 Twitter - twitter.com/gordic_aleksa
πŸ‘¨β€πŸ‘©β€πŸ‘§β€πŸ‘¦ Discord - discord.gg/peBrCpheKE

πŸ“Ί YouTube - youtube.com/c/TheAIEpiphany
πŸ“š Medium - gordicaleksa.medium.com
πŸ’» GitHub - github.com/gordicaleksa
πŸ“’ AI Newsletter - aiepiphany.substack.com

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬

#scaling #deepspeed #megatron
Ultimate Guide To Scaling ML Models - Megatron-LM | ZeRO | DeepSpeed | Mixed PrecisionT0: Multitask Prompted Training Enables Zero-Shot Task Generalization | Paper ExplainedDay 17: Open NLLB - analyzing batch iterators (Pt 2)State Space Models w/ Albert Gu & Karan Goel (Cartesia AI)ConvNeXt: A ConvNet for the 2020s | Paper ExplainedVQ-GAN: Taming Transformers for High-Resolution Image Synthesis | Paper ExplainedDay 14: Open NLLB - exploring BLEU, chrF++, logging (Pt 3. cont.)DALL-E: Zero-Shot Text-to-Image Generation | Paper ExplainedDay 24: Open NLLB - back from China, analyzing spikes, preparing HBS run (Pt 2)OpenAI CLIP | Machine Learning Coding SeriesFake It Till You Make It (Microsoft) | Paper ExplainedLLaMA 2 w/ Thomas Scialom (LLaMA 2 lead)
Aleksa Gordić - The AI Epiphany |

Ultimate Guide To Scaling ML Models - Megatron-LM | ZeRO | DeepSpeed | Mixed Precision

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER