Stable Diffusion: High-Resolution Image Synthesis with Latent Diffusion Models | ML Coding Series @TheAIEpiphany
Stable Diffusion: High-Resolution Image Synthesis with Latent Diffusion Models | ML Coding Series  @TheAIEpiphany
Uploaded September 2022 | Updated September 2026, 1 week ago
❀️ Become The AI Epiphany Patreon ❀️
patreon.com/theaiepiphany

πŸ‘¨β€πŸ‘©β€πŸ‘§β€πŸ‘¦ Join our Discord community πŸ‘¨β€πŸ‘©β€πŸ‘§β€πŸ‘¦
discord.gg/peBrCpheKE

If you want to understand how stable diffusion exactly works behind the scenes this video is for you. I do a deep dive into the code behind Stable Diffusion explaining:
1. First stage autoencoder training (autoencoder with KL regularization)
2. Latent Diffusion Model training (UNet + conditioning model)
3. Sampling using PLMS scheduler

Stable diffusion directly builds upon the "High-Resolution Image Synthesis with Latent Diffusion Models" paper so I do a deep dive into the code behind this paper. Let me know how you like this one - feedback is welcome as always!

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬
βœ… Stable diffusion repo: github.com/CompVis/stable-diffusion
βœ… LDM repo: github.com/CompVis/latent-diffusion

βœ… LDM paper: arxiv.org/abs/2112.10752
βœ… VQ-GAN (taming transformers) paper: arxiv.org/abs/2012.09841
βœ… PLMS paper: arxiv.org/abs/2202.09778

βœ… Imagenette dataset: github.com/fastai/imagenette
β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬

⌚️ Timetable:
00:00:00 Intro: why is Stable Diffusion important
00:03:50 Background knowledge: VQ-GAN, LDM, PLMS papers
00:09:20 Setup for a minimal code walk-through
00:13:30 Autoencoder with KL regularization training
00:17:15 LPIPS (perceptual loss) with discriminator loss
00:21:30 Loading ImageNet data and PyTorch Lightning training loop
00:26:35 Forward pass through the autoencoder
00:30:12 Loss calculation
00:32:08 Perceptual loss
00:36:30 KL and GAN generator loss
00:40:55 Discriminator loss
00:42:45 Summarizing the autoencoder training
00:45:44 LDM training
00:57:00 Encoding the image into the latent space
01:00:12 Forward pass through the LDM
01:01:22 LDM loss
01:04:08 Integrating conditioning via cross attention
01:10:34 Sampling using PLMS
01:16:02 CLIP
01:19:20 Classifier free guidance
01:22:19 Sampling code
01:26:42 Diffusion connection to differential equations (PLMS paper)
01:37:00 Quick glimpse into the safety check function
01:39:20 Outro

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬
πŸ’° BECOME A PATREON OF THE AI EPIPHANY ❀️

If these videos, GitHub projects, and blogs help you,
consider helping me out by supporting me on Patreon!

The AI Epiphany - patreon.com/theaiepiphany
One-time donation - paypal.com/paypalme/theaiepiphany

Huge thank you to these AI Epiphany patreons:
Eli Mahler
Petar VeličkoviΔ‡

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬

πŸ“„ Website - gordicaleksa.com
πŸ’Ό LinkedIn - linkedin.com/in/aleksagordic
🐦 Twitter - twitter.com/gordic_aleksa
πŸ‘¨β€πŸ‘©β€πŸ‘§β€πŸ‘¦ Discord - discord.gg/peBrCpheKE

πŸ“Ί YouTube - youtube.com/c/TheAIEpiphany
πŸ“š Medium - gordicaleksa.medium.com
πŸ’» GitHub - github.com/gordicaleksa
πŸ“’ AI Newsletter - aiepiphany.substack.com

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬

#stablediffusion #latentdiffusion #imagesynthesis
Stable Diffusion: High-Resolution Image Synthesis with Latent Diffusion Models | ML Coding SeriesDay 12: Open NLLB - on-boarding doc for new-joiners, eval (Pt 2.)Channel update: moving to London in 2 days, new MLOps seriesBuild AI Agents using Integrail (Halloween special)building the best RLHF (TRLX) library w/ Louis CastricatoALiBi | Train Short, Test Long: Attention With Linear Biases Enables Input Length ExtrapolationDay 23: Open NLLB - day before China trip! Refactoring & pushing  changes (Pt 1 cont.)Hyperbolic Graph Convolutional Networks | Geometric ML Paper ExplainedDiffusion Models Beat GANs on Image Synthesis | ML Coding Series | Part 2Facebook AIs DINO | PyTorch Code ExplainedUltimate Guide To Scaling ML Models - Megatron-LM | ZeRO | DeepSpeed | Mixed PrecisionT0: Multitask Prompted Training Enables Zero-Shot Task Generalization | Paper Explained
Aleksa Gordić - The AI Epiphany |

Stable Diffusion: High-Resolution Image Synthesis with Latent Diffusion Models | ML Coding Series

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER