DALL-E: Zero-Shot Text-to-Image Generation | Paper Explained @TheAIEpiphany
DALL-E: Zero-Shot Text-to-Image Generation | Paper Explained  @TheAIEpiphany
Uploaded July 2021 | Updated September 2026, 1 week ago
❤️ Become The AI Epiphany Patreon ❤️ ► patreon.com/theaiepiphany

In this video I cover DALL-E or "Zero-Shot Text-to-Image Generation" paper by OpenAI team.

They train a VQ-VAE to learn compressed image representations and then they train an autoregressive transformer on top of that discrete latent space and BPEd text.

The model learns to combine distinct concepts in a plausible way, image to image capabilities emerge, etc.

▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
✅ Paper: arxiv.org/abs/2102.12092

▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬

⌚️ Timetable:

00:00 What is DALL-E?
03:25 VQ-VAE blur problems
05:15 transformers, transformers, transformers!
07:10 Stage 1 and Stage 2 explained
07:30 Stage 1 VQ-VAE recap
10:00 Stage 2 autoregressive transformer
10:45 Some notes on ELBO
13:05 VQ-VAE modifications
17:20 Stage 2 in-depth
23:00 Results
24:25 Engineering, engineering, engineering
25:40 Automatic filtering via CLIP
27:40 More results
32:00 Additional image to image translation examples

▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
💰 BECOME A PATREON OF THE AI EPIPHANY ❤️

If these videos, GitHub projects, and blogs help you,
consider helping me out by supporting me on Patreon!

The AI Epiphany ► patreon.com/theaiepiphany
One-time donation:
paypal.com/paypalme/theaiepiphany

Much love! ❤️


Huge thank you to these AI Epiphany patreons:
Eli Mahler
Petar Veličković
Zvonimir Sabljic

▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬

💡 The AI Epiphany is a channel dedicated to simplifying the field of AI using creative visualizations and in general, a stronger focus on geometrical and visual intuition, rather than the algebraic and numerical "intuition".

▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
👋 CONNECT WITH ME ON SOCIAL
LinkedIn ► linkedin.com/in/aleksagordic
Twitter ► twitter.com/gordic_aleksa

Instagram ► instagram.com/aiepiphany
Facebook ► facebook.com/aiepiphany

👨‍👩‍👧‍👦 JOIN OUR DISCORD COMMUNITY:
Discord ► discord.gg/peBrCpheKE

📢 SUBSCRIBE TO MY MONTHLY AI NEWSLETTER:
Substack ► aiepiphany.substack.com

💻 FOLLOW ME ON GITHUB FOR COOL PROJECTS:
GitHub ► github.com/gordicaleksa

📚 FOLLOW ME ON MEDIUM:
Medium ► gordicaleksa.medium.com
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬

#dalle #openai #generativemodeling
DALL-E: Zero-Shot Text-to-Image Generation | Paper ExplainedDay 24: Open NLLB - back from China, analyzing spikes, preparing HBS run (Pt 2)OpenAI CLIP | Machine Learning Coding SeriesFake It Till You Make It (Microsoft) | Paper ExplainedLLaMA 2 w/ Thomas Scialom (LLaMA 2 lead)Day 24: Open NLLB - back from China, fuzzy dedup, preparing HBS run (Pt 1)Day 12: Open NLLB - on-boarding doc for new-joiners, eval (Pt 2.)Day 22: Open NLLB - HBS data analysis, split into Cyrillic & Latin (Pt 2)Day 4: Training 600M NLLB - data preps (Pt. 2)Day 4: Training 600M NLLB - data preps (Pt. 3)Day 15: Open NLLB - wrapping up detecting hallucinations paper (Pt 3)OpenAI GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
Aleksa Gordić - The AI Epiphany |

DALL-E: Zero-Shot Text-to-Image Generation | Paper Explained

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER