Make-A-Video: Text-To-Video Generation Without Text-Video Data | Paper Explained @TheAIEpiphany
Make-A-Video: Text-To-Video Generation Without Text-Video Data | Paper Explained  @TheAIEpiphany
Uploaded October 2022 | Updated September 2026, 1 week ago
πŸš€ Find out how to get started using Weights & Biases πŸš€
wandb.me/ai-epiphany

πŸ‘¨β€πŸ‘©β€πŸ‘§β€πŸ‘¦ Join our Discord community πŸ‘¨β€πŸ‘©β€πŸ‘§β€πŸ‘¦
discord.gg/peBrCpheKE

In this video I cover the latest text-to-video paper from Meta: "Make-A-Video: Text-To-Video Generation Without Text-Video Data".

I walk you through the 3-stage approach that consists of:
* Training a DALL-E 2 type of a model
* Integrating temporal information and tuning on unlabeled videos
* Fine-tuning the frame interpolation module.

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬
βœ… Paper: arxiv.org/abs/2209.14792
βœ… Website: https://makeavideo.studio/
β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬

⌚️ Timetable:
00:00 Intro
00:25 (sponsored) Weights & Biases
01:37 Going through the generations
06:15 High-level paper overview
10:50 Results
15:40 Limitations
16:30 Diving deep: DALL-E 2 backbone
23:35 Expanding to 3D - temporal info integration
32:39 Frame interpolation
37:24 3-stage training
41:28 Outro

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬
πŸ’° BECOME A PATREON OF THE AI EPIPHANY ❀️

If these videos, GitHub projects, and blogs help you,
consider helping me out by supporting me on Patreon!

The AI Epiphany - patreon.com/theaiepiphany
One-time donation - paypal.com/paypalme/theaiepiphany

Huge thank you to these AI Epiphany patreons:
Eli Mahler
Petar VeličkoviΔ‡

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬

πŸ’Ό LinkedIn - linkedin.com/in/aleksagordic
🐦 Twitter - twitter.com/gordic_aleksa
πŸ‘¨β€πŸ‘©β€πŸ‘§β€πŸ‘¦ Discord - discord.gg/peBrCpheKE

πŸ“Ί YouTube - youtube.com/c/TheAIEpiphany
πŸ“š Medium - gordicaleksa.medium.com
πŸ’» GitHub - github.com/gordicaleksa
πŸ“’ AI Newsletter - aiepiphany.substack.com

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬

#makeavideo #meta #texttovideo
Make-A-Video: Text-To-Video Generation Without Text-Video Data | Paper ExplainedBest LLM? Qwen 2 LLM w/ author Junyang LinDay 25: Open NLLB - filtering HBS (Pt 3)Day 19: Open NLLB - OPUS, downloading HBS parallel corpora (Pt 2)Day 20: Open NLLB - downloading HBS parallel corpora (Croatian - wrap up) (Pt 1)Day 29: Open NLLB - testing & improving fasttext HBS LID (Pt 3)Text Style Brush - Transfer of text aesthetics from a single example | Paper ExplainedHow to Build a Deep Learning Machine - Everything You Need To KnowDay 28: Open NLLB - debugging fuzzy dedup, training fasttext LID (Pt 2)Day 27: Open NLLB - filtering stage hyperparams, HBS LID detector (Pt 3 cont.)Day 21: Open NLLB - data work (Serbian, Croatian, Bosnian) (Pt 1 cont.)Day 30: Open NLLB - iterating on fasttext HBS LID, filtering (Pt 1)
Aleksa Gordić - The AI Epiphany |

Make-A-Video: Text-To-Video Generation Without Text-Video Data | Paper Explained

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER