AudioGen: Textually Guided Audio Generation | Text To Audio | Paper Explained @TheAIEpiphany
AudioGen: Textually Guided Audio Generation | Text To Audio | Paper Explained  @TheAIEpiphany
Uploaded October 2022 | Updated September 2026, 1 week ago
❀️ Become The AI Epiphany Patreon ❀️
patreon.com/theaiepiphany

πŸ‘¨β€πŸ‘©β€πŸ‘§β€πŸ‘¦ Join our Discord community πŸ‘¨β€πŸ‘©β€πŸ‘§β€πŸ‘¦
discord.gg/peBrCpheKE

In this video I do a deep dive of the recent "AudioGen: Textually Guided Audio Generation | Paper Explained" paper that introduced text-guided audio synthesis.

In a nutshell, it's the VQ-VAE/GAN idea applied to the audio modality.

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬
βœ… Paper: felixkreuk.github.io/text2audio_arxiv_samples/paper.pdf
βœ… Site: felixkreuk.github.io/text2audio_arxiv_samples

βœ… 3B1B on Fourier transform: youtube.com/watch?v=spUNpyF58BY
β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬

⌚️ Timetable:
00:00 Intro
01:17 Why is text-to-audio hard?
02:51 Comparison with VQ-GAN
05:15 Comparison with SoundStream
06:20 AudioGen overview
09:10 Deep dive: audio representation, LSTM
14:05 Losses explained
17:40 Complex-valued STFTs
21:57 Audio Language Modeling
23:37 Multi-stream audio inputs
25:32 Data and augmentations
29:05 Results
35:28 Outro

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬
πŸ’° BECOME A PATREON OF THE AI EPIPHANY ❀️

If these videos, GitHub projects, and blogs help you,
consider helping me out by supporting me on Patreon!

The AI Epiphany - patreon.com/theaiepiphany
One-time donation - paypal.com/paypalme/theaiepiphany

Huge thank you to these AI Epiphany patreons:
Eli Mahler
Petar VeličkoviΔ‡

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬

πŸ’Ό LinkedIn - linkedin.com/in/aleksagordic
🐦 Twitter - twitter.com/gordic_aleksa
πŸ‘¨β€πŸ‘©β€πŸ‘§β€πŸ‘¦ Discord - discord.gg/peBrCpheKE

πŸ“Ί YouTube - youtube.com/c/TheAIEpiphany
πŸ“š Medium - gordicaleksa.medium.com
πŸ’» GitHub - github.com/gordicaleksa
πŸ“’ AI Newsletter - aiepiphany.substack.com

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬

#audiogen #audiosynthesis #multimodal
AudioGen: Textually Guided Audio Generation | Text To Audio | Paper ExplainedDay 8: Meta NLLB - analyzing filtered data, data preps (Pt. 2)Day 28: Open NLLB - debugging fuzzy dedup, training fasttext LID (Pt 1)Day 9: Open NLLB - improving data download scripts (Pt 2.)How I Got a Job at DeepMind as a Research Engineer (without a Machine Learning Degree!)Day 3: LASER & GPT-NeoX (Pt. 3 resume)Machine Learning with JAX - From Zero to Hero | Tutorial #1Day 21: Open NLLB - data work (Serbian, Croatian, Bosnian) (Pt 1)Day 21: Open NLLB - data work (Serbian, Croatian, Bosnian) (Pt 2)Day 1 - Replicating Metas NLLB - SeamlessM4T paper (Pt. 2)Day 16: Open NLLB - Weights & Biases debugging session :) (Pt 3)GANs N Roses: Stable, Controllable, Diverse Image to Image Translation | Paper Explained
Aleksa Gordić - The AI Epiphany |

AudioGen: Textually Guided Audio Generation | Text To Audio | Paper Explained

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER