Uploaded October 2022 | Updated September 2026, 1 week ago
β€οΈ Become The AI Epiphany Patreon β€οΈ
patreon.com/theaiepiphany
π¨βπ©βπ§βπ¦ Join our Discord community π¨βπ©βπ§βπ¦
discord.gg/peBrCpheKE
In this video I do a deep dive of the recent "AudioGen: Textually Guided Audio Generation | Paper Explained" paper that introduced text-guided audio synthesis.
In a nutshell, it's the VQ-VAE/GAN idea applied to the audio modality.
β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬
β Paper: felixkreuk.github.io/text2audio_arxiv_samples/paper.pdf
β Site: felixkreuk.github.io/text2audio_arxiv_samples
β 3B1B on Fourier transform: youtube.com/watch?v=spUNpyF58BY
β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬
βοΈ Timetable:
00:00 Intro
01:17 Why is text-to-audio hard?
02:51 Comparison with VQ-GAN
05:15 Comparison with SoundStream
06:20 AudioGen overview
09:10 Deep dive: audio representation, LSTM
14:05 Losses explained
17:40 Complex-valued STFTs
21:57 Audio Language Modeling
23:37 Multi-stream audio inputs
25:32 Data and augmentations
29:05 Results
35:28 Outro
β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬
π° BECOME A PATREON OF THE AI EPIPHANY β€οΈ
If these videos, GitHub projects, and blogs help you,
consider helping me out by supporting me on Patreon!
The AI Epiphany - patreon.com/theaiepiphany
One-time donation - paypal.com/paypalme/theaiepiphany
Huge thank you to these AI Epiphany patreons:
Eli Mahler
Petar VeliΔkoviΔ
β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬
πΌ LinkedIn - linkedin.com/in/aleksagordic
π¦ Twitter - twitter.com/gordic_aleksa
π¨βπ©βπ§βπ¦ Discord - discord.gg/peBrCpheKE
πΊ YouTube - youtube.com/c/TheAIEpiphany
π Medium - gordicaleksa.medium.com
π» GitHub - github.com/gordicaleksa
π’ AI Newsletter - aiepiphany.substack.com
β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬
#audiogen #audiosynthesis #multimodal
β€οΈ Become The AI Epiphany Patreon β€οΈ
patreon.com/theaiepiphany
π¨βπ©βπ§βπ¦ Join our Discord community π¨βπ©βπ§βπ¦
discord.gg/peBrCpheKE
In this video I do a deep dive of the recent "AudioGen: Textually Guided Audio Generation | Paper Explained" paper that introduced text-guided audio synthesis.
In a nutshell, it's the VQ-VAE/GAN idea applied to the audio modality.
β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬
β Paper: felixkreuk.github.io/text2audio_arxiv_samples/paper.pdf
β Site: felixkreuk.github.io/text2audio_arxiv_samples
β 3B1B on Fourier transform: youtube.com/watch?v=spUNpyF58BY
β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬
βοΈ Timetable:
00:00 Intro
01:17 Why is text-to-audio hard?
02:51 Comparison with VQ-GAN
05:15 Comparison with SoundStream
06:20 AudioGen overview
09:10 Deep dive: audio representation, LSTM
14:05 Losses explained
17:40 Complex-valued STFTs
21:57 Audio Language Modeling
23:37 Multi-stream audio inputs
25:32 Data and augmentations
29:05 Results
35:28 Outro
β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬
π° BECOME A PATREON OF THE AI EPIPHANY β€οΈ
If these videos, GitHub projects, and blogs help you,
consider helping me out by supporting me on Patreon!
The AI Epiphany - patreon.com/theaiepiphany
One-time donation - paypal.com/paypalme/theaiepiphany
Huge thank you to these AI Epiphany patreons:
Eli Mahler
Petar VeliΔkoviΔ
β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬
πΌ LinkedIn - linkedin.com/in/aleksagordic
π¦ Twitter - twitter.com/gordic_aleksa
π¨βπ©βπ§βπ¦ Discord - discord.gg/peBrCpheKE
πΊ YouTube - youtube.com/c/TheAIEpiphany
π Medium - gordicaleksa.medium.com
π» GitHub - github.com/gordicaleksa
π’ AI Newsletter - aiepiphany.substack.com
β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬β¬
#audiogen #audiosynthesis #multimodal










