Uploaded December 2023 | Updated September 2026, 2 weeks ago
Learn about the intuition, theory, and mathematics of the transformer. I focus on the decoder component. I dive deep into masked multi-head attention and all the other sublayers. Learn how to use transformers for music generation. I share tips and tricks from the trenches. I focus on the importance of music representation and music data for generation. I also offer insights into future research in neuro-symbolic integration that can help transformers become more robust for music generation.
Get the lecture slides:
github.com/musikalkemist/generativemusicaicourse/blob/main/18.%20Transformers%20-%20Part%202/Slides/18.%20Transformers%20Part%202.pdf
Website of the Generative Music AI Workshop in Barcelona:
https://www.upf.edu/web/mtg/generative-music-ai-workshop
Sign up to The Sound of AI Slack Community to join the discussion:
valeriovelardo.com/the-sound-of-ai-community
======================================
Interested in music AI consulting?
thesoundofai.com/consulting.html
Interested in music AI recruitment?
thesoundofai.com/recruitment.html
Become a Python ninja with my Advanced Python Programming course:
the-sound-of-ai-academy.teachable.com/p/advanced-python-programming
Connect with Valerio on LinkedIn:
linkedin.com/in/valeriovelardo
Follow Valerio on Twitter:
twitter.com/musikalkemist
======================================
Content
0:00 Intro
1:37 Decoder intuition
5:24 Decoder input
6:51 Decoder block
7:21 Training / inference discrepancy
13:10 Masked multi-head attention
20:04 Add & norm
20:25 Multi-head attention
32:40 Feedforward
33:30 Decoder block
35:09 Linear & softmax
38:11 Decoder step-by-step
42:14 Training a transformer
44:05 Music generation with transformers
47:45 Valerio's music generation transformer routine
51:51 Music data is key
53:15 Pros and cons
57:04 Most promising research
1:00:32 Key takeaways
1:02:58 What's up next?
Learn about the intuition, theory, and mathematics of the transformer. I focus on the decoder component. I dive deep into masked multi-head attention and all the other sublayers. Learn how to use transformers for music generation. I share tips and tricks from the trenches. I focus on the importance of music representation and music data for generation. I also offer insights into future research in neuro-symbolic integration that can help transformers become more robust for music generation.
Get the lecture slides:
github.com/musikalkemist/generativemusicaicourse/blob/main/18.%20Transformers%20-%20Part%202/Slides/18.%20Transformers%20Part%202.pdf
Website of the Generative Music AI Workshop in Barcelona:
https://www.upf.edu/web/mtg/generative-music-ai-workshop
Sign up to The Sound of AI Slack Community to join the discussion:
valeriovelardo.com/the-sound-of-ai-community
======================================
Interested in music AI consulting?
thesoundofai.com/consulting.html
Interested in music AI recruitment?
thesoundofai.com/recruitment.html
Become a Python ninja with my Advanced Python Programming course:
the-sound-of-ai-academy.teachable.com/p/advanced-python-programming
Connect with Valerio on LinkedIn:
linkedin.com/in/valeriovelardo
Follow Valerio on Twitter:
twitter.com/musikalkemist
======================================
Content
0:00 Intro
1:37 Decoder intuition
5:24 Decoder input
6:51 Decoder block
7:21 Training / inference discrepancy
13:10 Masked multi-head attention
20:04 Add & norm
20:25 Multi-head attention
32:40 Feedforward
33:30 Decoder block
35:09 Linear & softmax
38:11 Decoder step-by-step
42:14 Training a transformer
44:05 Music generation with transformers
47:45 Valerio's music generation transformer routine
51:51 Music data is key
53:15 Pros and cons
57:04 Most promising research
1:00:32 Key takeaways
1:02:58 What's up next?








![MusicLM Generates Music From Text [Paper Breakdown]
MusicLM has taken the Music AI community by storm. The model published by Google is a step ahead towards text-based music generation in the audio realm. By leveraging a clever combination of deep learning base models, MusicLM generates convincing short music clips with good audio fidelity.
Join The Sound of AI Slack Community:
https://valeriovelardo.com/the-sound-of-ai-community/
MusicLM paper:
https://arxiv.org/abs/2301.11325
MusicLM demo:
https://google-research.github.io/seanet/musiclm/examples/
Music AI talent recruitment:
https://thesoundofai.com/
Interested in hiring me as a consultant/freelancer?
https://thesoundofai.com/consulting.html
The Sound of AI Academy:
https://the-sound-of-ai-academy.teachable.com/
Advanced Python Programming:
https://the-sound-of-ai-academy.teachable.com/p/advanced-python-programming
Connect with Valerio on Linkedin:
https://www.linkedin.com/in/valeriovelardo
Follow Valerio on Facebook:
https://www.facebook.com/TheSoundOfAI
Follow Valerio on Twitter:
https://twitter.com/musikalkemist
Content:
0:00 Intro
0:45 Text-to-music
2:15 MusicLM demo
3:31 Riffusion and Mubert AI
4:37 MusicLM architecture
6:08 Components overview
7:40 SoundStream
8:20 w2v-BERT
8:51 MuLan
10:51 Training
16:18 Inference
19:52 Experiments
21:52 Limitations
24:40 Thoughts on research procedure MusicLM Generates Music From Text [Paper Breakdown]](https://i.ytimg.com/vi/eaqO5uT0G9Q/mqdefault.jpg)

