Uploaded June 2020 | Updated September 2026, 2 weeks ago
In this video, I break down a research paper that introduces a new DL architecture to recognise moods in songs 🎧 🎧 and explain its predictions.
The paper is called “Towards Explainable Music Emotion Recognition: The Route via Mid-Level Features” and was published at ISMIR in 2019. It was authored by researchers at the University of Linz.
Continue the conversation on The Sound o AI Slack community:
valeriovelardo.com/the-sound-of-ai-community
Read “Towards Explainable Music Emotion Recognition: The Route via Mid-Level Features” :
archives.ismir.net/ismir2019/paper/000027.pdf
A simple intro to VGG networks:
becominghuman.ai/what-is-the-vgg-neural-network-a590caa72643
Interested in hiring me as a consultant/freelancer?
valeriovelardo.com
Follow Valerio on Facebook:
facebook.com/TheSoundOfAI
Connect with Valerio on Linkedin:
linkedin.com/in/valeriovelardo
Follow Valerio on Twitter:
twitter.com/musikalkemist
In this video, I break down a research paper that introduces a new DL architecture to recognise moods in songs 🎧 🎧 and explain its predictions.
The paper is called “Towards Explainable Music Emotion Recognition: The Route via Mid-Level Features” and was published at ISMIR in 2019. It was authored by researchers at the University of Linz.
Continue the conversation on The Sound o AI Slack community:
valeriovelardo.com/the-sound-of-ai-community
Read “Towards Explainable Music Emotion Recognition: The Route via Mid-Level Features” :
archives.ismir.net/ismir2019/paper/000027.pdf
A simple intro to VGG networks:
becominghuman.ai/what-is-the-vgg-neural-network-a590caa72643
Interested in hiring me as a consultant/freelancer?
valeriovelardo.com
Follow Valerio on Facebook:
facebook.com/TheSoundOfAI
Connect with Valerio on Linkedin:
linkedin.com/in/valeriovelardo
Follow Valerio on Twitter:
twitter.com/musikalkemist


![MusicLM Generates Music From Text [Paper Breakdown]
MusicLM has taken the Music AI community by storm. The model published by Google is a step ahead towards text-based music generation in the audio realm. By leveraging a clever combination of deep learning base models, MusicLM generates convincing short music clips with good audio fidelity.
Join The Sound of AI Slack Community:
https://valeriovelardo.com/the-sound-of-ai-community/
MusicLM paper:
https://arxiv.org/abs/2301.11325
MusicLM demo:
https://google-research.github.io/seanet/musiclm/examples/
Music AI talent recruitment:
https://thesoundofai.com/
Interested in hiring me as a consultant/freelancer?
https://thesoundofai.com/consulting.html
The Sound of AI Academy:
https://the-sound-of-ai-academy.teachable.com/
Advanced Python Programming:
https://the-sound-of-ai-academy.teachable.com/p/advanced-python-programming
Connect with Valerio on Linkedin:
https://www.linkedin.com/in/valeriovelardo
Follow Valerio on Facebook:
https://www.facebook.com/TheSoundOfAI
Follow Valerio on Twitter:
https://twitter.com/musikalkemist
Content:
0:00 Intro
0:45 Text-to-music
2:15 MusicLM demo
3:31 Riffusion and Mubert AI
4:37 MusicLM architecture
6:08 Components overview
7:40 SoundStream
8:20 w2v-BERT
8:51 MuLan
10:51 Training
16:18 Inference
19:52 Experiments
21:52 Limitations
24:40 Thoughts on research procedure MusicLM Generates Music From Text [Paper Breakdown]](https://i.ytimg.com/vi/eaqO5uT0G9Q/mqdefault.jpg)







