Uploaded August 2026 | Updated September 2026, 2 weeks ago
Why does AI-generated music still sometimes sound so... off? A hint: the number of parameters in a generative model doesn't have much to do with it. The problem most of the time is representation.
In this video, I break down one of the most overlooked ideas in generative AI: that a model is only as good as the way we encode reality for it. We'll go from the fundamentals of representation learning all the way to why most music generation systems are working with fundamentally broken inputs.
KEY TOPICS:
🧠 What representation actually encodes: explicit information, learnable relationships, invariances, and abstractions
🖼️ The representation spectrum: raw pixels vs. feature embeddings (like ResNet latents) vs. symbolic/semantic representations
⚙️ How feedforward networks, CNNs (VGG-16), and Transformers each bake in different inductive biases (convolution, recurrence, attention)
🎹 Why music representation is bad: piano rolls, MIDI, MusicXML, and event-based encodings like REMI all strip away musicological structure
🎼 What real music theory brings to the table: Lerdahl & Jackendoff's Generative Theory of Tonal Music, prolongation trees, time-span trees, and hierarchical grouping structure
🧩 Schema theory: how humans (and models) compress information into reusable patterns, from Gjerdingen's galant schemas to the Alberti bass
🌐 Beyond symbolic generation: combining score, audio, and rich representation for multimodal music models
📈 The big debate: is intelligence just brute-force scaling, or do we need smarter representations and inductive biases (a nod to JEPA and self-supervised world models)?
CONSULTING:
🚀 AI Music + Audio Consulting: valeriovelardoadvisor.com
📩 Get my AI Music content in your inbox for free: valeriovelardo.substack.com
JOIN THE SOUND OF AI COMMUNITY ON SLACK:
valeriovelardo.com/the-sound-of-ai-community
Why does AI-generated music still sometimes sound so... off? A hint: the number of parameters in a generative model doesn't have much to do with it. The problem most of the time is representation.
In this video, I break down one of the most overlooked ideas in generative AI: that a model is only as good as the way we encode reality for it. We'll go from the fundamentals of representation learning all the way to why most music generation systems are working with fundamentally broken inputs.
KEY TOPICS:
🧠 What representation actually encodes: explicit information, learnable relationships, invariances, and abstractions
🖼️ The representation spectrum: raw pixels vs. feature embeddings (like ResNet latents) vs. symbolic/semantic representations
⚙️ How feedforward networks, CNNs (VGG-16), and Transformers each bake in different inductive biases (convolution, recurrence, attention)
🎹 Why music representation is bad: piano rolls, MIDI, MusicXML, and event-based encodings like REMI all strip away musicological structure
🎼 What real music theory brings to the table: Lerdahl & Jackendoff's Generative Theory of Tonal Music, prolongation trees, time-span trees, and hierarchical grouping structure
🧩 Schema theory: how humans (and models) compress information into reusable patterns, from Gjerdingen's galant schemas to the Alberti bass
🌐 Beyond symbolic generation: combining score, audio, and rich representation for multimodal music models
📈 The big debate: is intelligence just brute-force scaling, or do we need smarter representations and inductive biases (a nod to JEPA and self-supervised world models)?
CONSULTING:
🚀 AI Music + Audio Consulting: valeriovelardoadvisor.com
📩 Get my AI Music content in your inbox for free: valeriovelardo.substack.com
JOIN THE SOUND OF AI COMMUNITY ON SLACK:
valeriovelardo.com/the-sound-of-ai-community










