Uploaded April 2021 | Updated September 2026, 2 weeks ago
Learn how to build an audio preprocessing pipeline for AI applications in Python. The pipeline batch preprocess audio files applying Short-Time Fourier Transform, zero-padding, min-max normalization all in one go!
Code:
github.com/musikalkemist/generating-sound-with-neural-networks/blob/main/12%20Preprocessing%20pipeline/preprocess.py
Free Spoken Digit Dataset:
github.com/Jakobovski/free-spoken-digit-dataset
===============================
Interested in hiring me as a consultant/freelancer?
valeriovelardo.com/
Join The Sound Of AI Slack community:
valeriovelardo.com/the-sound-of-ai-community
Follow Valerio on Facebook:
facebook.com/TheSoundOfAI
Connect with Valerio on Linkedin:
linkedin.com/in/valeriovelardo
Follow Valerio on Twitter:
twitter.com/musikalkemist
===============================
Content:
0:00 Intro
0:48 The Free Spoken Digit Dataset
1:37 Pipeline intuition + design
5:24 Implementating Loader
9:02 Implementing Padder
15:16 Implementing LogSpectrogramExtractor
20:21 Implementing MinMaxNormaliser
25:38 Implementing Preprocessing Pipeline
44:01 Implementing Saver
51:35 Recap of implemented classes
53:04 Prunning the preprocessing pipeline
58:17 Outro
Learn how to build an audio preprocessing pipeline for AI applications in Python. The pipeline batch preprocess audio files applying Short-Time Fourier Transform, zero-padding, min-max normalization all in one go!
Code:
github.com/musikalkemist/generating-sound-with-neural-networks/blob/main/12%20Preprocessing%20pipeline/preprocess.py
Free Spoken Digit Dataset:
github.com/Jakobovski/free-spoken-digit-dataset
===============================
Interested in hiring me as a consultant/freelancer?
valeriovelardo.com/
Join The Sound Of AI Slack community:
valeriovelardo.com/the-sound-of-ai-community
Follow Valerio on Facebook:
facebook.com/TheSoundOfAI
Connect with Valerio on Linkedin:
linkedin.com/in/valeriovelardo
Follow Valerio on Twitter:
twitter.com/musikalkemist
===============================
Content:
0:00 Intro
0:48 The Free Spoken Digit Dataset
1:37 Pipeline intuition + design
5:24 Implementating Loader
9:02 Implementing Padder
15:16 Implementing LogSpectrogramExtractor
20:21 Implementing MinMaxNormaliser
25:38 Implementing Preprocessing Pipeline
44:01 Implementing Saver
51:35 Recap of implemented classes
53:04 Prunning the preprocessing pipeline
58:17 Outro










