High Fidelity Neural Audio Compression | Paper & Code Explained @TheAIEpiphany
High Fidelity Neural Audio Compression | Paper & Code Explained  @TheAIEpiphany
Uploaded November 2022 | Updated September 2026, 1 week ago
❀️ Become The AI Epiphany Patreon ❀️
patreon.com/theaiepiphany

πŸ‘¨β€πŸ‘©β€πŸ‘§β€πŸ‘¦ Join our Discord community πŸ‘¨β€πŸ‘©β€πŸ‘§β€πŸ‘¦
discord.gg/peBrCpheKE

In this video I cover the "High Fidelity Neural Audio Compression" paper and code.

With 6 kbps they already get the same audio quality (as measured by the subjective MUSHRA metric) as mp3 at 64 kbps! 10x compression rate! This is super important as streaming video+audio makes for ~82% of total internet traffic!

Lots of ideas we've already seen in previous paper overview videos such as VQ-VAE, VQ-GAN, and AudioGen applied to the problem of audio compression.

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬
βœ… Paper: arxiv.org/abs/2210.13438
βœ… Code: github.com/facebookresearch/encodec
β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬

⌚️ Timetable:
00:00 Intro
02:37 Paper walk-through: high level overview
12:05 Residual Vector Quantization
18:05 Reducing the BW using arithmetic coding and transformers
20:05 Loss formulations and results
23:40 Code walk-through
26:00 EnCodec architecture
28:20 Residual Vector Quantizer module
32:55 Loading the audio signal
34:35 Compression - a forward pass through the encoder
38:00 Quantization forward pass
42:35 Efficiently packing the bits
45:25 Using LM to further compress audio
57:50 Outro

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬
πŸ’° BECOME A PATREON OF THE AI EPIPHANY ❀️

If these videos, GitHub projects, and blogs help you,
consider helping me out by supporting me on Patreon!

The AI Epiphany - patreon.com/theaiepiphany
One-time donation - paypal.com/paypalme/theaiepiphany

Huge thank you to these AI Epiphany patreons:
Eli Mahler
Petar VeličkoviΔ‡

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬

πŸ’Ό LinkedIn - linkedin.com/in/aleksagordic
🐦 Twitter - twitter.com/gordic_aleksa
πŸ‘¨β€πŸ‘©β€πŸ‘§β€πŸ‘¦ Discord - discord.gg/peBrCpheKE

πŸ“Ί YouTube - youtube.com/c/TheAIEpiphany
πŸ“š Medium - gordicaleksa.medium.com
πŸ’» GitHub - github.com/gordicaleksa
πŸ“’ AI Newsletter - aiepiphany.substack.com

β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬β–¬

#neural #audio #compression
High Fidelity Neural Audio Compression | Paper & Code ExplainedDay 24: Open NLLB - back from China, filtering HBS data (Pt 3)When Vision Transformers Outperform ResNets without Pretraining | Paper ExplainedDeepMind DetCon: Efficient Visual Pretraining with Contrastive Detection | Paper ExplainedArchie: an engineering AGI for Dyson Spheres | P-1 AI | $23 million seed roundDay 13: Open NLLB - Aya, dedup sharding analysis, analyzing training (Pt 2.)BigScience BLOOM | 3D Parallelism Explained | Large Language Models | ML Coding SeriesOpenAI DALL-E 3 with James Betker (1st author)Day 10: Open NLLB - evaluation data, filtering (Pt 3.)Day 13: Open NLLB - FSDP paper, kicking off the 1st run (Pt 1.)Day 8: Meta NLLB - analyzing the training script, Jais paper (Pt. 1)Channel Update: vacation, leaving Microsoft, approaching 10k subs and more!
Aleksa Gordić - The AI Epiphany |

High Fidelity Neural Audio Compression | Paper & Code Explained

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER