Uploaded August 2021 | Updated September 2026, 1 week ago
❤️ Become The AI Epiphany Patreon ❤️ ► patreon.com/theaiepiphany
In this video I cover the "Do Vision Transformers See Like Convolutional Neural Networks?" paper. They dissect ViTs and ResNets and show the differences in the features learned as well as what contributes to those differences (like the amount of data used, skip connections, etc.).
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
✅ Paper: arxiv.org/abs/2108.08810
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
⌚️ Timetable:
00:00 Intro
00:45 Contrasting features in ViTs vs CNNs
06:45 Global vs Local receptive fields
13:55 Data matters, mr. obvious
17:40 Contrasting receptive fields
20:30 Data flow through CLS vs spatial tokens
23:30 Skip connections matter a lot in ViTs
24:20 Spatial information is preserved in ViTs
30:10 Features evolution with the amount of data
32:20 Outro
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
💰 BECOME A PATREON OF THE AI EPIPHANY ❤️
If these videos, GitHub projects, and blogs help you,
consider helping me out by supporting me on Patreon!
The AI Epiphany ► patreon.com/theaiepiphany
One-time donation:
paypal.com/paypalme/theaiepiphany
Much love! ❤️
Huge thank you to these AI Epiphany patreons:
Eli Mahler
Petar Veličković
Zvonimir Sabljic
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
💡 The AI Epiphany is a channel dedicated to simplifying the field of AI using creative visualizations and in general, a stronger focus on geometrical and visual intuition, rather than the algebraic and numerical "intuition".
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
👋 CONNECT WITH ME ON SOCIAL
LinkedIn ► linkedin.com/in/aleksagordic
Twitter ► twitter.com/gordic_aleksa
Instagram ► instagram.com/aiepiphany
Facebook ► facebook.com/aiepiphany
👨👩👧👦 JOIN OUR DISCORD COMMUNITY:
Discord ► discord.gg/peBrCpheKE
📢 SUBSCRIBE TO MY MONTHLY AI NEWSLETTER:
Substack ► aiepiphany.substack.com
💻 FOLLOW ME ON GITHUB FOR ML PROJECTS:
GitHub ► github.com/gordicaleksa
📚 FOLLOW ME ON MEDIUM:
Medium ► gordicaleksa.medium.com
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
#vision #transformers #cnns
❤️ Become The AI Epiphany Patreon ❤️ ► patreon.com/theaiepiphany
In this video I cover the "Do Vision Transformers See Like Convolutional Neural Networks?" paper. They dissect ViTs and ResNets and show the differences in the features learned as well as what contributes to those differences (like the amount of data used, skip connections, etc.).
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
✅ Paper: arxiv.org/abs/2108.08810
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
⌚️ Timetable:
00:00 Intro
00:45 Contrasting features in ViTs vs CNNs
06:45 Global vs Local receptive fields
13:55 Data matters, mr. obvious
17:40 Contrasting receptive fields
20:30 Data flow through CLS vs spatial tokens
23:30 Skip connections matter a lot in ViTs
24:20 Spatial information is preserved in ViTs
30:10 Features evolution with the amount of data
32:20 Outro
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
💰 BECOME A PATREON OF THE AI EPIPHANY ❤️
If these videos, GitHub projects, and blogs help you,
consider helping me out by supporting me on Patreon!
The AI Epiphany ► patreon.com/theaiepiphany
One-time donation:
paypal.com/paypalme/theaiepiphany
Much love! ❤️
Huge thank you to these AI Epiphany patreons:
Eli Mahler
Petar Veličković
Zvonimir Sabljic
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
💡 The AI Epiphany is a channel dedicated to simplifying the field of AI using creative visualizations and in general, a stronger focus on geometrical and visual intuition, rather than the algebraic and numerical "intuition".
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
👋 CONNECT WITH ME ON SOCIAL
LinkedIn ► linkedin.com/in/aleksagordic
Twitter ► twitter.com/gordic_aleksa
Instagram ► instagram.com/aiepiphany
Facebook ► facebook.com/aiepiphany
👨👩👧👦 JOIN OUR DISCORD COMMUNITY:
Discord ► discord.gg/peBrCpheKE
📢 SUBSCRIBE TO MY MONTHLY AI NEWSLETTER:
Substack ► aiepiphany.substack.com
💻 FOLLOW ME ON GITHUB FOR ML PROJECTS:
GitHub ► github.com/gordicaleksa
📚 FOLLOW ME ON MEDIUM:
Medium ► gordicaleksa.medium.com
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
#vision #transformers #cnns










