Welch Labs
How the Bizarre Path of Mars Reshaped Astronomy [Keplers Laws Part 1]
updated
Welch Labs Book: welchlabs.com/resources/ai-book-ezrzm
Welch Labs eBook: welchlabs.com/resources/the-welch-labs-illustrated-guide-to-ai-digital-download
Patreon: patreon.com/welchlabs
SECTIONS
0:00 - Harpy
4:39 - The Bitter Lesson
5:58 - Sutton Goes on a Podcast
8:22 - LLMs are Not Bitter Lesson Pilled?
9:19 - Supervised Learning
10:04 - Reinforcement Learning
10:32 - Work for Tufalabs!
11:50 - How AlphaGo Surpassed Humans
17:49 - RLHF and RLVR
18:41 - The Era of Experience
20:27 - My Take
21:05 - Welch Labs Book!
TECHNICAL NOTES
welchlabs.com/blog/2026/1/31/the-bitter-lesson-video-technical-notes
CODE
github.com/stephencwelch/manim_videos
REFERENCES
The Bitter Lesson: http://www.incompleteideas.net/IncIdeas/BitterLesson.html
Dwarkesh Patel's interview with Richard Sutton: youtube.com/watch?v=21EYKqUsPfg
AlphaGo vs Lee Sedol Match 4: youtube.com/watch?v=yCALyQRN3hw
Repurposed some board setups and heatmaps from: lesswrong.com/posts/FF8i6SLfKb4g7C4EL/inside-the-mind-of-a-superhuman-go-model-how-does-leela-zero-2?utm_source=chatgpt.com
Great HARPY video: youtube.com/watch?v=NiiDe2n-GeQ
Sutton, Richard S., and Andrew G. Barto. *Reinforcement learning: An introduction*. Vol. 1. No. 1. Cambridge: MIT press, 1998.
Averbuch, Amir, et al. "An IBM PC based large-vocabulary isolated-utterance speech recognizer." ICASSP'86. IEEE International Conference on Acoustics, Speech, and Signal Processing. Vol. 11. IEEE, 1986.
Radford, Alec, et al. "Language models are unsupervised multitask learners." OpenAI blog 1.8 (2019): 9.
Silver, David, et al. "Mastering the game of Go with deep neural networks and tree search." *nature* 529.7587 (2016): 484-489.
Silver, David, et al. "Mastering the game of go without human knowledge." *nature* 550.7676 (2017): 354-359.
Lowerre, Bruce. "The HARPY speech understanding system." *Readings in speech recognition*. 1990. 576-586.
Silver, David, and Richard S. Sutton. "Welcome to the era of experience." Google AI 1 (2025).
PATRONS
Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti, Brian Henry, Tim Palade, Petar Vecutin, Nicolas baumann, Jason Singh, Robert Riley, vornska, Barry Silverman, Jake Ehrlich, Mitch Jacobs, Lauren Steely, Jeff Eastman, Rodolfo Ibarra, Clark Barrus, Rob Napier, Andrew White, Richard B Johnston, abhiteja mandava, Burt Humburg, Kevin Mitchell, Daniel Sanchez, Ferdie Wang, Tripp Hill, Richard Harbaugh Jr, Prasad Raje, Kalle Aaltonen, Midori Switch Hound, Zach Wilson, Chris Seltzer, Ven Popov, Hunter Nelson, Amit Bueno, Scott Olsen, Johan Rimez, Shehryar Saroya, Tyler Christensen, Beckett Madden-Woods, Darrell Thomas, Javier Soto, U007D, Caleb Begly, Rick Rubenstein, Brent Hunsaker, Dan Patterson, Tchsurvives, Alex Adai, Walter Reade, Zyansheep, Walter Reade, Duncan Stannett, Reginald Carey, Jean-Manuel Izaret, dh71633, Adrian Rodriguez, Dimitar Stojanovski, Michael Harder, Peter Maldonado, Emily Pesce, David Johnston, Insang Song, FaeTheWolf, Stephen Taylor, KittenKaboodle, EMatter, PATRICKMCCORMACK, John Beahan, Cameron, Cole Jones, Garrett Thornburg, Jeroen W, Rohit Sharma, GlennB, Emmanuel Cortes, Katie Quinn, Karina C, Cakra WW, Mike Ton, Eric Gometz, MacCallister Higgins, Niko Drossos, David Eraso, Tom Zehle, Steve, Brian Lineburg, rjbl, Michael Loh, Perry Vais, Bengal0, Farhad Manjoo, Sara Chipps, Ellis Driscoll, William Taysom, Will Harmon, CK, Abdullah, Peter Cho, Leo Nikora, Griffin Smith, Ash Katnoria, Alex, Markus Hays Nielsen, Catherine H., Vi, David Dobáš, Peter Wang, Sina Sohangir, Danny Thomas, Julian Francis, Hans Adler, Jiayu Peng, Weston M, Youssouf da Silva, John Thomas, Samuel Costello, Sam Adams, Bryan Liles, Malaya Zemlya, Karl, Vahe Andonians, Mike Doughty, Larry Novelo, Jonas Acres, Ludicrum Rex, Robert Blumofe, Anthony Z
Created by: Sam Baskin, Matthew Cohen, Pranav Gundu, and Stephen Welch
Content ID: CFAQJOTYQHT7JYIT
ebook: welchlabs.com/resources/the-welch-labs-illustrated-guide-to-ai-digital-download
Patreon: patreon.com/welchlabs
Sections
0:00 - Intro
2:39 - Modular Addition
3:54 - The Model’s Perspective
6:52 - An Accidental Discovery at OpenAI
7:56 - It Groks!
8:49 - Some Clues
13:17 - New Welch Labs Book!
13:57 - Deeper into the model
15:13 - Linear Probes
16:59 - Clocks perform modular addition
19:17 - How do x and y interact exactly?
23:19 - It learns a trig identity?!
26:38 - Putting the pieces together & excluded loss
30:02 - Anthropic finds 6D manifolds
32:24 - Final thoughts
34:02 - Welch Labs update
Special thanks to Neel Nanda for discussing his work and Mech Interp with me, if you want to learn more about Mech Interp, check out Neel’s getting started post here: neelnanda.io/getting-started
Thanks to Emmanuel Ameisen and Wes Gurnee for discussing their work on Claude Haiku. Their paper is incredibly in depth and interesting: https://transformer-circuits.pub/2025/linebreaks/index.html
Really nice deeper breakdown on polynomial double descent from viewer Avaneesh Narla:
avaneeshnarla.com/blog/double-descent.html
OpenAI team’s grokking paper: arxiv.org/pdf/2201.02177. I wasn’t able to reach to team for comment on the origin story, but it is told here: youtube.com/watch?v=gYGWFjMf9JA&t=1236s
Nanda et al. arxiv.org/pdf/2301.05217v1
More on Grokking: quantamagazine.org/how-do-machines-grok-data-20240412
Code based on excellent these notebooks from Neel Nanda and collaborators:
colab.research.google.com/drive/1F6_1_cWXE5M7WocUcpQWp3v8z4b1jL20
Andrej Karapathy on “Summoning Ghosts”: https://karpathy.bearblog.dev/animals-vs-ghosts/
Code: github.com/stephencwelch/manim_videos/tree/master/_2025/grokking
Technical Notes
- It’s very natural for the attention layer to take the sum of it’s inputs (e.g. cos(kx)+cos(ky)), however we also find strong product terms. There’s a couple ways the network can compute products like cos(kx)cos(ky). One option is to approximate the product using ReLU activation functions (see Nanda’s notebooks for more). It’s also feasible for the attention block to do this, I found evidence of this is my own exploration
- In the first 2D fourier decomposition, we’re leaving out one component, specifically a “negative frequency component” → 0.26 * np.cos(2*np.pi*((4*i)/113)) * np.cos(2*np.pi*((109*j)/113)). We left this out to avoid digging into a discussion of negative/aliased frequencies, and having this 4th component doesn’t add to our intuition about what the network is doing here.
- at 28:00 we’re not showing removing the 8pi/113 frequency from the model’s final output surface.
Patrons
Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti, Brian Henry, Tim Palade, Petar Vecutin, Nicolas baumann, Jason Singh, Robert Riley, vornska, Barry Silverman, Jake Ehrlich, Mitch Jacobs, Lauren Steely, Jeff Eastman, Rodolfo Ibarra, Clark Barrus, Rob Napier, Andrew White, Richard B Johnston, abhiteja mandava, Burt Humburg, Kevin Mitchell, Daniel Sanchez, Ferdie Wang, Tripp Hill, Richard Harbaugh Jr, Prasad Raje, Kalle Aaltonen, Midori Switch Hound, Zach Wilson, Chris Seltzer, Ven Popov, Hunter Nelson, Amit Bueno, Scott Olsen, Johan Rimez, Shehryar Saroya, Tyler Christensen, Beckett Madden-Woods, Darrell Thomas, Javier Soto, U007D, Caleb Begly, Rick Rubenstein, Brent Hunsaker, Dan Patterson, Tchsurvives, Alex Adai, Walter Reade, Zyansheep, Walter Reade, Duncan Stannett, Reginald Carey, Jean-Manuel Izaret, dh71633, Adrian Rodriguez, Dimitar Stojanovski, Michael Harder, Peter Maldonado, Emily Pesce, David Johnston, Insang Song, FaeTheWolf, Stephen Taylor, KittenKaboodle, EMatter, PATRICKMCCORMACK, John Beahan, Cameron, Cole Jones, Garrett Thornburg, Jeroen W, Rohit Sharma, GlennB, Emmanuel Cortes, Katie Quinn, Karina C, Cakra WW, Mike Ton, Eric Gometz, MacCallister Higgins, Niko Drossos, David Eraso, Tom Zehle, Steve, Brian Lineburg, rjbl, Michael Loh, Perry Vais, Bengal0, Farhad Manjoo, Sara Chipps, Ellis Driscoll, William Taysom, Will Harmon, CK, Abdullah, Peter Cho, Leo Nikora, Griffin Smith, Ash Katnoria, Alex, Markus Hays Nielsen, Catherine H., Vi, David Dobáš, Peter Wang, Sina Sohangir, Danny Thomas, Julian Francis, Hans Adler, Jiayu Peng
Created by: Stephen Welch, Sam Baskin, and Pranav Gundu
CFAQJOTYQHT7JYIT
New Book Available for Preorder Now! The Welch Labs Illustrated Guide to AI (30:47):
welchlabs.com/resources/ai-book-ezrzm
Sections
0:00 - Intro
3:43 - AlexNet & Overfitting
5:19 - Overfitting
6:45 - Rethinking Generalization
11:05 - KiwiCo is Awesome
12:28 - The Double Descent Hypothesis
13:57 - Double Descent is Real!
16:01 - Double Descent with Polynomial Curvefitting?!
20:36 - But why?
22:35 - Should I throw out my books?
24:28 - The Bias-Variance Tradeoff
28:30 - My take
30:47 - I’ve written a new book on AI!
Books with U-shaped test set error curves:
Murphy, Kevin P. Probabilistic machine learning: an introduction. MIT press, 2022.
Goodfellow, Ian, et al. *Deep learning*. Vol. 1. No. 2. Cambridge: MIT press, 2016.
Russell, Stuart Jonathan, and Peter Norvig, eds. *Prentice Hall series in artificial intelligence*. Englewood Cliffs, NJ:: Prentice Hall, 1995.
Bishop, Christopher M., and Nasser M. Nasrabadi. *Pattern recognition and machine learning*. Vol. 4. No. 4. New York: springer, 2006.
Learning, Machine. "Tom mitchell." *Publisher: McGraw Hill* (1997): 31.
Hastie, Trevor, Robert Tibshirani, and Jerome Friedman. "The elements of statistical learning." (2009).
Hastie, Trevor, Robert Tibshirani, and Jerome Friedman. "An introduction to statistical learning." (2009).
Abu-Mostafa, Yaser S., Malik Magdon-Ismail, and Hsuan-Tien Lin. *Learning from data*. Vol. 4. New York: AMLBook, 2012.
MacKay, David JC. *Information theory, inference and learning algorithms*. Cambridge university press, 2003.
Harvard Team’s code & results:
gitlab.com/harvard-machine-learning/double-descent
Great repo showing polynomial double descent:
github.com/RylanSchaeffer/Stanford-AI-Alignment-Double-Descent-Tutorial
Technical Notes
- 26:25 For these linear fits, we’re using N=15 instead of N=5 points. This increases the bias and reduces the variance of these fits, making the bias variance trade-off more clear, but also pushes out the interpolation threshold. Full results are here: github.com/stephencwelch/manim_videos/blob/master/_2025/generalization/Final Video Polynomial Examples.ipynb
- 27:38 It’s tricky to show the full bias-variance results here, as the variance explodes ad Degree=4. Instead we’ve chosen to show qualitative breakdowns, showing which terms dominate the overall error at each degree. Full results can be seen here: github.com/stephencwelch/manim_videos/blob/master/_2025/generalization/Final%20Video%20Polynomial%20Examples.ipynb
Special Thanks to Patrons patreon.com/welchlabs
Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti, Brian Henry, Tim Palade, Petar Vecutin, Nicolas baumann, Jason Singh, Robert Riley, vornska, Barry Silverman, Jake Ehrlich, Mitch Jacobs, Lauren Steely, Jeff Eastman, Rodolfo Ibarra, Clark Barrus, Rob Napier, Andrew White, Richard B Johnston, abhiteja mandava, Burt Humburg, Kevin Mitchell, Daniel Sanchez, Ferdie Wang, Tripp Hill, Richard Harbaugh Jr, Prasad Raje, Kalle Aaltonen, Midori Switch Hound, Zach Wilson, Chris Seltzer, Ven Popov, Hunter Nelson, Amit Bueno, Scott Olsen, Johan Rimez, Shehryar Saroya, Tyler Christensen, Beckett Madden-Woods, Darrell Thomas, Javier Soto, U007D, Caleb Begly, Rick Rubenstein, Brent Hunsaker, Dan Patterson, Tchsurvives, Alex Adai, Walter Reade, Zyansheep, Walter Reade, Duncan Stannett, Reginald Carey, Jean-Manuel Izaret, dh71633, Adrian Rodriguez, Dimitar Stojanovski, Michael Harder, Peter Maldonado, Emily Pesce, David Johnston, Insang Song, FaeTheWolf, Stephen Taylor, KittenKaboodle, EMatter, PATRICKMCCORMACK, John Beahan, Cameron, Cole Jones, Garrett Thornburg, Jeroen W, Rohit Sharma, GlennB, Emmanuel Cortes, Katie Quinn, Karina C, Cakra WW, Mike Ton, Eric Gometz, MacCallister Higgins, Niko Drossos, David Eraso, Tom Zehle, Steve, Brian Lineburg, rjbl, Michael Loh, Perry Vais, Bengal0, Farhad Manjoo, Sara Chipps, Ellis Driscoll, William Taysom, Will Harmon, CK, Abdullah, Peter Cho, Leo Nikora, Griffin Smith, Ash Katnoria, Alex, Markus Hays Nielsen
Special thanks to: Mikhail Belkin, Preetum Nakkiran, Emily Zhang, Varun Reddy
Code for Welch Labs Videos: github.com/stephencwelch/manim_videos
Written by: Stephen Welch
Produced by: Stephen Welch, Sam Baskin, and Pranav Gundu
Premium Beat IDs
EEDYZ3FP44YX8OWT
CFAQJOTYQHT7JYIT
Discount code at checkout: WELCH10
Note that need to buy $15 or more in runpod credits for the discount code to apply, $10 will be deducted from your total. See screen recording at 3:31.
Subliminal Learning Poster at 31:09: welchlabs.com/resources/subliminal-learning-poster-17x22
Subliminal Learning Bundle: welchlabs.com/resources/subliminal-learning-poster-book-bundle
Subliminal Learning Poster - Digital Download: welchlabs.com/resources/subliminal-learning-poster-digital-download
Sections
0:00 - Intro
1:47 - Why Welch Labs uses runpod for AI infrastructure - sponsored ad
3:49 - The subliminal learning phenomenon
5:44 - In context learning
6:56 - Why can’t we just train a classifier?
7:45 - Other clues
9:28 - Small scale replication on MNIST
12:47 - Mathematical proof
23:01 - Proof Take-aways
25:38 - Solving the GPT 4.1/4o mystery
26:14 - My take on what’s going on
27:55 - The token entanglement hypothesis
29:11 - Final thoughts & take-aways
31:09 - Subliminal Learning Poster!
References
Subliminal Learning Paper and code: alignment.anthropic.com/2025/subliminal-learning
Generate Your Own Numbers: https://subliminaldata.streamlit.app/
Token Entanglement: lesswrong.com/posts/m5XzhbZjEuF9uRgGR/it-s-owl-in-the-numbers-token-entanglement-in-subliminal-1
Hinton et. al. 2015. Distilling the Knowledge in a Neural Network. arxiv.org/pdf/1503.02531
Full Video on Backpropagation: youtu.be/VkHfRKewkWw?si=PPONLc5j9Xwlv4Jw
Softmax Basics: youtu.be/VkHfRKewkWw?si=WWPlqu7y1nozl1Fo&t=377
Softmax Gradient: youtu.be/VkHfRKewkWw?si=hd63mRFFIlF3wT-A&t=836
Softmax Visualized: youtu.be/VkHfRKewkWw?si=QZmFau5DjjrFvMso&t=1418
Big thanks to Alex Cloud, Minh Le, Jacob Hilton, and Owain Evans for graciously answering my questions as I worked on the script.
Special Thanks to Patrons patreon.com/welchlabs
Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti, Brian Henry, Tim Palade, Petar Vecutin, Nicolas baumann, Jason Singh, Robert Riley, vornska, Barry Silverman, Jake Ehrlich, Mitch Jacobs, Lauren Steely, Jeff Eastman, Rodolfo Ibarra, Clark Barrus, Rob Napier, Andrew White, Richard B Johnston, abhiteja mandava, Burt Humburg, Kevin Mitchell, Daniel Sanchez, Ferdie Wang, Tripp Hill, Richard Harbaugh Jr, Prasad Raje, Kalle Aaltonen, Midori Switch Hound, Zach Wilson, Chris Seltzer, Ven Popov, Hunter Nelson, Amit Bueno, Scott Olsen, Johan Rimez, Shehryar Saroya, Tyler Christensen, Beckett Madden-Woods, Darrell Thomas, Javier Soto, U007D, Caleb Begly, Rick Rubenstein, Brent Hunsaker, Dan Patterson, Tchsurvives, Alex Adai, Walter Reade, Zyansheep, Walter Reade, Duncan Stannett, Reginald Carey, Jean-Manuel Izaret, dh71633, Adrian Rodriguez, Dimitar Stojanovski, Michael Harder, Peter Maldonado, Emily Pesce, David Johnston, Insang Song, FaeTheWolf, Stephen Taylor, KittenKaboodle, EMatter, PATRICKMCCORMACK, John Beahan, Cameron, Cole Jones, Garrett Thornburg, Jeroen W, Rohit Sharma, GlennB, Emmanuel Cortes, Katie Quinn, Karina C, Cakra WW, Mike Ton, Eric Gometz, MacCallister Higgins, Niko Drossos, David Eraso, Tom Zehle, Steve, Brian Lineburg, rjbl, Michael Loh, Perry Vais, Bengal0, Farhad Manjoo, Sara Chipps
Special thank you to these readers for helping improve the Imaginary Numbers Book!
Marwan Daar, Matt Ellis, Nico Weber, Rafa Barroso, Jacob Sorensen, Bob Hall, Evan Van Peursem, Phillipe Loher, Attila Medl, Abdul Wahid Tanner, A friendly critic, NuttySwiss, Dean Burdick, Paul Du Bois, Włodzimierz Bzyl
Code for Welch Labs Videos: github.com/stephencwelch/manim_videos
Written by: Stephen Welch
Produced by: Stephen Welch, Sam Baskin, and Pranav Gundu
Premium Beat IDs
EEDYZ3FP44YX8OWT
MWROXNAY0SPXCMBS
CFAQJOTYQHT7JYIT
New Patreon Rewards 33:31- own a piece of Welch Labs history!
patreon.com/welchlabs
Books & Posters
welchlabs.com/resources
Sections
0:00 - Intro
4:49 - How Incogni Saves Me Time
6:32 - Part 2 Recap
8:10 - Moving to Two Layers
9:15 - How Activation Functions Fold Space
11:45 - Numerical Walkthrough
13:42 - Universal Approximation Theorem
15:45 - The Geometry of Backpropagation
19:52 - The Geometry of Depth
24:27 - Exponentially Better?
30:23 - Neural Networks Demystifed
31:50 - The Time I Quit YouTube
33:31 - New Patreon Rewards!
Special Thanks to Patrons patreon.com/welchlabs
Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti, Brian Henry, Tim Palade, Petar Vecutin, Nicolas baumann, Jason Singh, Robert Riley, vornska, Barry Silverman, Jake Ehrlich, Mitch Jacobs, Lauren Steely, Jeff Eastman, Rodolfo Ibarra, Clark Barrus, Rob Napier, Andrew White, Richard B Johnston, abhiteja mandava, Burt Humburg, Kevin Mitchell, Daniel Sanchez, Ferdie Wang, Tripp Hill, Richard Harbaugh Jr, Prasad Raje, Kalle Aaltonen, Midori Switch Hound, Zach Wilson, Chris Seltzer, Ven Popov, Hunter Nelson, Amit Bueno, Scott Olsen, Johan Rimez, Shehryar Saroya, Tyler Christensen, Beckett Madden-Woods, Darrell Thomas, Javier Soto
References
Simon Prince, Understanding Deep Learning. udlbook.github.io/udlbook
Liang, Shiyu, and Rayadurgam Srikant. "Why deep neural networks for function approximation?." arXiv preprint arXiv:1610.04161 (2016).
Hanin, Boris, and David Rolnick. "Deep relu networks have surprisingly few activation patterns." *Advances in neural information processing systems* 32 (2019).
Hanin, Boris, and David Rolnick. "Complexity of linear regions in deep networks." *International Conference on Machine Learning*. PMLR, 2019.
Fan, Feng-Lei, et al. "Deep relu networks have surprisingly simple polytopes." *arXiv preprint arXiv:2305.09145* (2023).
All Code:
github.com/stephencwelch/manim_videos
100k neuron wide example training code: github.com/stephencwelch/manim_videos/blob/master/_2025/backprop_3/notebooks/Wide%20Training%20Example.ipynb
Written by: Stephen Welch
Produced by: Stephen Welch, Sam Baskin, and Pranav Gundu
Premium Beat IDs
EEDYZ3FP44YX8OWTe
MWROXNAY0SPXCMBS
CFAQJOTYQHT7JYIT
New Patreon Rewards 29:48 - own a piece of Welch Labs history! patreon.com/welchlabs
Books & Posters
welchlabs.com/resources
Sections
0:00 - Intro
2:08 - No more spam calls w/ Incogni
3:45 - Toy Model
5:20 - y=mx+b
6:17 - Softmax
7:48 - Cross Entropy Loss
9:08 - Computing Gradients
12:31 - Backpropagation
18:23 - Gradient Descent
20:17 - Watching our Model Learn
23:53 - Scaling Up
25:45 - The Map of Language
28:13 - The time I quit YouTube
29:48 - New Patreon Rewards!
Nice Implementation for a viewer in C++:
kirit.com/Tiny%20Classifiers/tiny-classifier.cpp
Special Thanks to Patrons patreon.com/welchlabs
Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti, Brian Henry, Tim Palade, Petar Vecutin, Nicolas baumann, Jason Singh, Robert Riley, vornska, Barry Silverman, Jake Ehrlich, Mitch Jacobs, Lauren Steely
References
Werbos, P. J. (1994). The roots of backpropagation : from ordered derivatives to neural networks and political forecasting. United Kingdom: Wiley. Newton quote is on p4, Werbos expands on the analogy on p4.
Olazaran, Mikel. "A sociological study of the official history of the perceptrons controversy." *Social Studies of Science* 26.3 (1996): 611-659. Minsky quote is on p 393.
Widrow, Bernard. "Generalization and information storage in networks of adaline neurons.” Self-organizing systems (1962): 435-461.
Historical Videos
http://youtube.com/watch?v=FwFduRA_L6Q
youtube.com/watch?v=ntIczNQKfjQ
Code:
github.com/stephencwelch/manim_videos
Technical Notes
Large Llama training animation shows 8/16 layers. Specifically layers 1, 2, 7, 8, 9, 10, 15, and 16. Every third attention pattern is shown, and special tokens are ignored. MLP neurons are downsampled using max pooling. Only the weights and gradients above a specific percentile based threshold are shown. Only query weights are shown going into each attention layer.
The coordinates of Paris are subtracted from all training examples in the 4 city example as a simple normalization - this helps with convergence.
In some scenes, math is happening at higher precision behind the scenes, and results are rounded, which may create apparent inconsistencies.
Written by: Stephen Welch
Produced by: Stephen Welch, Sam Baskin, and Pranav Gundu
Special thanks to: Emily Zhang
Premium Beat IDs
EEDYZ3FP44YX8OWT
MWROXNAY0SPXCMBS
CFAQJOTYQHT7JYIT
Loss Landscape Posters! 21:23
welchlabs.com/resources/loss-landscape-poster-17x19
welchlabs.com/resources/loss-landscape-poster-digital-download
Poster and Book Bundle
welchlabs.com/resources/loss-landscape-bundle-w-imaginary-numbers-book
Special Matte Black Edition Poster
welchlabs.com/resources/loss-landscape-poster-17x22-matte-black-special-edition
Welch Labs Book
welchlabs.com/resources/imaginary-numbers-book
Sections
0:00 - Intro
1:18 - How Incogni gets me more focus time
3:01 - What are we measuring again?
6:18 - How to make our loss go down?
7:32 - Tuning one parameter
9:11 - Tuning two parameters together
11:01 - Gradient descent
13:18 - Visualizing high dimensional surfaces
15:10 - Loss Landscapes
16:55 - Wormholes!
17:55 - Wikitext
18:55 - But where do the wormholes come from?
20:00 - Why local minima are not a problem
21:23 - Posters
Special Thanks to Patrons patreon.com/welchlabs
Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti, Brian Henry, Tim Palade, Petar Vecutin, Nicolas baumann, Jason Singh, Robert Riley, vornska, Barry Silverman, Jake Ehrlich, Mitch Jacobs
References
Li et al: Visualizing the Loss Landscape of Neural Nets. arxiv.org/abs/1712.09913
Talking Nets: An Oral History of Neural Networks. (2000). United Kingdom: MIT Press. Hinton quote is on p376.
Goodfellow, I., Bengio, Y., Courville, A. (2016). Deep Learning. United Kingdom: MIT Press.
Prince, S. J. (2023). Understanding Deep Learning. United Kingdom: MIT Press.
Manim Animations: github.com/stephencwelch/manim_videos
Premium Beat IDs
MWROXNAY0SPXCMBS
CFAQJOTYQHT7JYIT
MLA/DeepSeek Poster at 17:12 (Free shipping for a limited time with code DEEPSEEK):
welchlabs.com/resources/mladeepseek-attention-poster-13x19
Limited edition MLA Poster and Signed Book:
welchlabs.com/resources/deepseek-bundle-mla-poster-and-signed-book-limited-run
Imaginary Numbers book is back in stock!
welchlabs.com/resources/imaginary-numbers-book
Special Thanks to Patrons patreon.com/c/welchlabs
Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti, Brian Henry, Tim Palade, Petar Vecutin, Nicolas baumann, Jason Singh, Robert Riley, vornska, Barry Silverman, Jake Ehrlich
References
DeepSeek-V2 paper: arxiv.org/pdf/2405.04434
DeepSeek-R1 paper: arxiv.org/abs/2501.12948
Great Article by Ege Erdil: epoch.ai/gradient-updates/how-has-deepseek-improved-the-transformer-architecture
GPT-2 Visualizaiton: github.com/TransformerLensOrg/TransformerLens
Manim Animations: github.com/stephencwelch/manim_videos
Technical Notes
1. Note that DeepSeek-V2 paper claims a KV cache size reduction of 93.3%. They don’t exactly publish their methodology, but as far as I can tell it’s something likes this: start with Deepseek-v2 hyperparameters here: huggingface.co/deepseek-ai/DeepSeek-V2/blob/main/configuration_deepseek.py. num_hidden_layers=30, num_attention_heads=32, v_head_dim = 128. If DeepSeek-v2 was implemented with traditional MHA, then KV cache size would be 2*32*128*30*2=491,520 B/token. With MLA with a KV cache size of 576, we get a total cache size of 576*30=34,560 B/token. The percent reduction in KV cache size is then equal to (491,520-34,560)/492,520=92.8%. The numbers I present in this video follow the same approach but are for DeepSeek-v3/R1 architecture: huggingface.co/deepseek-ai/DeepSeek-V3/blob/main/config.json. num_hidden_layers=61, num_attention_heads=128, v_head_dim = 128. So traditional MHA cache would be 2*128*128*61*2 = 3,997,696 B/token. MLA reduces this to 576*61*2=70,272 B/token. Tor the DeepSeek-V3/R1 architecture, MLA reduces the KV cache size by a factor of 3,997,696/70,272 =56.9X.
2. I claim a couple times that MLA allows DeepSeek to generate tokens more than 6x faster than a vanilla transformer. The DeepSeek-V2 paper claims a slightly less than 6x throughput improvement with MLA, but since the V3/R1 architecture is heavier, we expect a larger lift, which is why i claim “more than 6x faster than a vanilla transformer” - in reality it’s probably significantly more than 6x for the V3/R1 architecture.
3. In all attention patterns and walkthroughs, we’re ignoring the |beginning of sentence| token. “The American flag is red, white, and” actually maps to 10 tokens if we include this starting token, and may attention patterns do assign high values to this token.
4. We’re ignoring bias terms matrix equations.
5. We’re ignoring positional embeddings. These are fascinating. See DeepSeek papers and ROPE.
Imaginary Numbers book is back in stock! Update at 23:11
welchlabs.com/resources/imaginary-numbers-book
Welch Labs Posters:
welchlabs.com/resources
Cool Interactive Perceptron Simulator made by viewer Priyangsu Banerjee!
priyangsubanerjee.github.io/perceptron-simulator
Special Thanks to Patrons patreon.com/welchlabs
Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti, Brian Henry, Tim Palade, Petar Vecutin, Nicolas baumann, Jason Singh, Robert Riley, vornska, Barry Silverman
References
Rumelhart, D. E., Mcclelland, J. L. (1987). Parallel Distributed Processing, Volume 1: Explorations in the Microstructure of Cognition: Foundations. United Kingdom: Penguin Random House LLC.
Talking Nets: An Oral History of Neural Networks. (2000). United Kingdom: MIT Press.
Prince, S. J. (2023). Understanding Deep Learning. United Kingdom: MIT Press.
Crevier, D. (1993). AI : the tumultuous history of the search for artificial intelligence. New York: Basic Books.
Cat and dog face dataset: kaggle.com/datasets/andrewmvd/animal-faces?resource=download
Minsky, M., Papert, S. (2017). Perceptrons: An Introduction to Computational Geometry. United Kingdom: MIT Press.
Widrow, Bernard, and Michael A. Lehr. "30 years of adaptive neural networks: perceptron, madaline, and backpropagation." *Proceedings of the IEEE* 78.9 (1990): 1415-1442.
Olazaran, Mikel. "A sociological history of the neural network controversy." *Advances in computers*. Vol. 37. Elsevier, 1993. 335-425.
Widrow, Bernard. "Generalization and information storage in networks of adaline neurons." *Self-organizing systems* (1962): 435-461.
Widrow, Bernard. "Thinking about thinking: the discovery of the LMS algorithm." *IEEE Signal Processing Magazine* 22.1 (2005): 100-106.
Technical Notes
Method for counting neurons in ChatGPT: Starting with GPT-2 implementation here: github.com/karpathy/build-nanogpt/blob/master/train_gpt2.py - keys, queries, and values are implemented in Linear layers with n_embd inputs and 3*n_embd outputs, where n_embd is the embedding dimension. Output projection layer has n_embd and n_embd outputs. So a single attention layer will have ~4*n_embd neurons. GPT-3 has an embedding dimension of 12,288, so each attention layer has ~49,152 neurons. Each MLP block has n_embd inputs, 4*n_embd hidden units, and n_embd outputs, so ~5*n_embd total neurons, or ~61,440. Total neuron count for GPT-3 is then 96*(49,152+61,440)=10,616,832, ignoring initial embedding and final unembedding. Finally, GPT-4 reportedly has ~1.8 Trillion parameters (semianalysis.com/2023/07/10/gpt-4-architecture-infrastructure), making it ~10x larger than GPT-3. Note that GPT-4 is reportedly a mixture of experts, and not all experts are used for each inference, so it appears that not all 1.8 trillion parameters are used for a given inference call. Assuming that ~10x the parameters means 10x the neurons, then GPT-4 should have ~100M neurons.
CFAQJOTYQHT7JYIT
Welch Labs Imaginary Numbers Book!
welchlabs.com/resources/imaginary-numbers-book
Welch Labs Posters:welchlabs.com/resources
Special Thanks to Patrons patreon.com/welchlabs
Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti, Brian Henry, Tim Palade, Petar Vecutin, Nicolas baumann, Jason Singh, Robert Riley, vornska, Barry Silverman
My Gemma walkthrough notebook: colab.research.google.com/drive/1Y68yNr5TcHr4G5RJ0QHZhKkDe55AUkVj?usp=sharing
Most animations made with Manim: github.com/3b1b/manim
References and Further Reading
Chris Olah’s original “Dark Matter of Neural Networks” post: https://transformer-circuits.pub/2024/july-update/index.html#dark-matter
Great recent interview with Chris Olah: youtube.com/watch?v=ugvHCXCOmm4
Gemma Scope: arxiv.org/pdf/2408.05147
Experiment with SAEs yourself here! neuronpedia.org
Relevant work from the Anthropic team:
https://transformer-circuits.pub/2022/toy_model/index.html
https://transformer-circuits.pub/2023/monosemantic-features
https://transformer-circuits.pub/2024/scaling-monosemanticity/
Excellent intro Mechanistic Interpretability: https://arena3-chapter1-transformer-interp.streamlit.app/%5B1.2%5D_Intro_to_Mech_Interp
Neel Nanda’s Mechanistic Interpretability Explainer: dynalist.io/d/n2ZWtnoYHrU1s4vnFSAQ519J
Transformer Lens: github.com/TransformerLensOrg/TransformerLens
SAE Lens: jbloomaus.github.io/SAELens
Technical Notes
1. There are more advanced and more meaningful ways to map mid layer vectors to outputs, see: arxiv.org/pdf/2303.08112, neuralblog.github.io/logit-prisms, lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens
2. The 6x2304 matrix is actually 7x2304, we’re ignoring the /bos token.
3. Gemma also includes positional embeddings and lots and lots of normalization layers, which we didn’t really cover
4. I’m conflating tokens and words sometimes, in this example each word is a token, so we don’t have to worry about it too much
5. The “_” characters represent spaces in the token strings
CFAQJOTYQHT7JYIT
Book Update at 23:28!
Welch Labs Imaginary Numbers Book!
welchlabs.com/resources/imaginary-numbers-book
Welch Labs Posters: welchlabs.com/resources
Huge thanks to Grant Sanderson for a quick manim crash course - I used manim for some of the more complex animations.
Schrodinger equation numerical simulations used for animations were computed with: github.com/quantum-visualizations/qmsolve
Special thanks to Patrons: Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti, Brian Henry, Tim Palade, Petar Vecutin, Nicolas baumann, Jason Singh, Robert Riley
Orbitals Graphics Credits:
Geek3, commons.wikimedia.org/wiki/File:Atomic_orbitals_spdf_m-eigenstates_and_superpositions.png, license: creativecommons.org/licenses/by-sa/4.0/deed.en
PoorLeno, commons.wikimedia.org/wiki/File:Hydrogen_Density_Plots.png
References
De Broglie, L. (1924). *Recherches sur la théorie des quanta* (Doctoral dissertation, Migration-université en cours d'affectation).
Callender, C. (2023). Quantum mechanics: Keeping it real?. *The British Journal for the Philosophy of Science*, *74*(4), 837-851.
Dyson, F. (2009). Birds and frogs. *Notices of the AMS*, *56*(2), 212-223.
Einstein, A. (2011). Letters on Wave Mechanics: Correspondence with H. A. Lorentz, Max Planck, and Erwin Schrödinger. United States: Philosophical Library/Open Road.
Fleisch, D. A. (2020). A Student's Guide to the Schrödinger Equation. India: Cambridge University Press.
Fleisch, D., Kinnaman, L. (2015). A Student's Guide to Waves. United Kingdom: Cambridge University Press.
Gamow, G. (2012). Thirty Years that Shook Physics: The Story of Quantum Theory. United States: Dover Publications.
Jammer, M. (1989). The Conceptual Development of Quantum Mechanics. United Kingdom: Tomash Publishers.
Karam, R. (2020). Schrödinger's original struggles with a complex wave function. *American Journal of Physics*, *88*(6), 433-438.
Moore, W. (2015). Schrödinger: Life and Thought. United Kingdom: Cambridge University Press.
Schrödinger: Centenary Celebration of a Polymath. (1987). United Kingdom: Cambridge University Press.
Schrödinger, E. (2003). Collected Papers on Wave Mechanics: Together with His Four Lectures on Wave Mechanics. United States: AMS Chelsea Pub..
CFAQJOTYQHT7JYIT
Welch Labs Imaginary Numbers Book!
welchlabs.com/resources/imaginary-numbers-book
Welch Labs Posters: welchlabs.com/resources
How the Bizarre Path of Mars Reshaped Astronomy: youtu.be/Phscjl0u6TI
Support Welch Labs on Patreon! patreon.com/welchlabs
Special thanks to Patrons: Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti, Brian Henry, Tim Palade, Petar Vecutin, Nicolas baumann
Learn more about WelchLabs!
welchlabs.com
TikTok: tiktok.com/@welchlabs
Instagram: instagram.com/welchlabs
REFERENCES
Colwell, P. (1993). Solving Kepler's Equation Over Three Centuries. United Kingdom: Willmann-Bell.
Needham, T. (1997). Visual Complex Analysis. United Kingdom: Clarendon Press.
Bate, R. R., Mueller, D. D., White, J. E. (1971). Fundamentals of Astrodynamics. Egypt: Dover Publications.
Vallado, D. (2001). Fundamentals of Astrodynamics and Applications. Netherlands: Springer Netherlands.
Borghi R. On the Bessel Solution of Kepler’s Equation. *Mathematics*. 2024; 12(1):154. doi.org/10.3390/math12010154
Tom Archibald, Craig Fraser, Ivor Grattan-Guinness, The History of Differential Equations, 1670–1950. Oberwolfach Rep. 1 (2004), no. 4, pp. 2729–2794
Francisco G. M. Orlando, C. Farina, Carlos A. D. Zarro, P. Terra; Kepler's equation and some of its pearls. Am. J. Phys. 1 November 2018; 86 (11): 849–858.
Arthur A. Rambaut, M.A. A Simple Method of obtaining an Approximate Solution of Kepler's Equation. *Monthly Notices of the Royal Astronomical Society*, Volume 50, Issue 5, March 1890, Pages 301–302.
Ben Coleman. How to Find the Taylor Series of an Inverse Function. randorithms.com/2021/08/31/Taylor-Series-Inverse.html
17 Recent papers on Kepler’s equation can be found in references 2-16 here: mdpi.com/2227-7390/12/1/154
CFAQJOTYQHT7JYIT
Welch Labs Imaginary Numbers Book!
welchlabs.com/resources/imaginary-numbers-book
Welch Labs Posters: welchlabs.com/resources
Support Welch Labs on Patreon! patreon.com/welchlabs
Special thanks to Patrons: Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti, Brian Henry, Tim Palade, Petar Vecutin
Learn more about WelchLabs! welchlabs.com
TikTok: tiktok.com/@welchlabs
Instagram: instagram.com/welchlabs
REFERENCES
A Neural Scaling Law from the Dimension of the Data Manifold: arxiv.org/pdf/2004.10802
First 2020 OpenAI Scaling Paper: arxiv.org/pdf/2001.08361
GPT-3 Paper: arxiv.org/pdf/2005.14165
Second 202 OpenAI Scaling Paper: arxiv.org/pdf/2010.14701
Google Deepmind “Chinchilla Scaling” Paper: arxiv.org/abs/2203.15556
Nice summary of Chinchilla Scaling: lesswrong.com/posts/6Fpvch8RR29qLEWNH/chinchilla-s-wild-implications
GPT-4 Technical Report: arxiv.org/pdf/2303.08774
Nice Neural Scaling Laws Summary: lesswrong.com/posts/Yt5wAXMc7D2zLpQqx/an-140-theoretical-models-that-predict-scaling-laws
Explaining Neural Scaling Laws: arxiv.org/pdf/2102.06701
High Cost of Training GPT-4: wired.com/story/openai-ceo-sam-altman-the-age-of-giant-ai-models-is-already-over
Nvidia V100 FLOPs: lambdalabs.com/blog/demystifying-gpt-3
Nvidia V100 Original Price: [microway.com/hpc-tech-tips/nvidia-tesla-v100-price-analysis/#:~:text=Tesla GPU model,Key Points](microway.com/hpc-tech-tips/nvidia-tesla-v100-price-analysis/#:~:text=Tesla%20GPU%20model,Key%20Points)
Great paper on scaling up training infrastructure: arxiv.org/pdf/2104.04473
Eight Things to Know about LLMs: arxiv.org/abs/2304.00612
Emergent Properties of LLMs: arxiv.org/abs/2206.07682
Theoretical Motivation for Cross Entropy (Section 6.2): deeplearningbook.org
Some papers that appear to pass the compute efficient frontier
arxiv.org/pdf/2206.14486
arxiv.org/abs/2210.11399
CFAQJOTYQHT7JYIT
Leaked GPT-4 training info
patmcguinness.substack.com/p/gpt-4-details-revealed
semianalysis.com/p/gpt-4-architecture-infrastructure
epochai.org/blog/tracking-large-scale-ai-models
welchlabs.com/resources/imaginary-numbers-book
Book Digital Version
welchlabs.com/resources/imaginary-numbers-are-real-book-digital-download
Euler’s Formula Poster!
welchlabs.com/resources/eulers-formula-poster-13x19
Poster Digital Version
welchlabs.com/resources/eulers-formula-dark-mode-poster-digital-download
Special thanks to the Patrons:
Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti
Tattoo by @themutemaker - thank you! Check out her awesome work here:
instagram.com/themutemaker
Welch Labs
Ad free videos and exclusive perks: patreon.com/welchlabs
Watch on TikTok: tiktok.com/@welchlabs
Learn More or Contact: welchlabs.com
Instagram: instagram.com/welchlabs
X: twitter.com/welchlabs
References & Notes
Welch Labs Imaginary Numbers Series: youtube.com/watch?v=T647CGsuOVU
Excellent History of Logarithms by Florian Cajori
Cajori, Florian. “History of the Exponential and Logarithmic Concepts.” *The American Mathematical Monthly*, vol. 20, no. 1, 1913, pp. 5–14. *JSTOR*, doi.org/10.2307/2973509. Accessed 22 July 2024.
Cajori, Florian. “History of the Exponential and Logarithmic Concepts.” *The American Mathematical Monthly*, vol. 20, no. 2, 1913, pp. 35–47. *JSTOR*, doi.org/10.2307/2974078. Accessed 22 July 2024.
Cajori, Florian. “History of the Exponential and Logarithmic Concepts:” *The American Mathematical Monthly*, vol. 20, no. 3, 1913, pp. 75–84. *JSTOR*, doi.org/10.2307/2973441. Accessed 22 July 2024.
Cajori, Florian. “History of the Exponential and Logarithmic Concepts.” *The American Mathematical Monthly*, vol. 20, no. 4, 1913, pp. 107–17. *JSTOR*, doi.org/10.2307/2972960. Accessed 22 July 2024.
Nice History of Euler’s Formula
Sandifer, Ed. *e, pi and i: Why is “Euler” in the Euler identity?* http://eulerarchive.maa.org/hedi/HEDI-2007-08.pdf
Much of the visual approach presented here comes from Needham’s incredible book:
Needham, T. (1997). Visual Complex Analysis. United Kingdom: Clarendon Press.
Other books referenced
Maor, E. (2011). E: The Story of a Number. Ukraine: Princeton University Press.
Penrose, R. (2021). The Road to Reality: A Complete Guide to the Laws of the Universe. United Kingdom: Knopf Doubleday Publishing Group.
Dunham, W. (2022). Euler: The Master of Us All. United States: AMM Press.
Wilson, R. (2019). Euler's Pioneering Equation: The Most Beautiful Theorem in Mathematics. United Kingdom: Oxford University Press.
Nahin, P. J. (2010). An Imaginary Tale: The Story of √-1. Ukraine: Princeton University Press.
Stillwell, J. (2013). Mathematics and Its History. United Kingdom: Springer New York.
Euler’s Amazing 1747 Paper
Euler, Leonard. *"Sur les logarithmes des nombres négatifs et imaginaires”* Written in 1747, but not published until 1862. Euler did publish a similar paper in 1749. See Cajori #3 above.
English Translation: https://scholarlycommons.pacific.edu/cgi/viewcontent.cgi?filename=0&article=1806&context=euler-works&type=additional
Note on Benroulli’s area of sectors
Euler’s counterexample using Bernoulli’s sector area comes in a couple flavors. The one presented in his 1747 paper "Sur les logarithmes des nombres négatifs et imaginaires” is a bit different than an earlier example in a letter to Bernoulli. I chose the earlier example for clarity. See Cajori vol 2 and Sandifer.
Feynman Lectures - Algebra
https://www.feynmanlectures.caltech.edu/I_22.html
CFAQJOTYQHT7JYIT
Activation Atlas Posters!
welchlabs.com/resources/5gtnaauv6nb9lrhoz9cp604padxp5o
welchlabs.com/resources/activation-atlas-poster-mixed5b-13x19
welchlabs.com/resources/large-activation-atlas-poster-mixed4c-24x36
welchlabs.com/resources/activation-atlas-poster-mixed4c-13x19
Special thanks to the Patrons:
Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti
Welch Labs
Ad free videos and exclusive perks: patreon.com/welchlabs
Watch on TikTok: tiktok.com/@welchlabs
Learn More or Contact: welchlabs.com
Instagram: instagram.com/welchlabs
X: twitter.com/welchlabs
References
AlexNet Paper
proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf
Original Activation Atlas Article- explore here - Great interactive Atlas! https://distill.pub/2019/activation-atlas/
Carter, et al., "Activation Atlas", Distill, 2019.
Feature Visualization Article: https://distill.pub/2017/feature-visualization/
`Olah, et al., "Feature Visualization", Distill, 2017.`
Great LLM Explainability work: https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html
Templeton, et al., "Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet", Transformer Circuits Thread, 2024.
“Deep Visualization Toolbox" by Jason Yosinski video inspired many visuals:
youtube.com/watch?v=AgkfIQ4IGaM
Great LLM/GPT Intro paper
arxiv.org/pdf/2304.10557
3B1Bs GPT Videos are excellent, as always:
youtube.com/watch?v=eMlx5fFNoYc
youtube.com/watch?v=wjZofJX0v4M
Andrej Kerpathy's walkthrough is amazing:
youtube.com/watch?v=kCc8FmEb1nY
Goodfellow’s Deep Learning Book
deeplearningbook.org
OpenAI’s 10,000 V100 GPU cluster (1+ exaflop) news.microsoft.com/source/features/innovation/openai-azure-supercomputer
GPT-3 size, etc: Language Models are Few-Shot Learners, Brown et al, 2020.
Unique token count for ChatGPT: cookbook.openai.com/examples/how_to_count_tokens_with_tiktoken
GPT-4 training size etc, speculative:
patmcguinness.substack.com/p/gpt-4-details-revealed
semianalysis.com/p/gpt-4-architecture-infrastructure
Historical Neural Network Videos
youtube.com/watch?v=FwFduRA_L6Q
youtube.com/watch?v=cNxadbrN_aI
Errata
1:40 should be: "word fragment is appended to the end of the original input". Thanks for Chris A for finding this one.
CFAQJOTYQHT7JYIT
Tycho Brahe Cellarius Poster: welchlabs.com/resources/tycho-brahe-cellarius-poster-13x16
Eclipse Posters! welchlabs.com/resources
This is the second of a two part series - part one here: youtube.com/watch?v=Phscjl0u6TI&feature=youtu.be
Special thanks to the Patrons:
Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti
Welch Labs
Ad free videos and exclusive perks: patreon.com/welchlabs
Watch on TikTok: tiktok.com/@welchlabs
Learn More or Contact: welchlabs.com
Instagram: instagram.com/welchlabs
X: twitter.com/welchlabs
References
On the Shoulders of Giants: The Great Works of Physics and Astronomy. (2003). Kiribati: Penguin.
Koestler, A. (2017). The Sleepwalkers: A History of Man's Changing Vision of the Universe. United Kingdom: Penguin Books Limited.
en.wikipedia.org/wiki/History_of_Mars_observation
keplersdiscovery.com/index.html
The Cambridge Concise History of Astronomy. (1999). United Kingdom: Cambridge University Press.
Mazer, A. (2011). Shifting the Earth: The Mathematical Quest to Understand the Motion of the Universe. Germany: Wiley.
Voelkel, J. R. (2021). The Composition of Kepler's Astronomia Nova. United Kingdom: Princeton University Press.
stellarium.org
Kepler, J. (2015). Astronomia Nova. United States: Green Lion Press.
Stephenson, B. (2012). Kepler’s Physical Astronomy. Switzerland: Springer New York.
Brahe, T., Dreyer, J. L. E. (1972). Tychonis Brahe Dani Opera omnia. Netherlands: Swets & Zeitlinger.
dn790003.ca.archive.org/0/items/Astronomianovaa00Kepl/Astronomianovaa00Kepl.pdf
Feynman Lecture on Gravity: https://www.feynmanlectures.caltech.edu/I_07.html
CFAQJOTYQHT7JYIT


