Uploaded August 2015 | Updated September 2026, 1 week ago
For early access to new videos and other perks:
patreon.com/welchlabs
Want to learn more or teach this series? Check out the Imaginary Numbers are Real Workbook: welchlabs.com/resources.
Imaginary numbers are not some wild invention, they are the deep and natural result of extending our number system. Imaginary numbers are all about the discovery of numbers existing not in one dimension along the number line, but in full two dimensional space. Accepting this not only gives us more rich and complete mathematics, but also unlocks a ridiculous amount of very real, very tangible problems in science and engineering.
Part 1: Introduction
Part 2: A Little History
Part 3: Cardan's Problem
Part 4: Bombelli's Solution
Part 5: Numbers are Two Dimensional
Part 6: The Complex Plane
Part 7: Complex Multiplication
Part 8: Math Wizardry
Part 9: Closure
Part 10: Complex Functions
Part 11: Wandering in Four Dimensions
Part 12: Riemann's Solution
Part 13: Riemann Surfaces
Thanks to viewer "David O" for this correction:
The claim that Euler didn’t know what to do with negative numbers, or thought they were greater than infinity, is a misinterpretation of his On Divergent Series paper. Euler argued that, for infinite series, the word “sum” should mean the original finite expression from which the series originates. For example, applying the geometric series formula "1/(1–r) = 1 + r + r^2 + r^3 + …" with r = 2 gives the formal result "–1 = 1 + 2 + 4 + 8 + …”. However, he did not mean that negative numbers themselves are infinite, nor that such “sums” are equal in the ordinary arithmetic sense for divergent series.
Source: Paragraphs 1-12 of arxiv.org/pdf/1808.02841
Want to learn more or teach this series? Check out the Imaginary Numbers are Real Workbook: welchlabs.com/resources.
For early access to new videos and other perks:
patreon.com/welchlabs
Want to learn more or teach this series? Check out the Imaginary Numbers are Real Workbook: welchlabs.com/resources.
Imaginary numbers are not some wild invention, they are the deep and natural result of extending our number system. Imaginary numbers are all about the discovery of numbers existing not in one dimension along the number line, but in full two dimensional space. Accepting this not only gives us more rich and complete mathematics, but also unlocks a ridiculous amount of very real, very tangible problems in science and engineering.
Part 1: Introduction
Part 2: A Little History
Part 3: Cardan's Problem
Part 4: Bombelli's Solution
Part 5: Numbers are Two Dimensional
Part 6: The Complex Plane
Part 7: Complex Multiplication
Part 8: Math Wizardry
Part 9: Closure
Part 10: Complex Functions
Part 11: Wandering in Four Dimensions
Part 12: Riemann's Solution
Part 13: Riemann Surfaces
Thanks to viewer "David O" for this correction:
The claim that Euler didn’t know what to do with negative numbers, or thought they were greater than infinity, is a misinterpretation of his On Divergent Series paper. Euler argued that, for infinite series, the word “sum” should mean the original finite expression from which the series originates. For example, applying the geometric series formula "1/(1–r) = 1 + r + r^2 + r^3 + …" with r = 2 gives the formal result "–1 = 1 + 2 + 4 + 8 + …”. However, he did not mean that negative numbers themselves are infinite, nor that such “sums” are equal in the ordinary arithmetic sense for divergent series.
Source: Paragraphs 1-12 of arxiv.org/pdf/1808.02841
Want to learn more or teach this series? Check out the Imaginary Numbers are Real Workbook: welchlabs.com/resources.

![The Dark Matter of AI [Mechanistic Interpretability]
Take your personal data back with Incogni! Use code WELCHLABS at the link below and get 60% off an annual plan: http://incogni.com/welchlabs
Welch Labs Imaginary Numbers Book!
https://www.welchlabs.com/resources/imaginary-numbers-book
Welch Labs Posters:https://www.welchlabs.com/resources
Special Thanks to Patrons https://www.patreon.com/welchlabs
Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti, Brian Henry, Tim Palade, Petar Vecutin, Nicolas baumann, Jason Singh, Robert Riley, vornska, Barry Silverman
My Gemma walkthrough notebook: https://colab.research.google.com/drive/1Y68yNr5TcHr4G5RJ0QHZhKkDe55AUkVj?usp=sharing
Most animations made with Manim: https://github.com/3b1b/manim
References and Further Reading
Chris Olah’s original “Dark Matter of Neural Networks” post: https://transformer-circuits.pub/2024/july-update/index.html#dark-matter
Great recent interview with Chris Olah: https://www.youtube.com/watch?v=ugvHCXCOmm4
Gemma Scope: https://arxiv.org/pdf/2408.05147
Experiment with SAEs yourself here! https://www.neuronpedia.org/
Relevant work from the Anthropic team:
https://transformer-circuits.pub/2022/toy_model/index.html
https://transformer-circuits.pub/2023/monosemantic-features
https://transformer-circuits.pub/2024/scaling-monosemanticity/
Excellent intro Mechanistic Interpretability: https://arena3-chapter1-transformer-interp.streamlit.app/%5B1.2%5D_Intro_to_Mech_Interp
Neel Nanda’s Mechanistic Interpretability Explainer: https://dynalist.io/d/n2ZWtnoYHrU1s4vnFSAQ519J
Transformer Lens: https://github.com/TransformerLensOrg/TransformerLens
SAE Lens: https://jbloomaus.github.io/SAELens/
Technical Notes
1. There are more advanced and more meaningful ways to map mid layer vectors to outputs, see: https://arxiv.org/pdf/2303.08112, https://neuralblog.github.io/logit-prisms/, https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens
2. The 6x2304 matrix is actually 7x2304, we’re ignoring the /bos token.
3. Gemma also includes positional embeddings and lots and lots of normalization layers, which we didn’t really cover
4. I’m conflating tokens and words sometimes, in this example each word is a token, so we don’t have to worry about it too much
5. The “_” characters represent spaces in the token strings
CFAQJOTYQHT7JYIT The Dark Matter of AI [Mechanistic Interpretability]](https://i.ytimg.com/vi/UGO_Ehywuxc/mqdefault.jpg)

![Neural Networks Demystified [Part 2: Forward Propagation]
Neural Networks Demystified
@stephencwelch
Supporting Code:
https://github.com/stephencwelch/Neural-Networks-Demystified
In this short series, we will build and train a complete Artificial Neural Network in python. New videos every other friday.
Part 1: Data + Architecture
Part 2: Forward Propagation
Part 3: Gradient Descent
Part 4: Backpropagation
Part 5: Numerical Gradient Checking
Part 6: Training
Part 7: Overfitting, Testing, and Regularization Neural Networks Demystified [Part 2: Forward Propagation]](https://i.ytimg.com/vi/UJwK6jAStmg/mqdefault.jpg)
![Learning to See [Part 8: More Assumptions...Fewer Problems?]
In this series, well explore the complex landscape of machine learning and artificial intelligence through one example from the field of computer vision: using a decision tree to count the number of fingers in an image. Its gonna be crazy.
Supporting Code: https://github.com/stephencwelch/LearningToSee
welchlabs.com
@welchlabs Learning to See [Part 8: More Assumptions...Fewer Problems?]](https://i.ytimg.com/vi/UVwwYZMFocg/mqdefault.jpg)
![The moment we stopped understanding AI [AlexNet]
Thanks to KiwiCo for sponsoring todays video! Go to https://www.kiwico.com/welchlabs and use code WELCHLABS for 50% off your first month of monthly lines and/or for 20% off your first Panda Crate.
Activation Atlas Posters!
https://www.welchlabs.com/resources/5gtnaauv6nb9lrhoz9cp604padxp5o
https://www.welchlabs.com/resources/activation-atlas-poster-mixed5b-13x19
https://www.welchlabs.com/resources/large-activation-atlas-poster-mixed4c-24x36
https://www.welchlabs.com/resources/activation-atlas-poster-mixed4c-13x19
Special thanks to the Patrons:
Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti
Welch Labs
Ad free videos and exclusive perks: https://www.patreon.com/welchlabs
Watch on TikTok: https://www.tiktok.com/@welchlabs
Learn More or Contact: https://www.welchlabs.com/
Instagram: https://www.instagram.com/welchlabs
X: https://twitter.com/welchlabs
References
AlexNet Paper
https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf
Original Activation Atlas Article- explore here - Great interactive Atlas! https://distill.pub/2019/activation-atlas/
Carter, et al., Activation Atlas, Distill, 2019.
Feature Visualization Article: https://distill.pub/2017/feature-visualization/
`Olah, et al., Feature Visualization, Distill, 2017.`
Great LLM Explainability work: https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html
Templeton, et al., Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet, Transformer Circuits Thread, 2024.
“Deep Visualization Toolbox by Jason Yosinski video inspired many visuals:
https://www.youtube.com/watch?v=AgkfIQ4IGaM
Great LLM/GPT Intro paper
https://arxiv.org/pdf/2304.10557
3B1Bs GPT Videos are excellent, as always:
https://www.youtube.com/watch?v=eMlx5fFNoYc
https://www.youtube.com/watch?v=wjZofJX0v4M
Andrej Kerpathys walkthrough is amazing:
https://www.youtube.com/watch?v=kCc8FmEb1nY
Goodfellow’s Deep Learning Book
https://www.deeplearningbook.org/
OpenAI’s 10,000 V100 GPU cluster (1+ exaflop) https://news.microsoft.com/source/features/innovation/openai-azure-supercomputer/
GPT-3 size, etc: Language Models are Few-Shot Learners, Brown et al, 2020.
Unique token count for ChatGPT: https://cookbook.openai.com/examples/how_to_count_tokens_with_tiktoken
GPT-4 training size etc, speculative:
https://patmcguinness.substack.com/p/gpt-4-details-revealed
https://www.semianalysis.com/p/gpt-4-architecture-infrastructure
Historical Neural Network Videos
https://www.youtube.com/watch?v=FwFduRA_L6Q
https://www.youtube.com/watch?v=cNxadbrN_aI
Errata
1:40 should be: word fragment is appended to the end of the original input. Thanks for Chris A for finding this one.
CFAQJOTYQHT7JYIT The moment we stopped understanding AI [AlexNet]](https://i.ytimg.com/vi/UZDiGooFs54/mqdefault.jpg)
![The F=ma of Artificial Intelligence [Backpropagation, How Models Learn Part 2]
Take your personal data back with Incogni! Use code WELCHLABS and get 60% off an annual plan: http://incogni.com/welchlabs
New Patreon Rewards 29:48 - own a piece of Welch Labs history! https://www.patreon.com/welchlabs
Books & Posters
https://www.welchlabs.com/resources
Sections
0:00 - Intro
2:08 - No more spam calls w/ Incogni
3:45 - Toy Model
5:20 - y=mx+b
6:17 - Softmax
7:48 - Cross Entropy Loss
9:08 - Computing Gradients
12:31 - Backpropagation
18:23 - Gradient Descent
20:17 - Watching our Model Learn
23:53 - Scaling Up
25:45 - The Map of Language
28:13 - The time I quit YouTube
29:48 - New Patreon Rewards!
Nice Implementation for a viewer in C++:
https://kirit.com/Tiny%20Classifiers/tiny-classifier.cpp
Special Thanks to Patrons https://www.patreon.com/welchlabs
Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti, Brian Henry, Tim Palade, Petar Vecutin, Nicolas baumann, Jason Singh, Robert Riley, vornska, Barry Silverman, Jake Ehrlich, Mitch Jacobs, Lauren Steely
References
Werbos, P. J. (1994). The roots of backpropagation : from ordered derivatives to neural networks and political forecasting. United Kingdom: Wiley. Newton quote is on p4, Werbos expands on the analogy on p4.
Olazaran, Mikel. A sociological study of the official history of the perceptrons controversy. *Social Studies of Science* 26.3 (1996): 611-659. Minsky quote is on p 393.
Widrow, Bernard. Generalization and information storage in networks of adaline neurons.” Self-organizing systems (1962): 435-461.
Historical Videos
http://youtube.com/watch?v=FwFduRA_L6Q
https://www.youtube.com/watch?v=ntIczNQKfjQ
Code:
https://github.com/stephencwelch/manim_videos
Technical Notes
Large Llama training animation shows 8/16 layers. Specifically layers 1, 2, 7, 8, 9, 10, 15, and 16. Every third attention pattern is shown, and special tokens are ignored. MLP neurons are downsampled using max pooling. Only the weights and gradients above a specific percentile based threshold are shown. Only query weights are shown going into each attention layer.
The coordinates of Paris are subtracted from all training examples in the 4 city example as a simple normalization - this helps with convergence.
In some scenes, math is happening at higher precision behind the scenes, and results are rounded, which may create apparent inconsistencies.
Written by: Stephen Welch
Produced by: Stephen Welch, Sam Baskin, and Pranav Gundu
Special thanks to: Emily Zhang
Premium Beat IDs
EEDYZ3FP44YX8OWT
MWROXNAY0SPXCMBS
CFAQJOTYQHT7JYIT The F=ma of Artificial Intelligence [Backpropagation, How Models Learn Part 2]](https://i.ytimg.com/vi/VkHfRKewkWw/mqdefault.jpg)

![How To Science [Part 4: Science]
PDF: http://www.welchlabs.com/guides
Support Welch Labs: www.patreon.com/welchlabs
Music:
https://www.premiumbeat.com/royalty-free-tracks/jazz-manouche-forever
https://www.premiumbeat.com/royalty-free-tracks/out-of-the-woods
https://www.premiumbeat.com/royalty-free-tracks/truth-to-tell How To Science [Part 4: Science]](https://i.ytimg.com/vi/Y7kCJtWFpUU/mqdefault.jpg)
![Imaginary Numbers Are Real [Part 7: Complex Multiplication]
More information and resources: http://www.welchlabs.com
Imaginary numbers are not some wild invention, they are the deep and natural result of extending our number system. Imaginary numbers are all about the discovery of numbers existing not in one dimension along the number line, but in full two dimensional space. Accepting this not only gives us more rich and complete mathematics, but also unlocks a ridiculous amount of very real, very tangible problems in science and engineering.
Part 1: Introduction
Part 2: A Little History
Part 3: Cardans Problem
Part 4: Bombellis Solution
Part 5: Numbers are Two Dimensional
Part 6: The Complex Plane
Part 7: Complex Multiplication
Part 8: Math Wizardry
Part 9: Closure
Part 10: Complex Functions
Part 11: Wandering in Four Dimensions
Part 12: Riemanns Solution
Part 13: Riemann Surfaces
Want to learn more or teach this series? Check out the Imaginary Numbers are Real Workbook: http://www.welchlabs.com/resources.
Want to learn more or teach this series? Check out the Imaginary Numbers are Real Workbook: http://www.welchlabs.com/resources. Imaginary Numbers Are Real [Part 7: Complex Multiplication]](https://i.ytimg.com/vi/YHvR8siIiD0/mqdefault.jpg)
