Neural Networks Are Elastic Origami! [Prof. Randall Balestriero] @MachineLearningStreetTalk
Neural Networks Are Elastic Origami! [Prof. Randall Balestriero]  @MachineLearningStreetTalk
Uploaded February 2025 | Updated September 2026, 1 week ago
Professor Randall Balestriero joins us to discuss neural network geometry, spline theory, and emerging phenomena in deep learning, based on research presented at ICML. Topics include the delayed emergence of adversarial robustness in neural networks ("grokking"), geometric interpretations of neural networks via spline theory, and challenges in reconstruction learning. We also cover geometric analysis of Large Language Models (LLMs) for toxicity detection and the relationship between intrinsic dimensionality and model control in RLHF.

SPONSOR MESSAGES:
***
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments.
centml.ai/pricing

Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. Are you interested in working on reasoning, or getting involved in their events?

Goto tufalabs.ai
***

Show notes and transcript: dropbox.com/scl/fi/3lufge4upq5gy0ug75j4a/RANDALLSHOW.pdf?rlkey=nbemgpa0jhawt1e86rx7372e4&dl=0

TOC:
[00:00:00] Introduction

1. Neural Network Geometry and Spline Theory
[00:01:41] 1.1 Neural Network Geometry and Spline Theory
[00:07:41] 1.2 Deep Networks Always Grok
[00:11:39] 1.3 Grokking and Adversarial Robustness
[00:16:09] 1.4 Double Descent and Catastrophic Forgetting

2. Reconstruction Learning
[00:18:49] 2.1 Reconstruction Learning
[00:24:15] 2.2 Frequency Bias in Neural Networks

3. Geometric Analysis of Neural Networks
[00:29:02] 3.1 Geometric Analysis of Neural Networks
[00:34:41] 3.2 Adversarial Examples and Region Concentration

4. LLM Safety and Geometric Analysis
[00:40:05] 4.1 LLM Safety and Geometric Analysis
[00:46:11] 4.2 Toxicity Detection in LLMs
[00:52:24] 4.3 Intrinsic Dimensionality and Model Control
[00:58:07] 4.4 RLHF and High-Dimensional Spaces

5. Conclusion
[01:02:13] 5.1 Neural Tangent Kernel
[01:08:07] 5.2 Conclusion

REFS:
[00:01:35] Balestriero/Humayun – Deep network geometry & input space partitioning
arxiv.org/html/2408.04809v1

[00:03:55] Balestriero & Paris – Linking deep networks to adaptive spline operators
https://proceedings.mlr.press/v80/balestriero18b/balestriero18b.pdf

[00:13:55] Song et al. – Gradient-based white-box adversarial attacks
arxiv.org/abs/2012.14965

[00:16:05] Humayun, Balestriero & Baraniuk – Grokking phenomenon & emergent robustness
arxiv.org/abs/2402.15555

[00:18:25] Humayun – Training dynamics & double descent via linear region evolution
arxiv.org/abs/2310.12977

[00:20:15] Balestriero – Power diagram partitions in DNN decision boundaries
arxiv.org/abs/1905.08443

[00:23:00] Frankle & Carbin – Lottery Ticket Hypothesis for network pruning
arxiv.org/abs/1803.03635

[00:24:00] Belkin et al. – Double descent phenomenon in modern ML
arxiv.org/abs/1812.11118

[00:25:55] Balestriero et al. – Batch normalization’s regularization effects
arxiv.org/pdf/2209.14778

[00:29:35] EU – EU AI Act 2024 with compute restrictions
lw.com/admin/upload/SiteAttachments/EU-AI-Act-Navigating-a-Brave-New-World.pdf

[00:39:30] Humayun, Balestriero & Baraniuk – SplineCam: Visualizing deep network geometry
openaccess.thecvf.com/content/CVPR2023/papers/Humayun_SplineCam_Exact_Visualization_and_Characterization_of_Deep_Network_Geometry_and_CVPR_2023_paper.pdf

[00:40:40] Carlini – Trade-offs between adversarial robustness and accuracy
arxiv.org/abs/1902.06705

[00:44:55] Balestriero & LeCun – Limitations of reconstruction-based learning methods
raw.githubusercontent.com/mlresearch/v235/main/assets/balestriero24b/balestriero24b.pdf

[00:47:20] Balestriero & LeCun – Spectral analysis of neural network learning
proceedings.neurips.cc/paper_files/paper/2022/file/aa56c74513a5e35768a11f4e82dd7ffb-Paper-Conference.pdf

[00:49:45] He et al. – MAE: Masked Autoencoders for self-supervised learning
arxiv.org/abs/2111.06377

[00:54:50] Balestriero et al. – Geometric analysis of LLM layers for toxicity detection
arxiv.org/abs/2309.12312

[00:59:35] Balestriero et al. – Superior toxicity detection via geometric features
arxiv.org/html/2312.01648v2

[01:04:45] UofT ML – Self-attention control & context length effects
arxiv.org/abs/2310.04444

[01:11:55] Roberts – Foundations of deep learning theory
arxiv.org/abs/2106.10165

[01:15:40] Balestriero & Cha – Kolmogorov GAM Networks via spline partition theory
arxiv.org/pdf/2501.00704

[01:16:40] Various – Graph Kolmogorov-Arnold Networks (GKAN) extension
nature.com/articles/s41598-024-85083-8
Neural Networks Are Elastic Origami! [Prof. Randall Balestriero]Aidan Gomez lessons building CohereCompositionality - Prof. Kevin EllisConnor Leahy - e/acc, AGI and the future.Not real reasoning?The ARC Prize 2024 Winning Algorithm [Daniel Franzen and Jan Disselhoff]Rethinking the Mind - Prof. Mark SolmsPanel discussion on ARC Prize 2024 (Zurich)Dont invent faster horses - Prof. Jeff CluneHow Researchers Test AI for Hidden Goals — Apollo ResearchLanguage Models are Modelling The World [Nicholas Carlini]Cohere is not an AGI company - Nick Frosst (co-founder)
Machine Learning Street Talk |

Neural Networks Are Elastic Origami! [Prof. Randall Balestriero]

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER