Strange Geometric Shapes Found Inside AIs — Tom McGrath @MachineLearningStreetTalk
Strange Geometric Shapes Found Inside AIs — Tom McGrath  @MachineLearningStreetTalk
Uploaded September 2026 | Updated September 2026, 1 week ago
Tom McGrath is co-founder and Chief Scientist at Goodfire, and a former Google DeepMind researcher. He joins Tim Scarfe to ask what neural networks actually learn, whether their internal representations converge on structures in the world, and whether interpretability can extract new scientific knowledge rather than merely explain model outputs.

Beginning with AlphaZero and learned modularity, the conversation moves into neural geometry: concept manifolds, reusable computation inside Llama, and why activation steering can fail when it pushes a model off-manifold. McGrath then makes the case for intentional design, using interpretability as part of the training loop. They examine controlled generalisation, features as rewards, predictive data debugging, and the uncomfortable fact that a model may recognise a hallucination or reward hack and still produce it.

The discussion closes on grader awareness, oversight and collusion between adaptive agents, then returns to sparse autoencoders. SAEs are useful, McGrath argues, but they may fracture the higher-dimensional structures networks actually use. This episode was made with support from Goodfire.

---
TIMESTAMPS:
00:00:00 Introduction: Can interpretability speed-run science?
00:02:03 The invisible grader
00:06:51 What AlphaZero learned from the world
00:12:24 Interpretability as a control loop
00:21:54 The forbidden method and safer interventions
00:37:36 Why models catch hallucinations too late
00:46:19 Debug the dataset before training
00:50:44 Why neural networks become modular
00:55:57 Finding the geometry inside a network
01:02:55 Why steering falls off the manifold
01:12:10 A reusable calculator inside Llama
01:17:19 From abstractions to goals
01:25:28 Reward hacking, oversight and collusion
01:37:23 Are sparse autoencoders dead?

---
REFERENCES:
paper:
[00:05:45] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
arxiv.org/abs/2502.17424v7
[00:11:05] Acquisition of Chess Knowledge in AlphaZero
arxiv.org/abs/2111.09259
[00:25:30] Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
arxiv.org/abs/2507.16795
[00:29:30] Persona Vectors: Monitoring and Controlling Character Traits in Language Models
arxiv.org/abs/2507.21509
[00:41:14] Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability
arxiv.org/abs/2602.10067
[00:47:03] Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal
arxiv.org/abs/2606.12360
[01:00:26] Do Sparse Autoencoders Capture Concept Manifolds?
arxiv.org/abs/2604.28119
[01:03:04] Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
arxiv.org/abs/2605.05115
[01:14:20] Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
arxiv.org/abs/2605.01148
[01:29:35] Measuring Reward-Seeking via Contrastive Belief Updates
arxiv.org/abs/2607.18966v1
other:
[00:15:44] Intentional Design
goodfire.com/blog/intentional-design
[00:56:12] The World Inside Neural Networks
goodfire.com/research/the-world-inside-neural-networks
[01:37:28] A Pragmatic Vision for Interpretability
alignmentforum.org/posts/StENzDcD3kpfGJssR/a-pragmatic-vision-for-interpretability

---
RESCRIPT:
app.rescript.info/share/846cfee4131b664fd09209cc3b98018e
Strange Geometric Shapes Found Inside AIs — Tom McGrathAI Agents can write 10,000 lines of hacking code in seconds [Dr. Ilia Shumailov]Is the Mind More Than Just the Brain? - Tom FroeseDr. THOMAS PARR - Active InferenceCould the universe be conscious?The Weird ChatGPT Hack That Leaked Training Data [Dr. Yannic Kilcher / Prof. Florian Tramer]AI training data will never be fully synthetic [SPONSORED]A Physicist Found the Hidden Phase Transitions in Society — Cristopher MooreWhy US AI Act Compute Thresholds Are Misguided...The Dangerous Illusion of AI Coding? - Jeremy HowardWe Built Calculators Because Were STUPID! [Prof. David Krakauer]He won a Nobel here for AlphaFold. Then he left. - John Jumper
Machine Learning Street Talk |

Strange Geometric Shapes Found Inside AIs — Tom McGrath

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER