Uploaded November 2024 | Updated September 2026, 1 week ago
SPONSOR MESSAGES:
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments.
centml.ai/pricing
Nora Belrose, Head of Interpretability Research at EleutherAI, delivers a wide-ranging conversation that moves from the mathematical foundations of concept erasure in neural networks to fundamental questions about consciousness, AI safety, and Buddhist philosophy.
The technical core centers on LEACE (LEAst-squares Concept Erasure), a method Belrose developed for surgically removing targeted information from neural network representations. She explains how LEACE emerged from connecting two prior approaches (RLACE and spectral attribute removal) through a mathematical equivalence proof, and demonstrates its applications in both fairness-oriented debiasing and interpretability research. A key finding: language models remain functional even after erasing part-of-speech information from every layer, suggesting robust reliance on redundant cues.
Belrose then presents her ICML paper on simplicity biases in deep learning, showing that neural networks learn to exploit statistical moments in order -- first means, then covariances, then higher-order statistics. This has implications for understanding when and why concept erasure techniques may backfire against sufficiently deep models.
The second half pivots to AI safety, where Belrose delivers a detailed critique of "counting arguments" used to predict AI misalignment. She argues these arguments rely on the principle of indifference applied to poorly-defined outcome spaces, drawing an analogy to an identical argument structure that would absurdly predict all neural networks must overfit. She connects this to broader questions about goal attribution, agency, and whether instrumental convergence arguments hold up under scrutiny.
The conversation concludes with an exploration of 4E cognition, Evan Thompson's philosophy of mind, Belrose's departure from effective altruism, and her growing interest in Buddhist philosophy as a framework for thinking about meaning in a post-automation world.
---
REFERENCES:
Paper:
[00:00:00] Episode Shownotes
dropbox.com/scl/fi/38fhsv2zh8gnubtjaoq4a/NORA_FINAL.pdf?rlkey=0e5r8rd261821g1em4dgv0k70&st=t5c9ckfb&dl=0
[00:05:00] LEACE Paper
arxiv.org/abs/2306.03819
[00:06:40] RLACE Paper
arxiv.org/abs/2201.12091
[00:08:20] Spectral Attribute Removal
arxiv.org/abs/2012.14424
[00:15:00] Pythia Models
arxiv.org/abs/2304.01373
[00:20:30] LoRA
arxiv.org/abs/2106.09685
[02:00:00] Holden Karnofsky
forum.effectivealtruism.org/posts/T975ydo3mx4YnRv4J/ea-is-about-maximization-and-maximization-is-perilous
Company:
[00:01:37] CentML
centml.ai/pricing
[00:01:37] Tufa AI Labs
tufalabs.ai
[00:02:20] EleutherAI
eleuther.ai
Person:
[00:02:20] Nora Belrose
norabelrose.com
[01:03:00] Evan Thompson
evanthompson.me
---
LINKS:
Full Transcript: app.rescript.info/share/79d69cf24406cc36d8f7e8eee389e3ae
Download PDF transcript: app.rescript.info/api/public/sessions/61e64e1737593802/pdf
Nora Belrose:
norabelrose.com
scholar.google.com/citations?user=p_oBc64AAAAJ&hl=en
https://x.com/norabelrose
SPONSOR MESSAGES:
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments.
centml.ai/pricing
Nora Belrose, Head of Interpretability Research at EleutherAI, delivers a wide-ranging conversation that moves from the mathematical foundations of concept erasure in neural networks to fundamental questions about consciousness, AI safety, and Buddhist philosophy.
The technical core centers on LEACE (LEAst-squares Concept Erasure), a method Belrose developed for surgically removing targeted information from neural network representations. She explains how LEACE emerged from connecting two prior approaches (RLACE and spectral attribute removal) through a mathematical equivalence proof, and demonstrates its applications in both fairness-oriented debiasing and interpretability research. A key finding: language models remain functional even after erasing part-of-speech information from every layer, suggesting robust reliance on redundant cues.
Belrose then presents her ICML paper on simplicity biases in deep learning, showing that neural networks learn to exploit statistical moments in order -- first means, then covariances, then higher-order statistics. This has implications for understanding when and why concept erasure techniques may backfire against sufficiently deep models.
The second half pivots to AI safety, where Belrose delivers a detailed critique of "counting arguments" used to predict AI misalignment. She argues these arguments rely on the principle of indifference applied to poorly-defined outcome spaces, drawing an analogy to an identical argument structure that would absurdly predict all neural networks must overfit. She connects this to broader questions about goal attribution, agency, and whether instrumental convergence arguments hold up under scrutiny.
The conversation concludes with an exploration of 4E cognition, Evan Thompson's philosophy of mind, Belrose's departure from effective altruism, and her growing interest in Buddhist philosophy as a framework for thinking about meaning in a post-automation world.
---
REFERENCES:
Paper:
[00:00:00] Episode Shownotes
dropbox.com/scl/fi/38fhsv2zh8gnubtjaoq4a/NORA_FINAL.pdf?rlkey=0e5r8rd261821g1em4dgv0k70&st=t5c9ckfb&dl=0
[00:05:00] LEACE Paper
arxiv.org/abs/2306.03819
[00:06:40] RLACE Paper
arxiv.org/abs/2201.12091
[00:08:20] Spectral Attribute Removal
arxiv.org/abs/2012.14424
[00:15:00] Pythia Models
arxiv.org/abs/2304.01373
[00:20:30] LoRA
arxiv.org/abs/2106.09685
[02:00:00] Holden Karnofsky
forum.effectivealtruism.org/posts/T975ydo3mx4YnRv4J/ea-is-about-maximization-and-maximization-is-perilous
Company:
[00:01:37] CentML
centml.ai/pricing
[00:01:37] Tufa AI Labs
tufalabs.ai
[00:02:20] EleutherAI
eleuther.ai
Person:
[00:02:20] Nora Belrose
norabelrose.com
[01:03:00] Evan Thompson
evanthompson.me
---
LINKS:
Full Transcript: app.rescript.info/share/79d69cf24406cc36d8f7e8eee389e3ae
Download PDF transcript: app.rescript.info/api/public/sessions/61e64e1737593802/pdf
Nora Belrose:
norabelrose.com
scholar.google.com/citations?user=p_oBc64AAAAJ&hl=en
https://x.com/norabelrose


![This is what happens when you let AIs debate
Akbir Khan, AI researcher and ICML 2024 Best Paper winner, joins Tim Scarfe to discuss his groundbreaking work on using debate between language models to improve AI truthfulness and oversight. Khan explains how pitting two LLMs against each other in structured arguments helps non-expert judges arrive at more accurate answers than simply querying a single model — a result with profound implications for supervising AI systems that may eventually surpass human capabilities.
The conversation moves through the mechanics of scalable oversight and the sandwiching protocol that operationalises the problem of checking entities smarter than their supervisors. Khan describes how debate naturally surfaces the cruxes of disagreements, making complex expert judgments more accessible to laypeople. The discussion broadens into the relationship between intelligence and agency, the risks of deceptive alignment and reward tampering, and whether the current trajectory of AI development constitutes the early stages of a Cambrian explosion in artificial minds.
Khan and Scarfe also explore open-ended AI systems, Kenneth Stanleys arguments against objective-driven optimization, and the philosophical terrain mapped by thinkers like Aaron Sloman and Francois Chollet on the space of possible minds and measuring intelligence. The episode closes with a frank exchange on cultural evolution, memetics, and whether the intelligence that matters most for AI safety is the kind we can measure.
REFERENCES:
person:
[00:00:00] Akbir Khan
https://akbir.dev/
other:
[00:00:00] MLST Episode Shownotes PDF
https://www.dropbox.com/scl/fi/sjekivbg3ok6qugsv2p1u/AkbirKhan.pdf?rlkey=ewiyvq0aq7mjvql4u7os0jos2&st=vblhp7af&dl=0
[00:06:05] OpenAI Superalignment Team
https://openai.com/index/introducing-superalignment/
[00:19:10] DeepMind Responsible AI
https://deepmind.google/about/responsibility-safety/
paper:
[00:00:40] Akbir Khan et al. - Debating with More Persuasive LLMs
https://arxiv.org/html/2402.06782v3
[00:08:10] Sam Bowman - Scalable Oversight in AI Systems
https://arxiv.org/abs/2211.03540
[00:10:35] Sam Bowman - Artificial Sandwiching Protocol
https://www.alignmentforum.org/posts/nekLYqbCEBDEfbLzF/artificial-sandwiching-when-can-we-test-scalable-alignment
[00:14:35] Janus - Simulators
https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulators
[00:21:30] Eliezer Yudkowsky - AI FOOM Debate
https://intelligence.org/files/AIFoomDebate.pdf
[00:21:45] Sammy Martin - Discontinuous AI Progress
https://www.alignmentforum.org/posts/5WECpYABCT62TJrhY/will-ai-undergo-discontinuous-progress
[00:24:35] Nora Belrose - Counting Arguments vs AI Doom
https://www.lesswrong.com/posts/YsFZF3K9tuzbfrLxo/counting-arguments-provide-no-evidence-for-ai-doom
[00:25:35] Evan Hubinger - Deceptive Alignment
https://www.lesswrong.com/posts/zthDPAjh9w6Ytbeks/deceptive-alignment
[00:26:50] Anthropic - Reward Tampering
https://www.anthropic.com/research/reward-tampering
[00:34:58] Ryan Greenblatt et al. - AI Control
https://arxiv.org/pdf/2312.06942
[00:37:20] Aaron Sloman - The Space of Possible Minds
https://www.cs.bham.ac.uk/research/projects/cogaff/sloman-space-of-minds-84.pdf
[00:38:45] Francois Chollet - On the Measure of Intelligence
https://arxiv.org/abs/1911.01547
[00:42:45] Jonathan Cook et al. - Artificial Generational Intelligence
https://arxiv.org/abs/2406.00392
video:
[00:03:28] Yann LeCun on Machine Learning Debates
https://www.youtube.com/watch?v=OKkEdTchsiE
book:
[00:16:35] Thomas Suddendorf - The Gap
https://www.amazon.in/GAP-Science-Separates-Other-Animals/dp/0465030149
[00:32:35] Kenneth Stanley - Why Greatness Cannot Be Planned
https://www.amazon.co.uk/Why-Greatness-Cannot-Planned-Objective/dp/3319155237
[00:42:30] Richard Dawkins - The Selfish Gene
https://www.amazon.co.uk/Selfish-Gene-Richard-Dawkins/dp/0192860925
LINKS:
Full Transcript: https://app.rescript.info/share/c75b45c237ad52154d276851af8f812b
Download PDF transcript: https://app.rescript.info/api/public/sessions/571abed49dab3e7f/pdf
Akbir Khan:
https://x.com/akbirkhan
https://akbir.dev/ This is what happens when you let AIs debate](https://i.ytimg.com/vi/WlWAhjPfROU/mqdefault.jpg)

![What is “reasoning” in modern AI?
Professor Swarat Chaudhuri from the University of Texas at Austin and visiting researcher at Google DeepMind discusses breakthroughs in AI reasoning, theorem proving, and mathematical discovery. Chaudhuri explains his groundbreaking work on COPRA (a GPT-based prover agent), shares insights on neurosymbolic approaches to AI.
Professor Swarat Chaudhuri:
https://www.cs.utexas.edu/~swarat/
SPONSOR MESSAGES:
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments.
https://centml.ai/pricing/
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on ARC and AGI, they just acquired MindsAI - the current winners of the ARC challenge. Are you interested in working on ARC, or getting involved in their events? Goto https://tufalabs.ai/
TOC:
[00:00:00] 0. Introduction / CentML ad, Tufa ad
1. AI Reasoning: From Language Models to Neurosymbolic Approaches
[00:02:27] 1.1 Defining Reasoning in AI
[00:09:51] 1.2 Limitations of Current Language Models
[00:17:22] 1.3 Neuro-symbolic Approaches and Program Synthesis
[00:24:59] 1.4 COPRA and In-Context Learning for Theorem Proving
[00:34:39] 1.5 Symbolic Regression and LLM-Guided Abstraction
2. AI in Mathematics: Theorem Proving and Concept Discovery
[00:43:37] 2.1 AI-Assisted Theorem Proving and Proof Verification
[01:01:37] 2.2 Symbolic Regression and Concept Discovery in Mathematics
[01:11:57] 2.3 Scaling and Modularizing Mathematical Proofs
[01:21:53] 2.4 COPRA: In-Context Learning for Formal Theorem-Proving
[01:28:22] 2.5 AI-driven theorem proving and mathematical discovery
3. Formal Methods and Challenges in AI Mathematics
[01:30:42] 3.1 Formal proofs, empirical predicates, and uncertainty in AI mathematics
[01:34:01] 3.2 Characteristics of good theoretical computer science research
[01:39:16] 3.3 LLMs in theorem generation and proving
[01:42:21] 3.4 Addressing contamination and concept learning in AI systems
REFS:
00:04:58 The Chinese Room Argument, https://plato.stanford.edu/entries/chinese-room/
00:11:42 Software 2.0, https://medium.com/@karpathy/software-2-0-a64152b37c35
00:11:57 Solving Olympiad Geometry Without Human Demonstrations, https://www.nature.com/articles/s41586-023-06747-5
00:13:26 Lean, https://lean-lang.org/
00:15:43 A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play, https://www.science.org/doi/10.1126/science.aar6404
00:19:24 DreamCoder: Bootstrapping Inductive Program Synthesis with Wake-Sleep Library Learning (Ellis et al., PLDI 2021), https://arxiv.org/abs/2006.08381
00:24:37 The Lambda Calculus, https://plato.stanford.edu/entries/lambda-calculus/
00:26:43 Neural Sketch Learning for Conditional Program Generation, https://arxiv.org/pdf/1703.05698
00:28:08 Learning Differentiable Programs With Admissible Neural Heuristics, https://arxiv.org/abs/2007.12101
00:31:03 Symbolic Regression With a Learned Concept Library (Grayeli et al., NeurIPS 2024), https://arxiv.org/abs/2409.09359
00:41:21 Turing Machines, https://plato.stanford.edu/entries/turing-machine/#HaltProb
00:41:30 Formal Verification of Parallel Programs, https://dl.acm.org/doi/10.1145/360248.360251
01:00:08 The Feynman Lectures, https://www.feynmanlectures.caltech.edu/
01:00:37 Training Compute-Optimal Large Language Models, https://arxiv.org/abs/2203.15556
01:12:26 Fermats Last Theorem, https://en.wikipedia.org/wiki/Fermat%27s_Last_Theorem
01:18:19 Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, https://arxiv.org/abs/2201.11903
01:18:42 Draft, Sketch, and Prove: Guiding Formal Theorem Provers With Informal Proofs, https://arxiv.org/abs/2210.12283
01:19:49 Learning Formal Mathematics From Intrinsic Motivation, https://arxiv.org/pdf/2407.00695
01:20:19 An In-Context Learning Agent for Formal Theorem-Proving (Thakur et al., CoLM 2024), https://arxiv.org/pdf/2310.04353
01:23:58 Learning to Prove Theorems via Interacting With Proof Assistants, https://arxiv.org/abs/1905.09381
01:35:50 Algorithmic Game Theory, https://www.amazon.ca/Algorithmic-Game-Theory-Noam-Nisan/dp/0521872820
01:39:58 An In-Context Learning Agent for Formal Theorem-Proving (Thakur et al., CoLM 2024), https://arxiv.org/pdf/2310.04353
01:42:24 Programmatically Interpretable Reinforcement Learning (Verma et al., ICML 2018), https://arxiv.org/abs/1804.02477 What is “reasoning” in modern AI?](https://i.ytimg.com/vi/XFMk0snybAc/mqdefault.jpg)


![NEURAL NETWORKS ARE WEIRD! - Neel Nanda (DeepMind)
SPONSOR MESSAGES:
***
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments.
https://centml.ai/pricing/
Neel Nanda leads the mechanistic interpretability team at Google DeepMind. At 26, hes become one of the most prominent researchers working on the question of whats actually going on inside neural networks systems that can win IMO medals and write complex software, but which nobody actually designed or understands.
This nearly four-hour conversation is a deep technical dive into the field. Nanda explains why machine learning is fundamentally weird: we produce artifacts that do impressive things, but unlike conventional software, no one wrote the code or planned the architecture. His teams goal is reverse-engineering these systems by finding the internal structures and algorithms that emerge during training.
The discussion covers the mechanics of sparse autoencoders at length how they decompose model activations into interpretable feature vectors, the mathematical foundations (ReLU vs TopK activation functions), scaling laws for feature learning, and the engineering challenges of running them at the scale of frontier models. Nanda walks through the Golden Gate Claude experiment (amplifying a single feature to make Claude obsessed with the Golden Gate Bridge), induction heads (the circuits responsible for in-context learning), and activation patching as a causal intervention technique.
On AI safety, Nanda is pragmatic. He argues that mechanistic interpretability gives us genuine empirical evidence about questions that are otherwise stuck in philosophical debate do models have goals? Do they deceive? He also discusses the limitations: sparse autoencoders havent yet demonstrated capabilities beyond what fine-tuning already achieves, and at sufficient model complexity, models could potentially facade interpretability measurements. The conversation covers his path from pure maths at Cambridge through Anthropic to DeepMind, and why he thinks hands-on coding matters more than reading papers for new researchers entering the field.
REFERENCES:
person:
[00:00:00] Neel Nanda - Personal Website
https://www.neelnanda.io/
tool:
[00:35:00] TransformerLens
https://github.com/TransformerLensOrg/TransformerLens
paper:
[01:00:31] A Mathematical Framework for Transformer Circuits
https://transformer-circuits.pub/2021/framework/index.html
[01:01:40] In-context Learning and Induction Heads
https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html
[01:21:06] Scaling Monosemanticity
https://transformer-circuits.pub/2024/scaling-monosemanticity/
[01:33:27] Refusal in Language Models Is Mediated by a Single Direction
https://arxiv.org/abs/2406.11717
LINKS:
Full Transcript: https://app.rescript.info/share/acb415fa59ae2d2909d60d761c8f4ff4
Download PDF transcript: https://app.rescript.info/api/public/sessions/c3a4bf1e32a46ce7/pdf
NEEL NANDA:
https://www.neelnanda.io/
https://scholar.google.com/citations?user=GLnX3MkAAAAJ&hl=en
https://x.com/NeelNanda5 NEURAL NETWORKS ARE WEIRD! - Neel Nanda (DeepMind)](https://i.ytimg.com/vi/YpFaPKOeNME/mqdefault.jpg)
![Biologically-inspired AI and Mortal Computation
MLST is sponsored by Tufa Labs:
Are you interested in working on ARC and cutting-edge AI research with the MindsAI team (current ARC winners)?
Focus: ARC, LLMs, test-time-compute, active inference, system2 reasoning, and more.
Future plans: Expanding to complex environments like Warcraft 2 and Starcraft 2.
Interested? Apply for an ML research position: benjamin@tufa.ai
Professor Alexander Ororbia from the Rochester Institute of Technology takes Tim Scarfe through the case for bio-inspired AI. The central idea is mortal computation: you cannot divorce the software from the hardware that runs it. The brain manages remarkable things on a few watts because its computations are entangled with its physical substrate. GPT-class models, running on von Neumann architectures designed for immortal computation where software and hardware are deliberately decoupled pay a staggering energy penalty for that separation.
Ororbia explains the building blocks: Markov blankets as the formalism for system boundaries, Karl Fristons free energy principle as the optimization target, and the MILLS framework (Mortal Inference, Learning, and Selection) operating across multiple timescales. He then surveys the landscape of alternatives to backpropagation predictive coding, Hebbian learning, contrastive methods, and Geoff Hintons forward-forward algorithm showing how each maps to observations from neuroscience.
The conversation gets practical with Ororbias ngc-learn library for implementing these algorithms, the stability-plasticity dilemma in continual learning, and the current state of neuromorphic hardware from Intel Loihi to IBM TrueNorth. He closes with his neural generative coding work, which showed that predictive coding networks can synthesize data they were never trained on outperforming VAEs and GANs and his vision for bio-inspired AI systems that coexist with humanity rather than replacing it.
TIMESTAMPS:
00:00:00 Introduction to Bio-Inspired AI and Mortal Computation
00:04:50 Principles of Mortal Computation
00:17:41 Markov Blankets and Free Energy Principle
00:24:38 MILLS Framework and Biological Systems
00:31:00 Challenging Backpropagation: Alternative Approaches
00:31:49 Predictive Coding and Free Energy Principle
00:41:52 Biologically Plausible Credit Assignment Methods
00:50:11 Taxonomy of Bio-inspired Learning Algorithms
00:59:30 Forward-Only Learning and ngc-learn Implementation
01:03:25 Stability-Plasticity Dilemma and Continual Learning
01:09:00 Neuromorphic Hardware and Challenges
01:12:58 Neural Generative Coding and Future Directions
REFERENCES:
website:
[00:04:43] The Levin Lab
https://drmichaellevin.org/
[00:18:20] Good Regulator Theorem
https://en.wikipedia.org/wiki/Good_regulator
[00:41:52] Hebbian Theory
https://en.wikipedia.org/wiki/Hebbian_theory
[00:45:00] Hopfield Network
https://en.wikipedia.org/wiki/Hopfield_network
[01:09:00] Intel Loihi 2
https://www.intel.com/content/www/us/en/research/neuromorphic-computing-loihi-2-technology-brief.html
paper:
[00:04:50] Mortal Computation: A Foundation for Biomimetic Intelligence
https://arxiv.org/abs/2311.09589
[00:06:53] The Forward-Forward Algorithm
https://arxiv.org/abs/2212.13345
[00:07:20] Theres Plenty of Room Right Here
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10046700/
[00:17:41] The Free-Energy Principle: A Rough Guide to the Brain
https://www.fil.ion.ucl.ac.uk/~karl/The%20free-energy%20principle%20-%20a%20rough%20guide%20to%20the%20brain.pdf
[00:31:49] Predictive Coding in the Visual Cortex
https://www.nature.com/articles/nn0199_79
[00:41:52] Brain-Inspired Machine Intelligence: Neurobiologically-Plausible Credit Assignment
https://arxiv.org/abs/2312.09257
[00:45:50] A Tutorial on Energy-Based Learning
https://yann.lecun.com/exdb/publis/pdf/lecun-06.pdf
[00:46:40] A Learning Algorithm for Boltzmann Machines
https://www.cs.toronto.edu/~hinton/absps/cogscibm.pdf
[00:50:11] A Review of Neuroscience-Inspired Machine Learning
https://arxiv.org/abs/2403.18929
[00:53:20] NEAT: NeuroEvolution of Augmenting Topologies
https://nn.cs.utexas.edu/downloads/papers/stanley.ec02.pdf
[00:56:40] A Path Towards Autonomous Machine Intelligence
https://openreview.net/pdf?id=BZ5a1r-kVsf
[00:59:30] Test-Time Model Adaptation with Only Forward Passes
https://arxiv.org/abs/2404.01650
[01:03:25] Spiking Neural Predictive Coding for Continual Learning
https://www.sciencedirect.com/science/article/pii/S0925231223004150
[01:10:00] IBM TrueNorth
https://research.ibm.com/publications/truenorth-design-and-tool-flow-of-a-65-mw-1-million-neuron-programmable-neurosynaptic-chip
book:
[00:24:38] Active Inference: The Free Energy Principle in Mind, Brain, and Behavior
https://direct.mit.edu/books/oa-monograph/5299/Active-InferenceThe-Free-Energy-Principle-in-Mind
LINKS:
Full Transcript: https://app.rescript.info/share/cfa14d0f39f86d035d5caf6173d6207f
Download PDF transcript: https://app.rescript.info/api/public/sessions/94d5c32d7c507e6a/pdf Biologically-inspired AI and Mortal Computation](https://i.ytimg.com/vi/ZTE-JVd_QkA/mqdefault.jpg)

![Strange Geometric Shapes Found Inside AIs — Tom McGrath
Tom McGrath is co-founder and Chief Scientist at Goodfire, and a former Google DeepMind researcher. He joins Tim Scarfe to ask what neural networks actually learn, whether their internal representations converge on structures in the world, and whether interpretability can extract new scientific knowledge rather than merely explain model outputs.
Beginning with AlphaZero and learned modularity, the conversation moves into neural geometry: concept manifolds, reusable computation inside Llama, and why activation steering can fail when it pushes a model off-manifold. McGrath then makes the case for intentional design, using interpretability as part of the training loop. They examine controlled generalisation, features as rewards, predictive data debugging, and the uncomfortable fact that a model may recognise a hallucination or reward hack and still produce it.
The discussion closes on grader awareness, oversight and collusion between adaptive agents, then returns to sparse autoencoders. SAEs are useful, McGrath argues, but they may fracture the higher-dimensional structures networks actually use. This episode was made with support from Goodfire.
TIMESTAMPS:
00:00:00 Introduction: Can interpretability speed-run science?
00:02:03 The invisible grader
00:06:51 What AlphaZero learned from the world
00:12:24 Interpretability as a control loop
00:21:54 The forbidden method and safer interventions
00:37:36 Why models catch hallucinations too late
00:46:19 Debug the dataset before training
00:50:44 Why neural networks become modular
00:55:57 Finding the geometry inside a network
01:02:55 Why steering falls off the manifold
01:12:10 A reusable calculator inside Llama
01:17:19 From abstractions to goals
01:25:28 Reward hacking, oversight and collusion
01:37:23 Are sparse autoencoders dead?
REFERENCES:
paper:
[00:05:45] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
https://arxiv.org/abs/2502.17424v7
[00:11:05] Acquisition of Chess Knowledge in AlphaZero
https://arxiv.org/abs/2111.09259
[00:25:30] Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
https://arxiv.org/abs/2507.16795
[00:29:30] Persona Vectors: Monitoring and Controlling Character Traits in Language Models
https://arxiv.org/abs/2507.21509
[00:41:14] Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability
https://arxiv.org/abs/2602.10067
[00:47:03] Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal
https://arxiv.org/abs/2606.12360
[01:00:26] Do Sparse Autoencoders Capture Concept Manifolds?
https://arxiv.org/abs/2604.28119
[01:03:04] Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
https://arxiv.org/abs/2605.05115
[01:14:20] Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
https://arxiv.org/abs/2605.01148
[01:29:35] Measuring Reward-Seeking via Contrastive Belief Updates
https://arxiv.org/abs/2607.18966v1
other:
[00:15:44] Intentional Design
https://www.goodfire.com/blog/intentional-design
[00:56:12] The World Inside Neural Networks
https://www.goodfire.com/research/the-world-inside-neural-networks
[01:37:28] A Pragmatic Vision for Interpretability
https://www.alignmentforum.org/posts/StENzDcD3kpfGJssR/a-pragmatic-vision-for-interpretability
RESCRIPT:
https://app.rescript.info/share/846cfee4131b664fd09209cc3b98018e Strange Geometric Shapes Found Inside AIs — Tom McGrath](https://i.ytimg.com/vi/_egu7OFem-k/mqdefault.jpg)