Uploaded January 2026 | Updated September 2026, 1 week ago
What makes something truly *intelligent?* Is a rock an agent? Could a perfect simulation of your brain actually *be* you? In this fascinating conversation, Dr. Jeff Beck takes us on a journey through the philosophical and technical foundations of agency, intelligence, and the future of AI.
Jeff doesn't hold back on the big questions. He argues that from a purely mathematical perspective, there's no structural difference between an agent and a rock – both execute policies that map inputs to outputs. The real distinction lies in *sophistication* – how complex are the internal computations? Does the system engage in planning and counterfactual reasoning, or is it just a lookup table that happens to give the right answers?
*Key topics explored in this conversation:*
*The Black Box Problem of Agency* – How can we tell if something is truly planning versus just executing a pre-computed response? Jeff explains why this question is nearly impossible to answer from the outside, and why the best we can do is ask which model gives us the simplest explanation.
*Energy-Based Models Explained* – A masterclass on how EBMs differ from standard neural networks. The key insight: traditional networks only optimize weights, while energy-based models optimize *both* weights and internal states – a subtle but profound distinction that connects to Bayesian inference.
*Why Your Brain Might Have Evolved from Your Nose* – One of the most surprising moments in the conversation. Jeff proposes that the complex, non-smooth nature of olfactory space may have driven the evolution of our associative cortex and planning abilities.
*The JEPA Revolution* – A deep dive into Yann LeCun's Joint Embedding Prediction Architecture and why learning in latent space (rather than predicting every pixel) might be the key to more robust AI representations.
*AI Safety Without Skynet Fears* – Jeff takes a refreshingly grounded stance on AI risk. He's less worried about rogue superintelligences and more concerned about humans becoming "reward function selectors" – couch potatoes who just approve or reject AI outputs. His proposed solution? Use inverse reinforcement learning to derive AI goals from observed human behavior, then make *small* perturbations rather than naive commands like "end world hunger."
Whether you're interested in the philosophy of mind, the technical details of modern machine learning, or just want to understand what makes intelligence *tick,* this conversation delivers insights you won't find anywhere else.
---
TIMESTAMPS:
00:00:00 Geometric Deep Learning & Physical Symmetries
00:00:56 Defining Agency: From Rocks to Planning
00:05:25 The Black Box Problem & Counterfactuals
00:08:45 Simulated Agency vs. Physical Reality
00:12:55 Energy-Based Models & Test-Time Training
00:17:30 Bayesian Inference & Free Energy
00:20:07 JEPA, Latent Space, & Non-Contrastive Learning
00:27:07 Evolution of Intelligence & Modular Brains
00:34:00 Scientific Discovery & Automated Experimentation
00:38:04 AI Safety, Enfeeblement & The Future of Work
---
REFERENCES:
Concept:
[00:00:58] Free Energy Principle (FEP)
en.wikipedia.org/wiki/Free_energy_principle
[00:06:00] Monte Carlo Tree Search
en.wikipedia.org/wiki/Monte_Carlo_tree_search
Book:
[00:09:00] The Intentional Stance
https://mitpress.mit.edu/9780262540537/the-intentional-stance/
Paper:
[00:13:00] A Tutorial on Energy-Based Learning (LeCun 2006)
yann.lecun.com/exdb/publis/pdf/lecun-06.pdf
[00:15:00] Auto-Encoding Variational Bayes (VAE)
arxiv.org/abs/1312.6114
[00:20:15] JEPA (Joint Embedding Prediction Architecture)
openreview.net/forum?id=BZ5a1r-kVsf
[00:22:30] The Wake-Sleep Algorithm
https://www.cs.toronto.edu/~hinton/absps/ws.pdf
[00:22:45] Barlow Twins: Self-Supervised Learning
arxiv.org/abs/2103.03230
[00:30:40] GFlowNets (Generative Flow Networks)
arxiv.org/abs/2111.09266
[00:45:00] Maximum Entropy Inverse Reinforcement Learning
aaai.org/Papers/AAAI/2008/AAAI08-227.pdf
Challenge:
[00:27:15] ARC Prize (Abstraction and Reasoning Corpus)
arcprize.org
---
RESCRIPT:
app.rescript.info/public/share/DJlSbJ_Qx080q315tWaqMWn3PixCQsOcM4Kf1IW9_Eo
PDF:
app.rescript.info/api/public/sessions/0efec296b9b6e905/pdf
What makes something truly *intelligent?* Is a rock an agent? Could a perfect simulation of your brain actually *be* you? In this fascinating conversation, Dr. Jeff Beck takes us on a journey through the philosophical and technical foundations of agency, intelligence, and the future of AI.
Jeff doesn't hold back on the big questions. He argues that from a purely mathematical perspective, there's no structural difference between an agent and a rock – both execute policies that map inputs to outputs. The real distinction lies in *sophistication* – how complex are the internal computations? Does the system engage in planning and counterfactual reasoning, or is it just a lookup table that happens to give the right answers?
*Key topics explored in this conversation:*
*The Black Box Problem of Agency* – How can we tell if something is truly planning versus just executing a pre-computed response? Jeff explains why this question is nearly impossible to answer from the outside, and why the best we can do is ask which model gives us the simplest explanation.
*Energy-Based Models Explained* – A masterclass on how EBMs differ from standard neural networks. The key insight: traditional networks only optimize weights, while energy-based models optimize *both* weights and internal states – a subtle but profound distinction that connects to Bayesian inference.
*Why Your Brain Might Have Evolved from Your Nose* – One of the most surprising moments in the conversation. Jeff proposes that the complex, non-smooth nature of olfactory space may have driven the evolution of our associative cortex and planning abilities.
*The JEPA Revolution* – A deep dive into Yann LeCun's Joint Embedding Prediction Architecture and why learning in latent space (rather than predicting every pixel) might be the key to more robust AI representations.
*AI Safety Without Skynet Fears* – Jeff takes a refreshingly grounded stance on AI risk. He's less worried about rogue superintelligences and more concerned about humans becoming "reward function selectors" – couch potatoes who just approve or reject AI outputs. His proposed solution? Use inverse reinforcement learning to derive AI goals from observed human behavior, then make *small* perturbations rather than naive commands like "end world hunger."
Whether you're interested in the philosophy of mind, the technical details of modern machine learning, or just want to understand what makes intelligence *tick,* this conversation delivers insights you won't find anywhere else.
---
TIMESTAMPS:
00:00:00 Geometric Deep Learning & Physical Symmetries
00:00:56 Defining Agency: From Rocks to Planning
00:05:25 The Black Box Problem & Counterfactuals
00:08:45 Simulated Agency vs. Physical Reality
00:12:55 Energy-Based Models & Test-Time Training
00:17:30 Bayesian Inference & Free Energy
00:20:07 JEPA, Latent Space, & Non-Contrastive Learning
00:27:07 Evolution of Intelligence & Modular Brains
00:34:00 Scientific Discovery & Automated Experimentation
00:38:04 AI Safety, Enfeeblement & The Future of Work
---
REFERENCES:
Concept:
[00:00:58] Free Energy Principle (FEP)
en.wikipedia.org/wiki/Free_energy_principle
[00:06:00] Monte Carlo Tree Search
en.wikipedia.org/wiki/Monte_Carlo_tree_search
Book:
[00:09:00] The Intentional Stance
https://mitpress.mit.edu/9780262540537/the-intentional-stance/
Paper:
[00:13:00] A Tutorial on Energy-Based Learning (LeCun 2006)
yann.lecun.com/exdb/publis/pdf/lecun-06.pdf
[00:15:00] Auto-Encoding Variational Bayes (VAE)
arxiv.org/abs/1312.6114
[00:20:15] JEPA (Joint Embedding Prediction Architecture)
openreview.net/forum?id=BZ5a1r-kVsf
[00:22:30] The Wake-Sleep Algorithm
https://www.cs.toronto.edu/~hinton/absps/ws.pdf
[00:22:45] Barlow Twins: Self-Supervised Learning
arxiv.org/abs/2103.03230
[00:30:40] GFlowNets (Generative Flow Networks)
arxiv.org/abs/2111.09266
[00:45:00] Maximum Entropy Inverse Reinforcement Learning
aaai.org/Papers/AAAI/2008/AAAI08-227.pdf
Challenge:
[00:27:15] ARC Prize (Abstraction and Reasoning Corpus)
arcprize.org
---
RESCRIPT:
app.rescript.info/public/share/DJlSbJ_Qx080q315tWaqMWn3PixCQsOcM4Kf1IW9_Eo
PDF:
app.rescript.info/api/public/sessions/0efec296b9b6e905/pdf

![ARC-AGI-3 winning team - Millennia of minds, compressed into words.
Tim Scarfe travels to Zurich to sit down with the Tufa Labs ARC-AGI-3 team — founder Benjamin Crouzier, with Jeroen Cottaar, Dries Smit, Stefano Viel and Michal Tesnar — to work out what their leaderboard-topping system does and what the benchmark is really testing.
The cut opens on the games: a walkthrough of the Locksmith game, where you read the rules of an unfamiliar world straight from raw frames. ARC-AGI-3 makes ARC interactive and agentic, so the model has to *discover* the goal rather than transduce a static grid. It stays easy for humans and breaks LLMs, and it runs through everything that follows. Dries traces his StochasticGoose preview win — brute force that only searched actions which changed the frame — and why it collapsed once the organisers added action-efficiency scoring and unseen games.
Induction and transduction run through the middle of the conversation — how much of an answer is really priors leaking back the moment a model recognises a maze. The abstraction mountain, and Tims case that LLMs reach the right answer through fractured, entangled representations — performance, not competence. Whether transformers plan at all or just fake it well enough. Why the score really measures action efficiency, not games solved, and why agents lock onto the wrong goal and cannot climb back out.
Crouzier closes on the Tufa Labs thesis — a small lab against the giants, the bitter lesson against hand-built harnesses, and safety — and Tim ties it back to Kenneth Stanley, deep constraints, and creativity as competence.
Disclosure: Tufa Labs sponsors MLST.
TIMESTAMPS:
00:00:00 Meet the Tufa team and what makes ARC-AGI-3 hard
00:02:11 Locksmith game: reading the rules from raw frames
00:03:10 Why build an independent research lab
00:04:11 StochasticGoose: a preview win, then the hardened games
00:07:58 Induction, transduction, and priors inside LLMs
00:10:31 Curiosity, world models, and exploring by frame change
00:14:32 Understanding debt and losing sight of your own code
00:15:53 Requirements-based agents and human-AI co-creativity
00:19:22 Why auto-research misses the big picture
00:21:54 The abstraction mountain and fractured representations
00:27:36 Constraints and making LLMs act as if they understand
00:34:51 Human difficulty calibration, esports priors, and emergence
00:41:35 Agency, goal acquisition, and two kinds of planning
00:47:31 Harnesses, the 36% number, and wrong-goal loops
00:52:33 Rewards, goals, and why ARC-AGI-3 resists brute force
01:00:46 Would solving ARC-AGI-3 prove AGI?
01:07:53 Stripping language away, then priors leak back
01:14:06 Representation and whether language is necessary
01:18:04 The bitter lesson versus specialised harnesses
01:22:20 Capability research, safety, and the software singularity
REFERENCES:
organization:
[00:02:11] ARC-AGI-3
https://arcprize.org/arc-agi/3
[00:03:10] Tufa Labs
https://tufalabs.ai/team/
[00:04:20] ARC-AGI-3 Preview Agent Competition
https://arcprize.org/competitions/arc-agi-3-preview-agents
tool:
[00:04:55] StochasticGoose ARC-AGI-3 solution
https://github.com/DriesSmit/ARC3-solution
[00:07:42] ArcGentica
https://github.com/symbolica-ai/arcgentica
[00:07:49] RGB-Agent
https://github.com/alexisfox7/RGB-Agent
[00:14:38] Claude Code
https://www.anthropic.com/claude-code
[01:03:42] Qwen 3.6 27B
https://huggingface.co/Qwen/Qwen3.6-27B
paper:
[00:13:03] On the Measure of Intelligence
https://arxiv.org/abs/1911.01547
[00:27:42] DreamCoder
https://arxiv.org/abs/2006.08381
[00:43:55] On the Biology of a Large Language Model
https://transformer-circuits.pub/2025/attribution-graphs/biology.html
[01:18:46] ImageNet Classification with Deep CNNs (AlexNet)
https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf
other:
[01:18:16] The Bitter Lesson
http://www.incompleteideas.net/IncIdeas/BitterLesson.html
ReScript:
https://app.rescript.info/public/share/463d7f031349b4b9db428553eed88230 ARC-AGI-3 winning team - Millennia of minds, compressed into words.](https://i.ytimg.com/vi/Vg6FBKTlfOw/mqdefault.jpg)
![AI Interpretability, Safety, and Meaning - Nora Belrose
SPONSOR MESSAGES:
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments.
https://centml.ai/pricing/
Nora Belrose, Head of Interpretability Research at EleutherAI, delivers a wide-ranging conversation that moves from the mathematical foundations of concept erasure in neural networks to fundamental questions about consciousness, AI safety, and Buddhist philosophy.
The technical core centers on LEACE (LEAst-squares Concept Erasure), a method Belrose developed for surgically removing targeted information from neural network representations. She explains how LEACE emerged from connecting two prior approaches (RLACE and spectral attribute removal) through a mathematical equivalence proof, and demonstrates its applications in both fairness-oriented debiasing and interpretability research. A key finding: language models remain functional even after erasing part-of-speech information from every layer, suggesting robust reliance on redundant cues.
Belrose then presents her ICML paper on simplicity biases in deep learning, showing that neural networks learn to exploit statistical moments in order first means, then covariances, then higher-order statistics. This has implications for understanding when and why concept erasure techniques may backfire against sufficiently deep models.
The second half pivots to AI safety, where Belrose delivers a detailed critique of counting arguments used to predict AI misalignment. She argues these arguments rely on the principle of indifference applied to poorly-defined outcome spaces, drawing an analogy to an identical argument structure that would absurdly predict all neural networks must overfit. She connects this to broader questions about goal attribution, agency, and whether instrumental convergence arguments hold up under scrutiny.
The conversation concludes with an exploration of 4E cognition, Evan Thompsons philosophy of mind, Belroses departure from effective altruism, and her growing interest in Buddhist philosophy as a framework for thinking about meaning in a post-automation world.
REFERENCES:
Paper:
[00:00:00] Episode Shownotes
https://www.dropbox.com/scl/fi/38fhsv2zh8gnubtjaoq4a/NORA_FINAL.pdf?rlkey=0e5r8rd261821g1em4dgv0k70&st=t5c9ckfb&dl=0
[00:05:00] LEACE Paper
https://arxiv.org/abs/2306.03819
[00:06:40] RLACE Paper
https://arxiv.org/abs/2201.12091
[00:08:20] Spectral Attribute Removal
https://arxiv.org/abs/2012.14424
[00:15:00] Pythia Models
https://arxiv.org/abs/2304.01373
[00:20:30] LoRA
https://arxiv.org/abs/2106.09685
[02:00:00] Holden Karnofsky
https://forum.effectivealtruism.org/posts/T975ydo3mx4YnRv4J/ea-is-about-maximization-and-maximization-is-perilous
Company:
[00:01:37] CentML
https://centml.ai/pricing/
[00:01:37] Tufa AI Labs
https://tufalabs.ai/
[00:02:20] EleutherAI
https://www.eleuther.ai/
Person:
[00:02:20] Nora Belrose
https://norabelrose.com/
[01:03:00] Evan Thompson
https://evanthompson.me/
LINKS:
Full Transcript: https://app.rescript.info/share/79d69cf24406cc36d8f7e8eee389e3ae
Download PDF transcript: https://app.rescript.info/api/public/sessions/61e64e1737593802/pdf
Nora Belrose:
https://norabelrose.com/
https://scholar.google.com/citations?user=p_oBc64AAAAJ&hl=en
https://x.com/norabelrose AI Interpretability, Safety, and Meaning - Nora Belrose](https://i.ytimg.com/vi/VgPrjHxIS0I/mqdefault.jpg)


![This is what happens when you let AIs debate
Akbir Khan, AI researcher and ICML 2024 Best Paper winner, joins Tim Scarfe to discuss his groundbreaking work on using debate between language models to improve AI truthfulness and oversight. Khan explains how pitting two LLMs against each other in structured arguments helps non-expert judges arrive at more accurate answers than simply querying a single model — a result with profound implications for supervising AI systems that may eventually surpass human capabilities.
The conversation moves through the mechanics of scalable oversight and the sandwiching protocol that operationalises the problem of checking entities smarter than their supervisors. Khan describes how debate naturally surfaces the cruxes of disagreements, making complex expert judgments more accessible to laypeople. The discussion broadens into the relationship between intelligence and agency, the risks of deceptive alignment and reward tampering, and whether the current trajectory of AI development constitutes the early stages of a Cambrian explosion in artificial minds.
Khan and Scarfe also explore open-ended AI systems, Kenneth Stanleys arguments against objective-driven optimization, and the philosophical terrain mapped by thinkers like Aaron Sloman and Francois Chollet on the space of possible minds and measuring intelligence. The episode closes with a frank exchange on cultural evolution, memetics, and whether the intelligence that matters most for AI safety is the kind we can measure.
REFERENCES:
person:
[00:00:00] Akbir Khan
https://akbir.dev/
other:
[00:00:00] MLST Episode Shownotes PDF
https://www.dropbox.com/scl/fi/sjekivbg3ok6qugsv2p1u/AkbirKhan.pdf?rlkey=ewiyvq0aq7mjvql4u7os0jos2&st=vblhp7af&dl=0
[00:06:05] OpenAI Superalignment Team
https://openai.com/index/introducing-superalignment/
[00:19:10] DeepMind Responsible AI
https://deepmind.google/about/responsibility-safety/
paper:
[00:00:40] Akbir Khan et al. - Debating with More Persuasive LLMs
https://arxiv.org/html/2402.06782v3
[00:08:10] Sam Bowman - Scalable Oversight in AI Systems
https://arxiv.org/abs/2211.03540
[00:10:35] Sam Bowman - Artificial Sandwiching Protocol
https://www.alignmentforum.org/posts/nekLYqbCEBDEfbLzF/artificial-sandwiching-when-can-we-test-scalable-alignment
[00:14:35] Janus - Simulators
https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulators
[00:21:30] Eliezer Yudkowsky - AI FOOM Debate
https://intelligence.org/files/AIFoomDebate.pdf
[00:21:45] Sammy Martin - Discontinuous AI Progress
https://www.alignmentforum.org/posts/5WECpYABCT62TJrhY/will-ai-undergo-discontinuous-progress
[00:24:35] Nora Belrose - Counting Arguments vs AI Doom
https://www.lesswrong.com/posts/YsFZF3K9tuzbfrLxo/counting-arguments-provide-no-evidence-for-ai-doom
[00:25:35] Evan Hubinger - Deceptive Alignment
https://www.lesswrong.com/posts/zthDPAjh9w6Ytbeks/deceptive-alignment
[00:26:50] Anthropic - Reward Tampering
https://www.anthropic.com/research/reward-tampering
[00:34:58] Ryan Greenblatt et al. - AI Control
https://arxiv.org/pdf/2312.06942
[00:37:20] Aaron Sloman - The Space of Possible Minds
https://www.cs.bham.ac.uk/research/projects/cogaff/sloman-space-of-minds-84.pdf
[00:38:45] Francois Chollet - On the Measure of Intelligence
https://arxiv.org/abs/1911.01547
[00:42:45] Jonathan Cook et al. - Artificial Generational Intelligence
https://arxiv.org/abs/2406.00392
video:
[00:03:28] Yann LeCun on Machine Learning Debates
https://www.youtube.com/watch?v=OKkEdTchsiE
book:
[00:16:35] Thomas Suddendorf - The Gap
https://www.amazon.in/GAP-Science-Separates-Other-Animals/dp/0465030149
[00:32:35] Kenneth Stanley - Why Greatness Cannot Be Planned
https://www.amazon.co.uk/Why-Greatness-Cannot-Planned-Objective/dp/3319155237
[00:42:30] Richard Dawkins - The Selfish Gene
https://www.amazon.co.uk/Selfish-Gene-Richard-Dawkins/dp/0192860925
LINKS:
Full Transcript: https://app.rescript.info/share/c75b45c237ad52154d276851af8f812b
Download PDF transcript: https://app.rescript.info/api/public/sessions/571abed49dab3e7f/pdf
Akbir Khan:
https://x.com/akbirkhan
https://akbir.dev/ This is what happens when you let AIs debate](https://i.ytimg.com/vi/WlWAhjPfROU/mqdefault.jpg)

![What is “reasoning” in modern AI?
Professor Swarat Chaudhuri from the University of Texas at Austin and visiting researcher at Google DeepMind discusses breakthroughs in AI reasoning, theorem proving, and mathematical discovery. Chaudhuri explains his groundbreaking work on COPRA (a GPT-based prover agent), shares insights on neurosymbolic approaches to AI.
Professor Swarat Chaudhuri:
https://www.cs.utexas.edu/~swarat/
SPONSOR MESSAGES:
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments.
https://centml.ai/pricing/
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on ARC and AGI, they just acquired MindsAI - the current winners of the ARC challenge. Are you interested in working on ARC, or getting involved in their events? Goto https://tufalabs.ai/
TOC:
[00:00:00] 0. Introduction / CentML ad, Tufa ad
1. AI Reasoning: From Language Models to Neurosymbolic Approaches
[00:02:27] 1.1 Defining Reasoning in AI
[00:09:51] 1.2 Limitations of Current Language Models
[00:17:22] 1.3 Neuro-symbolic Approaches and Program Synthesis
[00:24:59] 1.4 COPRA and In-Context Learning for Theorem Proving
[00:34:39] 1.5 Symbolic Regression and LLM-Guided Abstraction
2. AI in Mathematics: Theorem Proving and Concept Discovery
[00:43:37] 2.1 AI-Assisted Theorem Proving and Proof Verification
[01:01:37] 2.2 Symbolic Regression and Concept Discovery in Mathematics
[01:11:57] 2.3 Scaling and Modularizing Mathematical Proofs
[01:21:53] 2.4 COPRA: In-Context Learning for Formal Theorem-Proving
[01:28:22] 2.5 AI-driven theorem proving and mathematical discovery
3. Formal Methods and Challenges in AI Mathematics
[01:30:42] 3.1 Formal proofs, empirical predicates, and uncertainty in AI mathematics
[01:34:01] 3.2 Characteristics of good theoretical computer science research
[01:39:16] 3.3 LLMs in theorem generation and proving
[01:42:21] 3.4 Addressing contamination and concept learning in AI systems
REFS:
00:04:58 The Chinese Room Argument, https://plato.stanford.edu/entries/chinese-room/
00:11:42 Software 2.0, https://medium.com/@karpathy/software-2-0-a64152b37c35
00:11:57 Solving Olympiad Geometry Without Human Demonstrations, https://www.nature.com/articles/s41586-023-06747-5
00:13:26 Lean, https://lean-lang.org/
00:15:43 A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play, https://www.science.org/doi/10.1126/science.aar6404
00:19:24 DreamCoder: Bootstrapping Inductive Program Synthesis with Wake-Sleep Library Learning (Ellis et al., PLDI 2021), https://arxiv.org/abs/2006.08381
00:24:37 The Lambda Calculus, https://plato.stanford.edu/entries/lambda-calculus/
00:26:43 Neural Sketch Learning for Conditional Program Generation, https://arxiv.org/pdf/1703.05698
00:28:08 Learning Differentiable Programs With Admissible Neural Heuristics, https://arxiv.org/abs/2007.12101
00:31:03 Symbolic Regression With a Learned Concept Library (Grayeli et al., NeurIPS 2024), https://arxiv.org/abs/2409.09359
00:41:21 Turing Machines, https://plato.stanford.edu/entries/turing-machine/#HaltProb
00:41:30 Formal Verification of Parallel Programs, https://dl.acm.org/doi/10.1145/360248.360251
01:00:08 The Feynman Lectures, https://www.feynmanlectures.caltech.edu/
01:00:37 Training Compute-Optimal Large Language Models, https://arxiv.org/abs/2203.15556
01:12:26 Fermats Last Theorem, https://en.wikipedia.org/wiki/Fermat%27s_Last_Theorem
01:18:19 Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, https://arxiv.org/abs/2201.11903
01:18:42 Draft, Sketch, and Prove: Guiding Formal Theorem Provers With Informal Proofs, https://arxiv.org/abs/2210.12283
01:19:49 Learning Formal Mathematics From Intrinsic Motivation, https://arxiv.org/pdf/2407.00695
01:20:19 An In-Context Learning Agent for Formal Theorem-Proving (Thakur et al., CoLM 2024), https://arxiv.org/pdf/2310.04353
01:23:58 Learning to Prove Theorems via Interacting With Proof Assistants, https://arxiv.org/abs/1905.09381
01:35:50 Algorithmic Game Theory, https://www.amazon.ca/Algorithmic-Game-Theory-Noam-Nisan/dp/0521872820
01:39:58 An In-Context Learning Agent for Formal Theorem-Proving (Thakur et al., CoLM 2024), https://arxiv.org/pdf/2310.04353
01:42:24 Programmatically Interpretable Reinforcement Learning (Verma et al., ICML 2018), https://arxiv.org/abs/1804.02477 What is “reasoning” in modern AI?](https://i.ytimg.com/vi/XFMk0snybAc/mqdefault.jpg)


![NEURAL NETWORKS ARE WEIRD! - Neel Nanda (DeepMind)
SPONSOR MESSAGES:
***
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments.
https://centml.ai/pricing/
Neel Nanda leads the mechanistic interpretability team at Google DeepMind. At 26, hes become one of the most prominent researchers working on the question of whats actually going on inside neural networks systems that can win IMO medals and write complex software, but which nobody actually designed or understands.
This nearly four-hour conversation is a deep technical dive into the field. Nanda explains why machine learning is fundamentally weird: we produce artifacts that do impressive things, but unlike conventional software, no one wrote the code or planned the architecture. His teams goal is reverse-engineering these systems by finding the internal structures and algorithms that emerge during training.
The discussion covers the mechanics of sparse autoencoders at length how they decompose model activations into interpretable feature vectors, the mathematical foundations (ReLU vs TopK activation functions), scaling laws for feature learning, and the engineering challenges of running them at the scale of frontier models. Nanda walks through the Golden Gate Claude experiment (amplifying a single feature to make Claude obsessed with the Golden Gate Bridge), induction heads (the circuits responsible for in-context learning), and activation patching as a causal intervention technique.
On AI safety, Nanda is pragmatic. He argues that mechanistic interpretability gives us genuine empirical evidence about questions that are otherwise stuck in philosophical debate do models have goals? Do they deceive? He also discusses the limitations: sparse autoencoders havent yet demonstrated capabilities beyond what fine-tuning already achieves, and at sufficient model complexity, models could potentially facade interpretability measurements. The conversation covers his path from pure maths at Cambridge through Anthropic to DeepMind, and why he thinks hands-on coding matters more than reading papers for new researchers entering the field.
REFERENCES:
person:
[00:00:00] Neel Nanda - Personal Website
https://www.neelnanda.io/
tool:
[00:35:00] TransformerLens
https://github.com/TransformerLensOrg/TransformerLens
paper:
[01:00:31] A Mathematical Framework for Transformer Circuits
https://transformer-circuits.pub/2021/framework/index.html
[01:01:40] In-context Learning and Induction Heads
https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html
[01:21:06] Scaling Monosemanticity
https://transformer-circuits.pub/2024/scaling-monosemanticity/
[01:33:27] Refusal in Language Models Is Mediated by a Single Direction
https://arxiv.org/abs/2406.11717
LINKS:
Full Transcript: https://app.rescript.info/share/acb415fa59ae2d2909d60d761c8f4ff4
Download PDF transcript: https://app.rescript.info/api/public/sessions/c3a4bf1e32a46ce7/pdf
NEEL NANDA:
https://www.neelnanda.io/
https://scholar.google.com/citations?user=GLnX3MkAAAAJ&hl=en
https://x.com/NeelNanda5 NEURAL NETWORKS ARE WEIRD! - Neel Nanda (DeepMind)](https://i.ytimg.com/vi/YpFaPKOeNME/mqdefault.jpg)