Uploaded July 2024 | Updated September 2026, 1 week ago
Prof. Subbarao Kambhampati argues that while LLMs are impressive and useful tools, especially for creative tasks, they have fundamental limitations in logical reasoning and cannot provide guarantees about the correctness of their outputs. He advocates for hybrid approaches that combine LLMs with external verification systems.
MLST is sponsored by Brave:
The Brave Search API covers over 20 billion webpages, built from scratch without Big Tech biases or the recent extortionate price hikes on search API access. Perfect for AI model training and retrieval augmentated generation. Try it now - get 2,000 free queries monthly at brave.com/api.
This is 2/13 of our #ICML2024 series
TOC
[00:00:00] Intro
[00:02:06] Bio
[00:03:02] LLMs are n-gram models on steroids
[00:07:26] Is natural language a formal language?
[00:08:34] Natural language is formal?
[00:11:01] Do LLMs reason?
[00:19:13] Definition of reasoning
[00:31:40] Creativity in reasoning
[00:50:27] Chollet's ARC challenge
[01:01:31] Can we reason without verification?
[01:10:00] LLMs cant solve some tasks
[01:19:07] LLM Modulo framework
[01:29:26] Future trends of architecture
[01:34:48] Future research directions
Pod: podcasters.spotify.com/pod/show/machinelearningstreettalk/episodes/Prof--Subbarao-Kambhampati---LLMs-dont-reason--they-memorize-ICML2024-213-e2mjcse
Subbarao Kambhampati:
https://x.com/rao2z
Interviewer: Dr. Tim Scarfe
Refs:
Can LLMs Really Reason and Plan?
cacm.acm.org/blogcacm/can-llms-really-reason-and-plan
On the Planning Abilities of Large Language Models : A Critical Investigation
arxiv.org/pdf/2305.15771
Chain of Thoughtlessness? An Analysis of CoT in Planning
arxiv.org/pdf/2405.04776
On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks
arxiv.org/pdf/2402.08115
LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
arxiv.org/pdf/2402.01817
Embers of Autoregression: Understanding Large Language
Models Through the Problem They are Trained to Solve
arxiv.org/pdf/2309.13638
arxiv.org/abs/2402.04210
"Task Success" is not Enough
Faith and Fate: Limits of Transformers on Compositionality "finetuning multiplication with four digit numbers" (added after pub)
arxiv.org/pdf/2305.18654
Partition function (number theory) (Srinivasa Ramanujan and G.H. Hardy's work)
en.wikipedia.org/wiki/Partition_function_(number_theory)
Poincaré conjecture
en.wikipedia.org/wiki/Poincar%C3%A9_conjecture
Gödel's incompleteness theorems
en.wikipedia.org/wiki/G%C3%B6del%27s_incompleteness_theorems
ROT13 (Rotate13, "rotate by 13 places")
en.wikipedia.org/wiki/ROT13
A Mathematical Theory of Communication (C. E. SHANNON)
https://people.math.harvard.edu/~ctm/home/text/others/shannon/entropy/entropy.pdf
Sparks of AGI
arxiv.org/abs/2303.12712
Kambhampati thesis on speech recognition (1983)
https://rakaposhi.eas.asu.edu/rao-btech-thesis.pdf
PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change
arxiv.org/abs/2206.10498
Explainable human-AI interaction
link.springer.com/book/10.1007/978-3-031-03767-2
Tree of Thoughts
arxiv.org/abs/2305.10601
On the Measure of Intelligence (ARC Challenge)
arxiv.org/abs/1911.01547
Getting 50% (SoTA) on ARC-AGI with GPT-4o (Ryan Greenblatt ARC solution)
redwoodresearch.substack.com/p/getting-50-sota-on-arc-agi-with-gpt
PROGRAMS WITH COMMON SENSE (John McCarthy) - "AI should be an advice taker program"
https://www.cs.cornell.edu/selman/cs672/readings/mccarthy-upd.pdf
Original chain of thought paper
arxiv.org/abs/2201.11903
ICAPS 2024 Keynote: Dale Schuurmans on "Computing and Planning with Large Generative Models" (COT)
youtube.com/watch?v=YnMqbpdHcaY
The Hardware Lottery (Hooker)
arxiv.org/abs/2009.06489
A Path Towards Autonomous Machine Intelligence (JEPA/LeCun)
openreview.net/pdf?id=BZ5a1r-kVsf
AlphaGeometry
nature.com/articles/s41586-023-06747-5
FunSearch
nature.com/articles/s41586-023-06924-6
Emergent Abilities of Large Language Models
arxiv.org/abs/2206.07682
Language models are not naysayers (Negation in LLMs)
arxiv.org/abs/2306.08189
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
arxiv.org/abs/2309.12288
Embracing negative results
openreview.net/forum?id=3RXAiU7sss
Prof. Subbarao Kambhampati argues that while LLMs are impressive and useful tools, especially for creative tasks, they have fundamental limitations in logical reasoning and cannot provide guarantees about the correctness of their outputs. He advocates for hybrid approaches that combine LLMs with external verification systems.
MLST is sponsored by Brave:
The Brave Search API covers over 20 billion webpages, built from scratch without Big Tech biases or the recent extortionate price hikes on search API access. Perfect for AI model training and retrieval augmentated generation. Try it now - get 2,000 free queries monthly at brave.com/api.
This is 2/13 of our #ICML2024 series
TOC
[00:00:00] Intro
[00:02:06] Bio
[00:03:02] LLMs are n-gram models on steroids
[00:07:26] Is natural language a formal language?
[00:08:34] Natural language is formal?
[00:11:01] Do LLMs reason?
[00:19:13] Definition of reasoning
[00:31:40] Creativity in reasoning
[00:50:27] Chollet's ARC challenge
[01:01:31] Can we reason without verification?
[01:10:00] LLMs cant solve some tasks
[01:19:07] LLM Modulo framework
[01:29:26] Future trends of architecture
[01:34:48] Future research directions
Pod: podcasters.spotify.com/pod/show/machinelearningstreettalk/episodes/Prof--Subbarao-Kambhampati---LLMs-dont-reason--they-memorize-ICML2024-213-e2mjcse
Subbarao Kambhampati:
https://x.com/rao2z
Interviewer: Dr. Tim Scarfe
Refs:
Can LLMs Really Reason and Plan?
cacm.acm.org/blogcacm/can-llms-really-reason-and-plan
On the Planning Abilities of Large Language Models : A Critical Investigation
arxiv.org/pdf/2305.15771
Chain of Thoughtlessness? An Analysis of CoT in Planning
arxiv.org/pdf/2405.04776
On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks
arxiv.org/pdf/2402.08115
LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
arxiv.org/pdf/2402.01817
Embers of Autoregression: Understanding Large Language
Models Through the Problem They are Trained to Solve
arxiv.org/pdf/2309.13638
arxiv.org/abs/2402.04210
"Task Success" is not Enough
Faith and Fate: Limits of Transformers on Compositionality "finetuning multiplication with four digit numbers" (added after pub)
arxiv.org/pdf/2305.18654
Partition function (number theory) (Srinivasa Ramanujan and G.H. Hardy's work)
en.wikipedia.org/wiki/Partition_function_(number_theory)
Poincaré conjecture
en.wikipedia.org/wiki/Poincar%C3%A9_conjecture
Gödel's incompleteness theorems
en.wikipedia.org/wiki/G%C3%B6del%27s_incompleteness_theorems
ROT13 (Rotate13, "rotate by 13 places")
en.wikipedia.org/wiki/ROT13
A Mathematical Theory of Communication (C. E. SHANNON)
https://people.math.harvard.edu/~ctm/home/text/others/shannon/entropy/entropy.pdf
Sparks of AGI
arxiv.org/abs/2303.12712
Kambhampati thesis on speech recognition (1983)
https://rakaposhi.eas.asu.edu/rao-btech-thesis.pdf
PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change
arxiv.org/abs/2206.10498
Explainable human-AI interaction
link.springer.com/book/10.1007/978-3-031-03767-2
Tree of Thoughts
arxiv.org/abs/2305.10601
On the Measure of Intelligence (ARC Challenge)
arxiv.org/abs/1911.01547
Getting 50% (SoTA) on ARC-AGI with GPT-4o (Ryan Greenblatt ARC solution)
redwoodresearch.substack.com/p/getting-50-sota-on-arc-agi-with-gpt
PROGRAMS WITH COMMON SENSE (John McCarthy) - "AI should be an advice taker program"
https://www.cs.cornell.edu/selman/cs672/readings/mccarthy-upd.pdf
Original chain of thought paper
arxiv.org/abs/2201.11903
ICAPS 2024 Keynote: Dale Schuurmans on "Computing and Planning with Large Generative Models" (COT)
youtube.com/watch?v=YnMqbpdHcaY
The Hardware Lottery (Hooker)
arxiv.org/abs/2009.06489
A Path Towards Autonomous Machine Intelligence (JEPA/LeCun)
openreview.net/pdf?id=BZ5a1r-kVsf
AlphaGeometry
nature.com/articles/s41586-023-06747-5
FunSearch
nature.com/articles/s41586-023-06924-6
Emergent Abilities of Large Language Models
arxiv.org/abs/2206.07682
Language models are not naysayers (Negation in LLMs)
arxiv.org/abs/2306.08189
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
arxiv.org/abs/2309.12288
Embracing negative results
openreview.net/forum?id=3RXAiU7sss
![Moving Beyond Surface Statistics (Apple researcher) [Iman Mirzadeh]
Iman Mirzadeh is a machine learning research engineer at Apple and the lead author of the GSM-Symbolic paper, which exposed deep fragility in how large language models handle mathematical reasoning. In this conversation, he draws a sharp line between intelligence and achievement between what a system can score on a benchmark and what it actually understands.
The discussion starts with chess. Mirzadeh explains how grandmasters dont use engines to memorize moves; they use them to develop theory. AlphaZero discovered unprecedented strategies, but that knowledge stays trapped in the game. Humans, by contrast, extract abstract principles like control the center and transfer them to entirely different domains. That capacity for abstraction is what he thinks current AI architecturally lacks.
His critique of LLMs is structural. These systems are trained to minimize cross-entropy loss over a distribution, and by construction they cannot reason beyond what that distribution contains. Change the surface form of a problem swap names, add irrelevant clauses and performance varies wildly, even on grade-school math.
Thats the core finding of GSM-Symbolic. By generating templated variants of math word problems, Mirzadehs team showed that even frontier models exhibit large performance variance from changes that should be semantically irrelevant. The implication: what looks like reasoning is closer to sophisticated pattern matching across memorized distributions.
Mirzadeh proposes that intelligence should be measured by the slope of a systems scaling how fast it can learn novel things rather than its current benchmark position. The conversation also covers the connectionism-symbolism divide, active engagement and agency as necessary conditions for learning, and why we might need a fundamentally different vessel to reach genuine reasoning.
SPONSOR MESSAGES:
***
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich.
TIMESTAMPS:
00:00:00 Intelligence vs Achievement in AI
00:03:27 AlphaZero and Abstract Understanding in Chess
00:10:10 Language Models as Distribution Learners
00:14:47 The State of AI Research Methodology
00:24:24 Interpolation vs True Reasoning in LLMs
00:29:00 Measuring Intelligence: From Chollet to the Iman Moon Test
00:35:35 Agency, Active Learning, and World Models
00:47:15 Scaling Laws and the Connectionism-Symbolism Debate
00:58:09 GSM-Symbolic: Exposing LLM Reasoning Fragility
REFERENCES:
paper:
[00:00:55] Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
https://arxiv.org/abs/1712.01815
[00:17:15] GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
https://arxiv.org/abs/2410.05229
[00:21:20] Connectionism and Cognitive Architecture: A Critical Analysis
https://www.sciencedirect.com/science/article/pii/001002779090014B
[00:29:35] On the Measure of Intelligence
https://arxiv.org/abs/1911.01547
[00:33:25] On definition of intelligence
https://www.sciencedirect.com/science/article/pii/S0160289624000266
[00:35:25] Defining Intelligence
https://cis.temple.edu/~wangp/papers.html
[00:43:10] Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
https://arxiv.org/abs/2201.11903
[00:47:45] Scaling Laws for Neural Language Models
https://arxiv.org/abs/2001.08361
[00:55:10] Tensor Product Variable Binding and the Representation of Symbolic Structures in Connectionist Systems
https://www.sciencedirect.com/science/article/abs/pii/000437029090007M
book:
[00:07:05] Game Changer: AlphaZeros Groundbreaking Chess Strategies
https://www.amazon.com/Game-Changer-AlphaZeros-Groundbreaking-Strategies/dp/9056918184
[00:37:35] How We Learn: Why Brains Learn Better Than Any Machine... for Now
https://www.amazon.com/How-We-Learn-Brains-Machine/dp/0525559884
[00:39:30] Surfaces and Essences: Analogy as the Fuel and Fire of Thinking
https://www.amazon.com/Surfaces-Essences-Analogy-Fuel-Thinking/dp/0465018475
reference:
[00:11:30] NLP Course: Language Modeling
http://lena-voita.github.io/nlp_course/language_modeling.html
[01:08:40] GSM8K: Training Verifiers to Solve Math Word Problems
https://huggingface.co/datasets/openai/gsm8k
LINKS:
Full Transcript: https://app.rescript.info/share/72689965572c7fd460954f70e26b2eaa
Download PDF transcript: https://app.rescript.info/api/public/sessions/abf4f57b59fcbbd5/pdf Moving Beyond Surface Statistics (Apple researcher) [Iman Mirzadeh]](https://i.ytimg.com/vi/yQPduek-Q5s/mqdefault.jpg)

![Why AI Has a Plato Problem — Mazviita Chirimuuta
Professor Mazviita Chirimuuta joins us for a fascinating deep dive into the philosophy of neuroscience and what it really means to understand the mind.
*What can neuroscience actually tell us about how the mind works?* In this thought-provoking conversation, we explore the hidden assumptions behind computational theories of the brain, the limits of scientific abstraction, and why the question of machine consciousness might be more complicated than AI researchers assume.
Mazviita, author of *The Brain Abstracted,* brings a unique perspective shaped by her background in both neuroscience research and philosophy. She challenges us to think critically about the metaphors we use to understand cognition — from the reflex theory of the late 19th century to todays dominant view of the brain as a computer.
*Key topics explored:*
*The problem of oversimplification* — Why scientific models necessarily leave things out, and how this can sometimes lead entire fields astray. The cautionary tale of reflex theory shows how elegant explanations can blind us to biological complexity.
*Is the brain really a computer?* — Mazviita unpacks the philosophical assumptions behind computational neuroscience and asks: if we can model anything computationally, what makes brains special? The answer might challenge everything you thought you knew about AI.
*Haptic realism* — A fresh way of thinking about scientific knowledge that emphasizes interaction over passive observation. Knowledge isnt about reading the source code of the universe — its something we actively construct through engagement with the world.
*Why embodiment matters for understanding* — Can a disembodied language model truly understand? Mazviita makes a compelling case that human cognition is deeply entangled with our sensory-motor engagement and biological existence in ways that cant simply be abstracted away.
*Technology and human finitude* — Drawing on Heidegger, we discuss how the dream of transcending our physical limitations through technology might reflect a fundamental misunderstanding of what it means to be a knower.
This conversation is essential viewing for anyone interested in AI, consciousness, philosophy of mind, or the future of cognitive science. Whether youre skeptical of strong AI claims or a true believer in machine consciousness, Mazviitas careful philosophical analysis will give you new tools for thinking through these profound questions.
TIMESTAMPS:
00:00:00 The Problem of Generalizing Neuroscience
00:02:51 Abstraction vs. Idealization: The Kaleidoscope
00:05:39 Platonism in AI: Discovering or Inventing Patterns?
00:09:42 When Simplification Fails: The Reflex Theory
00:12:23 Behaviorism and the Black Box Trap
00:14:20 Haptic Realism: Knowledge Through Interaction
00:20:23 Is Nature Protean? The Myth of Converging Truth
00:23:23 The Computational Theory of Mind: A Useful Fiction?
00:27:25 Biological Constraints: Why Brains Arent Just Neural Nets
00:31:01 Agency, Distal Causes, and Dennetts Stances
00:37:13 Searles Challenge: Causal Powers and Understanding
00:41:58 Heideggers Warning & The Experiment on Children
REFERENCES:
Book:
[00:01:28] The Brain Abstracted
https://mitpress.mit.edu/9780262548045/the-brain-abstracted/
[00:11:05] The Integrated Action of the Nervous System
https://www.amazon.sg/integrative-action-nervous-system/dp/9354179029
[00:18:15] The Quest for Certainty (Dewey)
https://www.amazon.com/Quest-Certainty-Relation-Knowledge-Lectures/dp/0399501916
[00:19:45] Realism for Realistic People (Chang)
https://www.cambridge.org/core/books/realism-for-realistic-people/ACC93A7F03B15AA4D6F3A466E3FC5AB7
[00:38:15] The Rediscovery of the Mind (Searle)
https://mitpress.mit.edu/9780262691543/the-rediscovery-of-the-mind/
[00:47:18] So Youve Been Publicly Shamed (Ronson)
https://www.amazon.com/So-Youve-Been-Publicly-Shamed/dp/1594634017
[00:50:30] Reality+ (Chalmers)
https://consc.net/reality/
Person:
[00:05:00] Francois Chollet
https://arcprize.org/
Paper:
[00:08:03] Real Patterns (Dennett)
https://ruccs.rutgers.edu/images/personal-zenon-pylyshyn/class-info/FP2012/FP2012_readings/Dennett_RealPatterns.pdf
[00:25:30] A Logical Calculus of Ideas... (1943)
https://link.springer.com/article/10.1007/BF02478259
[00:29:17] The Lottery Ticket Hypothesis
https://arxiv.org/abs/1803.03635
Philosophy:
[00:17:30] Transcendental Idealism (Kant)
https://plato.stanford.edu/entries/kant-transcendental-idealism/
RESCRIPT:
https://app.rescript.info/public/share/A6cZ1TY35p8ORMmYCWNBI0no9ChU3-Kx7dPXGJURvZ0
PDF Transcript:
https://app.rescript.info/api/public/sessions/0fb7767e066cf712/pdf Why AI Has a Plato Problem — Mazviita Chirimuuta](https://i.ytimg.com/vi/yq318DIwPqw/mqdefault.jpg)

![Solving Chollets ARC-AGI with GPT4o
Ryan Greenblatt from Redwood Research recently published Getting 50% on ARC-AGI with GPT-4.0, where he used GPT4o to reach a state-of-the-art accuracy on Francois Chollets ARC Challenge by generating many Python programs.
Sponsor:
Sign up to Kalshi here https://kalshi.onelink.me/1r91/mlst the first 500 traders who deposit $100 will get a free $20 credit! Important disclaimer - In case its not obvious - this is basically gambling and a *high risk* activity - only trade what you can afford to lose.
We discuss:
- Ryans unique approach to solving the ARC Challenge and achieving impressive results.
- The strengths and weaknesses of current AI models.
- How AI and humans differ in learning and reasoning.
- Combining various techniques to create smarter AI systems.
- The potential risks and future advancements in AI, including the idea of agentic AI.
https://x.com/RyanPGreenblatt
https://www.redwoodresearch.org/
TOC
00:00:00 Intro
00:01:38 Prelude on goals in LLMs
00:02:42 Ryan intro
00:03:11 Ryans ARC Challenge Approach
00:38:15 Language models, reasoning and agency
01:14:14 Timelines on superintelligence
01:27:05 Growth of superintelligence
02:06:41 Reflections on ARC
02:11:49 Why wouldnt AI knowledge be subjective
Host: Dr. Tim Scarfe
Pod: https://podcasters.spotify.com/pod/show/machinelearningstreettalk/episodes/Ryan-Greenblatt Solving-ARC-with-GPT4o-e2lnplq
Refs:
Getting 50% (SoTA) on ARC-AGI with GPT-4o [Ryan Greenblatt]
https://redwoodresearch.substack.com/p/getting-50-sota-on-arc-agi-with-gpt
On the Measure of Intelligence [Chollet]
https://arxiv.org/abs/1911.01547
Connectionism and Cognitive Architecture: A Critical Analysis [Jerry A. Fodor and Zenon W. Pylyshyn]
https://ruccs.rutgers.edu/images/personal-zenon-pylyshyn/proseminars/Proseminar13/ConnectionistArchitecture.pdf
Software 2.0 [Andrej Karpathy]
https://karpathy.medium.com/software-2-0-a64152b37c35
Why Greatness Cannot Be Planned: The Myth of the Objective [Kenneth Stanley]
https://amzn.to/3Wfy2E0
Biographical account of Terence Tao’s mathematical development. [M.A.(KEN) CLEMENTS]
https://gwern.net/doc/iq/high/smpy/1984-clements.pdf
Model Evaluation and Threat Research (METR)
https://metr.org/
Why Tool AIs Want to Be Agent AIs
https://gwern.net/tool-ai
Simulators - Janus
https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulators
AI Control: Improving Safety Despite Intentional Subversion
https://www.lesswrong.com/posts/d9FJHawgkiMSPjagR/ai-control-improving-safety-despite-intentional-subversion
https://arxiv.org/abs/2312.06942
What a Compute-Centric Framework Says About Takeoff Speeds
https://www.openphilanthropy.org/research/what-a-compute-centric-framework-says-about-takeoff-speeds/
Global GDP over the long run
https://ourworldindata.org/grapher/global-gdp-over-the-long-run?yScale=log
Safety Cases: How to Justify the Safety of Advanced AI Systems
https://arxiv.org/abs/2403.10462
The Danger of a “Safety Case
http://sunnyday.mit.edu/The-Danger-of-a-Safety-Case.pdf
The Future Of Work Looks Like A UPS Truck (~02:15:50)
https://www.npr.org/sections/money/2014/05/02/308640135/episode-536-the-future-of-work-looks-like-a-ups-truck
SWE-bench
https://www.swebench.com/
Using DeepSpeed and Megatron to Train Megatron-Turing NLG
530B, A Large-Scale Generative Language Model
https://arxiv.org/pdf/2201.11990
Algorithmic Progress in Language Models
https://epochai.org/blog/algorithmic-progress-in-language-models Solving Chollets ARC-AGI with GPT4o](https://i.ytimg.com/vi/z9j3wB1RRGA/mqdefault.jpg)




![There are monsters in your LLM. (Murray Shanahan)
Murray Shanahan — professor of Cognitive Robotics at Imperial College London and senior research scientist at DeepMind — challenges how we think and talk about machine intelligence. He argues that the biggest danger of anthropomorphism is not emotional attachment but systematic misattribution of capabilities, in both directions.
The conversation spans simulators and simulacra (drawing on Janus from LessWrong), the shoggoth theory of what lies behind the RLHF mask, Wittgensteins private language argument applied to AI consciousness, and whether concepts like agency and embodiment are necessary for genuine understanding. Murray draws on his work as scientific advisor to Ex Machina and his papers on conscious exotica to articulate why our existing vocabulary for consciousness is inadequate for these new entities.
This is a philosophically rigorous two-hour conversation that never loses sight of the engineering reality — covering everything from the Turing test to Nagels bat to the ARC challenge, with Murray consistently pushing back on easy answers.
TIMESTAMPS:
00:00:00 Intro
00:05:49 Simulators and simulacra
00:11:04 The 20 questions game and simulacra stickiness
00:18:50 Murrays experience with Claude 3
00:30:04 RLHF and alignment
00:32:41 Anthropic Golden Gate Bridge experiment
00:37:05 Agency in language models
00:41:05 Embodiment and knowledge acquisition
00:57:51 ARC challenge and abstract reasoning
01:03:31 The conscious stance
01:13:58 Space of possible minds
01:17:45 Wittgenstein private language and subjectivity
01:29:58 Conscious exotica
01:33:23 Dennett and the intentional stance
01:40:58 Anthropomorphisation risks
01:46:47 Reasoning in language models
01:53:56 The Turing test revisited
02:04:41 Nagels bat and subjective experience
02:08:08 Mark Bishop idealism and Chinese Room
02:09:32 Panpsychism and consciousness
REFERENCES:
book:
[00:00:00] The Technological Singularity
https://www.doc.ic.ac.uk/~mpsha/
[00:41:05] Embodiment and the Inner Life
https://www.doc.ic.ac.uk/~mpsha/
[01:17:45] Philosophical Investigations
https://en.wikipedia.org/wiki/Philosophical_Investigations
[01:17:45] The Language Game
https://en.wikipedia.org/wiki/The_Language_Game
[01:33:23] The intentional stance
https://en.wikipedia.org/wiki/Intentional_stance
[01:40:58] Metaphors We Live By
https://en.wikipedia.org/wiki/Metaphors_We_Live_By
website:
[00:00:00] Ex Machina
https://en.wikipedia.org/wiki/Ex_Machina_(film)
[00:57:51] ARC-AGI Challenge
https://github.com/fchollet/ARC-AGI
person:
[00:00:00] Murray Shanahan academic page
https://www.doc.ic.ac.uk/~mpsha/
article:
[00:05:49] Simulators article
https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulators
[01:13:58] Space of possible minds
https://en.wikipedia.org/wiki/Aaron_Sloman
[02:08:08] Chinese Room Argument
https://en.wikipedia.org/wiki/Chinese_room
paper:
[00:05:49] Role play with large language models
https://arxiv.org/abs/2305.16367
[00:32:41] Scaling Monosemanticity
https://transformer-circuits.pub/2024/scaling-monosemanticity/
[01:29:58] Conscious Exotica
https://www.doc.ic.ac.uk/~mpsha/
[02:04:41] What Is It Like to Be a Bat
https://en.wikipedia.org/wiki/What_Is_It_Like_to_Be_a_Bat%3F
LINKS:
Full Transcript: https://app.rescript.info/share/a752f88b40bf0658f4e8608bd9feaa34
Download PDF transcript: https://app.rescript.info/api/public/sessions/7faad4e03ce9fd93/pdf
Prof Murray Shanahan:
https://www.doc.ic.ac.uk/~mpsha/ (look at his selected publications)
https://scholar.google.co.uk/citations?user=00bnGpAAAAAJ&hl=en
https://en.wikipedia.org/wiki/Murray_Shanahan
https://x.com/mpshanahan There are monsters in your LLM. (Murray Shanahan)](https://i.ytimg.com/vi/ztNdagyT8po/mqdefault.jpg)
![We need AIs with PHYSICAL experience (Jeff Beck)
Jeff Beck spent years studying computational neuroscience with Alexandre Pouget, Peter Latham, and Wei Ji Ma before founding Noumenal Labs. In this conversation with Tim Scarfe, he argues that language models are fundamentally limited because they manipulate symbols without the physical grounding that gives those symbols meaning.
Beck walks through the Bayesian brain hypothesis — the idea that our brains maintain probabilistic models of the world, built from direct sensory experience. Language, he says, is just a thin lossy summary of that richer internal model. He points to Markus Meisters work showing that human information output runs at roughly 10 bits per second, a tiny fraction of what comes in.
The conversation gets interesting around the question of whether AI systems can genuinely understand anything. Becks position is clear: until a language model produces something genuinely novel — not recombined from training data — he wont call it intelligent. The second half turns to alignment — Beck argues that you cannot separate someones beliefs from their reward function just by observing their behavior, a claim with serious implications for AI safety.
SPONSOR MESSAGES:
***
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich.
Goto https://tufalabs.ai/
***
TIMESTAMPS:
00:00:00 Bayesian brain and neural computation foundations
00:02:00 Belief-reward entanglement and AI alignment
00:02:19 Sponsor: Tufa Labs
00:02:45 Visual cortex discovery and the Hubel-Wiesel experiment
00:05:00 Theory of mind tests and ChatGPT limitations
00:08:50 The sum and product riddle: pattern recognition vs reasoning
00:09:30 Systems engineering and object decomposition
00:12:20 Vicarious experience and grounded understanding
00:15:20 Neural coding and choice of generative model
00:17:30 Line-of-sight legibility in AI reasoning
00:18:50 Language as lossy compression of cognition
00:22:50 10 bits per second: information processing bottleneck
00:25:00 Why language models cannot substitute for understanding
00:28:00 Scientific abstraction and idealization
00:29:40 Markov blankets and system partitioning
00:33:20 Scientific realism, noise, and the limits of models
00:36:00 Free energy principle as mathematical framework
00:38:40 Black-box prediction vs legible explanation
00:41:00 Reward functions and the impossibility of alignment
00:45:00 Modeling beliefs as prerequisite for value inference
REFERENCES:
paper:
[00:00:15] Bayesian inference in neural computation (Ma, Beck, Latham, Pouget)
https://www.nature.com/articles/nn1790
[00:01:50] Noumenal Labs research paper
https://arxiv.org/html/2502.13161v1
[00:05:25] LLM performance on theory of mind tasks (Kosinski)
https://arxiv.org/abs/2302.02083
[00:16:25] Bayesian Mechanics (Ramstead, Sakthivadivel, Heins et al.)
https://arxiv.org/abs/2205.11543
[00:17:10] Building Human-like Communicative Intelligence (Dubova)
https://arxiv.org/abs/2201.02734
[00:20:45] Why do we live at 10 bits/s? (Meister, Zheng)
https://www.sciencedirect.com/science/article/abs/pii/S0896627324008080
[00:28:30] Markov blankets in biological systems (Friston)
https://royalsocietypublishing.org/doi/10.1098/rsif.2017.0792
[00:32:15] Critique of Gabor patches in neuroscience (Tsao)
https://pmc.ncbi.nlm.nih.gov/articles/PMC9564096/
[00:35:55] MaxEnt and MaxCal principles (Presse et al.)
https://journals.aps.org/rmp/abstract/10.1103/RevModPhys.85.1115
[00:40:08] Reward is Enough (Silver, Singh, Precup, Sutton)
https://www.sciencedirect.com/science/article/pii/S0004370221000862
[00:45:40] Modeling Human Beliefs about AI Behavior (Lang, Forre)
https://www.arxiv.org/pdf/2502.21262
website:
[00:01:50] Noumenal Labs (Jeff Beck)
https://www.noumenal.ai/
[00:02:19] Tufa AI Labs
https://tufalabs.ai/
[00:22:25] Steven Piantadosi
https://colala.berkeley.edu/people/piantadosi/
[00:23:35] Mad Libs (Stern, Price)
https://en.wikipedia.org/wiki/Mad_Libs
[00:31:05] Mathematical Platonism (Linnebo, SEP)
https://plato.stanford.edu/entries/platonism-mathematics/
book:
[00:25:25] The Brain Abstracted (Chirimuuta)
https://mitpress.mit.edu/9780262548045/the-brain-abstracted/
LINKS:
Full Transcript: https://app.rescript.info/share/75dfad07d8cad3ed6ccd03253aa006be
Download PDF transcript: https://app.rescript.info/api/public/sessions/981caa3aeba083df/pdf
Extended version on patreon:
https://www.patreon.com/posts/jeff-beck-125455115 We need AIs with PHYSICAL experience (Jeff Beck)](https://i.ytimg.com/vi/zv6qzWecj5c/mqdefault.jpg)