Uploaded September 2024 | Updated September 2026, 1 week ago
Gary Marcus, cognitive scientist and one of AI's most prominent critics, sits down with Tim Scarfe for a nearly two-hour dissection of the current state of artificial intelligence and the tech industry that builds it. Timed to the release of his book Taming Silicon Valley, Marcus makes a sustained case that large language models remain fundamentally brittle — impressive on the surface, yet lacking the compositional semantics, object permanence, and genuine world models that would make them reliable.
The discussion ranges across AI safety theater and the gap between marketing claims and actual capability, the persistent failure modes of image generation and code synthesis, and why LLMs playing chess badly tells us something important about what they have and have not learned. Marcus is especially pointed on industry resistance to regulation, walking through the politics of California's SB-1047, the EU AI Act, and the structural reasons why voluntary self-governance by tech companies has not worked.
The second half turns to the societal damage already visible: copyright battles, the firehose of AI-generated misinformation, surveillance capitalism, and the erosion of trust in digital information. Marcus and Scarfe spar over x-risk and alignment, with Marcus arguing that near-term harms deserve more attention than speculative doom scenarios. The conversation closes on the prospects for neuro-symbolic AI, the legacy of thinkers like Chomsky and Piaget, and what it would actually take to build systems that reason rather than pattern-match.
---
REFERENCES:
book:
[00:00:00] Gary Marcus - Taming Silicon Valley
amzn.to/3XTlC5s
[00:57:26] Shoshana Zuboff - Surveillance Capitalism
amzn.to/3ZqHAxS
[01:23:14] Sayash Kapoor and Arvind Narayanan - AI Snake Oil
aisnakeoil.com
[01:23:14] Isaac Asimov - Three Laws of Robotics
amzn.to/3XTIwtl
[01:44:33] Gary Marcus - The Algebraic Mind
https://mitpress.mit.edu/books/algebraic-mind
person:
[00:00:00] Gary Marcus Substack
garymarcus.substack.com
[00:23:49] Jean Piaget - Object Permanence Theory
en.wikipedia.org/wiki/Object_permanence
[01:44:33] Seymour Papert - Logo Programming Language
https://el.media.mit.edu/logo-foundation/what_is_logo/logo_primer.html
paper:
[00:23:49] Evelina Leivada et al. - LLMs Understanding Human Language
arxiv.org/pdf/2308.00109
[00:31:09] Alan Turing - Computing Machinery and Intelligence
academic.oup.com/mind/article/LIX/236/433/986238
[00:34:45] Nicholas Carlini - Chess with LLMs
nicholas.carlini.com/writing/2023/chess-llm.html
[00:34:45] Mathieu Acher - GPT-4 Chess Analysis
blog.mathieuacher.com/ChessWinning7MovesGPT
[00:42:10] Jiexin Wang - AI-Assisted Coding Security
arxiv.org/pdf/2407.02395v1
[00:42:10] Rodney Brooks - Three Laws of AI
rodneybrooks.com/rodney-brooks-three-laws-of-artificial-intelligence
[00:48:10] Gary Marcus - Open Letter on SB-1047
garymarcus.substack.com/p/an-open-letter-to-fei-fei-li-concerning
[00:57:26] Jaron Lanier - Twitter Poisoning
nytimes.com/2022/11/11/opinion/trump-musk-kanye-twitter.html
[00:57:26] Chris Lu et al. - AI Scientist Paper
arxiv.org/abs/2408.06292
[01:23:14] Gary Marcus - p(doom) Analysis
garymarcus.substack.com/p/d28
[01:44:33] Gary Marcus - Neural Networks and Generalization
sciencedirect.com/science/article/pii/S0010028598906946
other:
[00:23:49] Gottlob Frege - Compositional Semantics
https://plato.stanford.edu/entries/compositionality/
[00:42:10] Cruise Teleoperation Revelation
nytimes.com/2023/11/03/technology/cruise-general-motors-self-driving-cars.html
[00:48:10] California SB-1047 AI Regulation
apcp.assembly.ca.gov/system/files/2024-06/sb-1047-wiener-apcp-analysis_0.pdf
[00:48:10] European Commission - EU AI Act
https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai
[00:54:55] A&M Records v. Napster - Copyright Precedent
en.wikipedia.org/wiki/A%26M_Records,_Inc._v._Napster,_Inc.
[00:57:26] Firehose of Falsehood - Russian Propaganda Model
en.wikipedia.org/wiki/Firehose_of_falsehood
[01:44:33] Laplace Demon Concept
en.wikipedia.org/wiki/Laplace%27s_demon
video:
[00:57:26] Adam Curtis - HyperNormalisation
imdb.com/title/tt6156350
[00:57:26] Luciano Floridi - Digital Ethics
youtube.com/watch?v=YLNGvvgq3eg
[01:44:33] MLST - Noam Chomsky Interview
youtube.com/watch?v=axuGfh4UR9Q
[01:44:33] Beff Jezos - Physics-Inspired Intelligence
youtube.com/watch?v=0zxi0xSBOaQ
---
LINKS:
Full Transcript: app.rescript.info/share/b2f49ea9ca469401e8f65b29156495c8
Download PDF transcript: app.rescript.info/api/public/sessions/530cbbb31f702d43/pdf
Gary Marcus:
garymarcus.substack.com
https://x.com/GaryMarcus
Gary Marcus, cognitive scientist and one of AI's most prominent critics, sits down with Tim Scarfe for a nearly two-hour dissection of the current state of artificial intelligence and the tech industry that builds it. Timed to the release of his book Taming Silicon Valley, Marcus makes a sustained case that large language models remain fundamentally brittle — impressive on the surface, yet lacking the compositional semantics, object permanence, and genuine world models that would make them reliable.
The discussion ranges across AI safety theater and the gap between marketing claims and actual capability, the persistent failure modes of image generation and code synthesis, and why LLMs playing chess badly tells us something important about what they have and have not learned. Marcus is especially pointed on industry resistance to regulation, walking through the politics of California's SB-1047, the EU AI Act, and the structural reasons why voluntary self-governance by tech companies has not worked.
The second half turns to the societal damage already visible: copyright battles, the firehose of AI-generated misinformation, surveillance capitalism, and the erosion of trust in digital information. Marcus and Scarfe spar over x-risk and alignment, with Marcus arguing that near-term harms deserve more attention than speculative doom scenarios. The conversation closes on the prospects for neuro-symbolic AI, the legacy of thinkers like Chomsky and Piaget, and what it would actually take to build systems that reason rather than pattern-match.
---
REFERENCES:
book:
[00:00:00] Gary Marcus - Taming Silicon Valley
amzn.to/3XTlC5s
[00:57:26] Shoshana Zuboff - Surveillance Capitalism
amzn.to/3ZqHAxS
[01:23:14] Sayash Kapoor and Arvind Narayanan - AI Snake Oil
aisnakeoil.com
[01:23:14] Isaac Asimov - Three Laws of Robotics
amzn.to/3XTIwtl
[01:44:33] Gary Marcus - The Algebraic Mind
https://mitpress.mit.edu/books/algebraic-mind
person:
[00:00:00] Gary Marcus Substack
garymarcus.substack.com
[00:23:49] Jean Piaget - Object Permanence Theory
en.wikipedia.org/wiki/Object_permanence
[01:44:33] Seymour Papert - Logo Programming Language
https://el.media.mit.edu/logo-foundation/what_is_logo/logo_primer.html
paper:
[00:23:49] Evelina Leivada et al. - LLMs Understanding Human Language
arxiv.org/pdf/2308.00109
[00:31:09] Alan Turing - Computing Machinery and Intelligence
academic.oup.com/mind/article/LIX/236/433/986238
[00:34:45] Nicholas Carlini - Chess with LLMs
nicholas.carlini.com/writing/2023/chess-llm.html
[00:34:45] Mathieu Acher - GPT-4 Chess Analysis
blog.mathieuacher.com/ChessWinning7MovesGPT
[00:42:10] Jiexin Wang - AI-Assisted Coding Security
arxiv.org/pdf/2407.02395v1
[00:42:10] Rodney Brooks - Three Laws of AI
rodneybrooks.com/rodney-brooks-three-laws-of-artificial-intelligence
[00:48:10] Gary Marcus - Open Letter on SB-1047
garymarcus.substack.com/p/an-open-letter-to-fei-fei-li-concerning
[00:57:26] Jaron Lanier - Twitter Poisoning
nytimes.com/2022/11/11/opinion/trump-musk-kanye-twitter.html
[00:57:26] Chris Lu et al. - AI Scientist Paper
arxiv.org/abs/2408.06292
[01:23:14] Gary Marcus - p(doom) Analysis
garymarcus.substack.com/p/d28
[01:44:33] Gary Marcus - Neural Networks and Generalization
sciencedirect.com/science/article/pii/S0010028598906946
other:
[00:23:49] Gottlob Frege - Compositional Semantics
https://plato.stanford.edu/entries/compositionality/
[00:42:10] Cruise Teleoperation Revelation
nytimes.com/2023/11/03/technology/cruise-general-motors-self-driving-cars.html
[00:48:10] California SB-1047 AI Regulation
apcp.assembly.ca.gov/system/files/2024-06/sb-1047-wiener-apcp-analysis_0.pdf
[00:48:10] European Commission - EU AI Act
https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai
[00:54:55] A&M Records v. Napster - Copyright Precedent
en.wikipedia.org/wiki/A%26M_Records,_Inc._v._Napster,_Inc.
[00:57:26] Firehose of Falsehood - Russian Propaganda Model
en.wikipedia.org/wiki/Firehose_of_falsehood
[01:44:33] Laplace Demon Concept
en.wikipedia.org/wiki/Laplace%27s_demon
video:
[00:57:26] Adam Curtis - HyperNormalisation
imdb.com/title/tt6156350
[00:57:26] Luciano Floridi - Digital Ethics
youtube.com/watch?v=YLNGvvgq3eg
[01:44:33] MLST - Noam Chomsky Interview
youtube.com/watch?v=axuGfh4UR9Q
[01:44:33] Beff Jezos - Physics-Inspired Intelligence
youtube.com/watch?v=0zxi0xSBOaQ
---
LINKS:
Full Transcript: app.rescript.info/share/b2f49ea9ca469401e8f65b29156495c8
Download PDF transcript: app.rescript.info/api/public/sessions/530cbbb31f702d43/pdf
Gary Marcus:
garymarcus.substack.com
https://x.com/GaryMarcus
![Why Program Synthesis Is Next (Kevin Ellis and Zenna Tavares)
Kevin Ellis (Cornell) and Zenna Tavares (BASIS) argue that the next wave of AI needs to learn like humans do: building abstract models from small amounts of data through active exploration, not just passive pattern matching at scale.
The conversation centers on their joint work comparing two fundamentally different ways of solving problems. Induction searches for an explicit program something you could write in Python that transforms inputs to outputs. Transduction skips the program and directly predicts the answer, the way a neural network would. On the Abstraction and Reasoning Corpus (ARC), these approaches turn out to be complementary: some problems yield to systematic symbolic search, others to neural intuition. The ensemble is stronger than either alone, and the reasons connect to findings in cognitive science about when explicit reasoning helps versus hurts.
Kevin explains how his DreamCoder work pioneered a wake-sleep cycle for program synthesis: dream up programs, run them to see what they do, learn the inverse mapping, then wake up and let real-world failures adjust the distribution of dreams. The modern version replaces explicit symbolic libraries with in-context learning over LLM-generated code, keeping the same iterative refinement loop.
Zenna introduces his Autumn system for synthesizing the source code of interactive environments from observed behavior a form of computational science where the model must also infer hidden state it cannot directly observe. Both researchers converge on the idea that abstraction is the key unsolved problem: real intelligence requires knowing what to ignore, not just what to represent. Zenna frames this through resource rationality choosing the right level of abstraction given your computational budget and expected tasks.
The discussion closes with Project MARA, their joint effort to build interactive benchmarks that go beyond ARCs static puzzles, requiring agents to actively explore and build world models from scratch.
SPONSOR MESSAGES:
***
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich.
REFERENCES:
Paper:
[00:00:25] DreamCoder: Growing Generalizable, Interpretable Knowledge with Wake-Sleep Bayesian Program Learning
https://arxiv.org/abs/2006.08381
[00:01:10] Mind Your Step: Active Search over Compositional Spaces
https://arxiv.org/abs/2410.21333
[00:06:05] Bayesian inference in the cognitive sciences
https://psycnet.apa.org/record/2008-06911-003
[00:13:00] Induction and Transduction
https://arxiv.org/abs/2411.02272
[00:23:15] Neurosymbolic AI: The 3rd Wave
https://arxiv.org/abs/2012.05876
[00:38:35] On the Measure of Intelligence (ARC)
https://arxiv.org/abs/1911.01547
[00:39:20] Causal Reactive Programs (Autumn)
http://www.zenna.org/publications/autumn2022.pdf
[00:42:50] MuZero
http://arxiv.org/pdf/1911.08265
[00:43:20] VisualPredicator
https://arxiv.org/abs/2410.23156
Book:
[00:48:55] Bayesian Models of Cognition
https://mitpress.mit.edu/9780262049412/bayesian-models-of-cognition/
Essay:
[00:49:30] The Bitter Lesson
http://www.incompleteideas.net/IncIdeas/BitterLesson.html
Project:
[01:11:55] Project MARA
https://www.basis.ai/blog/mara/
LINKS:
Full Transcript: https://app.rescript.info/share/e0a208e545cabae728a3d72f76fcd310
Download PDF transcript: https://app.rescript.info/api/public/sessions/f47975e800b064d9/pdf Why Program Synthesis Is Next (Kevin Ellis and Zenna Tavares)](https://i.ytimg.com/vi/oYTm0p3DCzg/mqdefault.jpg)
![#70 - LETITIA PARCALABESCU - Symbolics, Linguistics [UNPLUGGED]
#70 - LETITIA PARCALABESCU - Symbolics, Linguistics [UNPLUGGED] #70 - LETITIA PARCALABESCU - Symbolics, Linguistics [UNPLUGGED]](https://i.ytimg.com/vi/p2D2duT-R2E/mqdefault.jpg)
![Why Every Brain Metaphor in History Has Been Wrong [SPECIAL EDITION]
What if everything we think we know about the brain is just a really good metaphor that we forgot was a metaphor?
This episode takes you on a journey through the history of scientific simplification, from a young Karl Friston watching wood lice in his garden to the bold claims that your mind is literally software running on biological hardware.
We bring together some of the most brilliant minds weve interviewed — Professor Mazviita Chirimuuta, Francois Chollet, Joscha Bach, Professor Luciano Floridi, Professor Noam Chomsky, Nobel laureate John Jumper, and more — to wrestle with a deceptively simple question: *When scientists simplify reality to study it, what gets captured and what gets lost?*
*Key ideas explored:*
*The Spherical Cow Problem* — Science requires simplification. Were limited creatures trying to understand systems far more complex than our working memory can hold. But when does a useful model become a dangerous illusion?
*The Kaleidoscope Hypothesis* — Francois Chollets beautiful idea that beneath all the apparent chaos of reality lies simple, repeating patterns — like bits of colored glass in a kaleidoscope creating infinite complexity. Is this profound truth or Platonic wishful thinking?
*Is Software Really Spirit?* — Joscha Bach makes the provocative claim that software is literally spirit, not metaphorically. We push back hard on this, asking whether the sameness we see across different computers running the same program exists in nature or only in our descriptions.
*The Cultural Illusion of AGI* — Why does artificial general intelligence seem so inevitable to people in Silicon Valley? Professor Chirimuuta suggests we might be caught in a cultural historical illusion — our mechanistic assumptions about minds making AI seem like destiny when it might just be a bet.
*Prediction vs. Understanding* — Nobel Prize winner John Jumper: AI can predict and control, but understanding requires a human in the loop.
Throughout history, weve described the brain as hydraulic pumps, telegraph networks, telephone switchboards, and now computers. Each metaphor felt obviously true at the time. This episode asks: what will we think was naive about our current assumptions in fifty years?
Featuring insights from *The Brain Abstracted* by Mazviita Chirimuuta — possibly the most influential book on how we think about thinking in 2025.
TIMESTAMPS:
00:00:00 The Wood Louse & The Spherical Cow
00:02:04 The Necessity of Abstraction
00:04:42 Simplicius vs. Ignorantio: The Boxing Match
00:06:39 The Kaleidoscope Hypothesis
00:08:40 Is the Mind Software?
00:13:15 Critique of Causal Patterns
00:14:40 Temperature is Not a Thing
00:18:24 The Ship of Theseus & Ontology
00:23:45 Metaphors Hardening into Reality
00:25:41 The Illusion of AGI Inevitability
00:27:45 Prediction vs. Understanding
00:32:00 Climbing the Mountain vs. The Helicopter
00:34:53 Haptic Realism & The Limits of Knowledge
REFERENCES:
Person:
[00:00:00] Karl Friston (UCL)
https://profiles.ucl.ac.uk/1236-karl-friston
[00:06:30] Francois Chollet
https://fchollet.com/
[00:14:41] Cesar Hidalgo, MLST interview.
https://www.youtube.com/watch?v=vzpFOJRteeI
[00:30:30] Terence Taos Blog
https://terrytao.wordpress.com/
Book:
[00:02:25] The Brain Abstracted
https://mitpress.mit.edu/9780262548045/the-brain-abstracted/
[00:06:00] On Learned Ignorance
https://www.amazon.com/Nicholas-Cusa-learned-ignorance-translation/dp/0938060236
[00:24:15] Science and the Modern World
https://amazon.com/dp/0684836394
Interview.:
[00:02:43] The Brain Abstracted Patreon interview
https://www.patreon.com/posts/brain-abstracted-124479979
Interview:
[00:04:18] David Krakauers presentation on intelligence.
https://www.youtube.com/watch?v=dY46YsGWMIc
[00:06:45] Machine Learning Street Talk interview with Francois Chollet.
https://www.youtube.com/watch?v=JTU8Ha4Jyfc
[00:09:32] Joscha Bach Patreon interview.
https://www.patreon.com/posts/joscha-bach-deep-141884561
[00:18:24] Luciano Floridi, MLST interview.
https://www.youtube.com/watch?v=YLNGvvgq3eg
[00:25:02] Jeff Beck, MLST interview.
https://www.youtube.com/watch?v=9suqiofCiwM
[00:28:00] John Jumper
https://www.patreon.com/posts/john-jumper-on-144557652
[00:30:48] Noam Chomsky, MLST interview.
https://www.youtube.com/watch?v=axuGfh4UR9Q
[00:31:59] Anna Ciaunica
https://www.patreon.com/posts/dr-anna-ciaunica-144509970
[00:33:12] Mike Israetel debate on functionalism.
https://www.youtube.com/watch?v=4yYcN_mFi18
Company/Org:
[00:09:25] Neuralink
https://neuralink.com/
Paper:
[00:23:45] A Logical Calculus of Ideas Immanent in Nervous Activity
https://link.springer.com/article/10.1007/BF02478259
[00:28:00] Highly Accurate Protein Structure Prediction with AlphaFold
https://www.nature.com/articles/s41586-021-03819-2
Thank you to Dr. Maxwell Ramstead for early script work on this show (Ph.D student of Friston) and the woodlice story came from him! Why Every Brain Metaphor in History Has Been Wrong [SPECIAL EDITION]](https://i.ytimg.com/vi/pO0WZsN8Oiw/mqdefault.jpg)

![Google Researcher Shows Life Emerges From Code [Blaise Agüera y Arcas]
Blaise Agüera y Arcas explores some mind-bending ideas about what intelligence and life really are—and why they might be more similar than we think (filmed at ALIFE conference, 2025 - https://2025.alife.org/ ).
Life and intelligence are both fundamentally computational (he says). From the very beginning, living things have been running programs. Your DNA? Its literally a computer program, and the ribosomes in your cells are tiny universal computers building you according to those instructions.
**SPONSOR MESSAGES**
—
Prolific - Quality data. From real people. For faster breakthroughs.
https://www.prolific.com/?utm_source=mlst
—
cyber•Fund https://cyber.fund/?utm_source=mlst is a founder-led investment firm accelerating the cybernetic economy
Oct SF conference - https://dagihouse.com/?utm_source=mlst - Joscha Bach keynoting(!) + OAI, Anthropic, NVDA,++
Hiring a SF VC Principal: https://talent.cyber.fund/companies/cyber-fund-2/jobs/57674170-ai-investment-principal#content?utm_source=mlst
Submit investment deck: https://cyber.fund/contact?utm_source=mlst
—
Blaise argues that there is more to evolution than random mutations (like most people think). The secret to increasing complexity is *merging* i.e. when different organisms or systems come together and combine their histories and capabilities.
Blaise describes his BFF experiment where random computer code spontaneously evolved into self-replicating programs, showing how purpose and complexity can emerge from pure randomness through computational processes.
https://en.wikipedia.org/wiki/Blaise_Ag%C3%BCera_y_Arcas
https://x.com/blaiseaguera?lang=en
TRANSCRIPT:
https://app.rescript.info/public/share/VX7Gktfr3_wIn4Bj7cl9StPBO1MN4R5lcJ11NE99hLg
TOC:
00:00:00 Introduction - New book What is Intelligence?
00:01:45 Life as computation - Von Neumanns insights
00:12:00 BFF experiment - How purpose emerges
00:26:00 Symbiogenesis and evolutionary complexity
00:40:00 Functionalism and consciousness
00:49:45 AI as part of collective human intelligence
00:57:00 Comparing AI and human cognition
REFS:
What is intelligence [Blaise Agüera y Arcas]
https://whatisintelligence.antikythera.org/ [Read free online, interactive rich media]
https://mitpress.mit.edu/9780262049955/what-is-intelligence/ [MIT Press]
Large Language Models and Emergence: A Complex Systems Perspective
https://arxiv.org/abs/2506.11135
Our first Noam Chomsky MLST interview
https://www.youtube.com/watch?v=axuGfh4UR9Q
Chance and Necessity [Jacques Monod]
https://monoskop.org/images/9/99/Monod_Jacques_Chance_and_Necessity.pdf
Wonderful Life: The Burgess Shale and the History of Nature [Stephen Jay Gould]
https://www.amazon.co.uk/Wonderful-Life-Burgess-Nature-History/dp/0099273454
The major evolutionary transitions [E Szathmáry, J M Smith]
https://wiki.santafe.edu/images/0/0e/Szathmary.MaynardSmith_1995_Nature.pdf
Dont Sleep, There Are Snakes: Life and Language in the Amazonian Jungle [Dan Everett]
https://www.amazon.com/Dont-Sleep-There-Are-Snakes/dp/0307386120
The Nature of Technology: What It Is and How It Evolves [W. Brian Arthur]
https://www.amazon.com/Nature-Technology-What-How-Evolves-ebook/dp/B002RI9W16/
The MANIAC [Benjamin Labatut]
https://www.amazon.com/MANIAC-Benjam%C3%ADn-Labatut/dp/1782279814
When We Cease to Understand the World [Benjamin Labatut]
https://www.amazon.com/When-We-Cease-Understand-World/dp/1681375664/
The Boys in the Boat [Dan Brown]
https://www.amazon.com/Boys-Boat-Americans-Berlin-Olympics/dp/0143125478
How something can be said about Telling More Than We Can Know [Petter Johansson] (Split brain)
https://www.lucs.lu.se/fileadmin/user_upload/lucs/2011/01/Johansson-et-al.-2006-How-Something-Can-Be-Said-About-Telling-More-Than-We-Can-Know.pdf
If Anyone Builds It, Everyone Dies [Eliezer Yudkowsky, Nate Soares]
https://www.amazon.com/Anyone-Builds-Everyone-Dies-Superhuman/dp/0316595640
The science of cycology: Failures to understand how everyday objects work [Rebeca Lawson]
https://link.springer.com/content/pdf/10.3758/bf03195929.pdf
SEVA: Leveraging sketches to evaluate alignment between human and machine visual abstraction [Kushin Mukherjee, Judith Fan et al]
https://arxiv.org/pdf/2312.03035
Brain-Score [Martin Schrimpf]
https://www.biorxiv.org/content/10.1101/407007v1.full.pdf
Feature Visualization [Chris Olah]
https://distill.pub/2017/feature-visualization/
THE REVERSAL CURSE [Lukas Berglund]
https://arxiv.org/pdf/2309.12288 Google Researcher Shows Life Emerges From Code [Blaise Agüera y Arcas]](https://i.ytimg.com/vi/rMSEqJ_4EBk/mqdefault.jpg)
![AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart
This episode is sponsored by Notion. Learn more about Notions Developer Platform today at https://notion.com/mlst
Why can deep networks discover abstractions that shallow models miss? Statistical physicist Matthieu Wyart joins Tim Scarfe to argue that the answer lies in the hidden hierarchy of data. Language and images are built from parts within parts; depth lets a network recover those coarse-grained variables and escape the curse of dimensionality.
The conversation moves from jamming transitions and rough loss surfaces to Chomsky, context-free grammars and machine creativity. Wyart explains why next-token prediction can still recover compositional structure, where current systems fall short of genuine scientific invention, and why predicting latent representations rather than raw tokens could make learning far more sample-efficient.
They also examine diffusion models, neural scaling laws and the limits of physics-inspired theory. The final question is on a personal note: if mistakes are the price of leaving the beaten path, how much scientific risk is worth taking?
TIMESTAMPS:
00:00:00 Can machines learn abstractions from data?
00:02:00 Notion agentic workspace
00:02:49 From statistical physics to machine learning
00:06:40 What physics can explain about learning
00:16:37 From Carnot to Chomsky bulldozer
00:21:21 How deep networks recover hidden hierarchies
00:32:43 Where machine creativity still falls short
00:40:48 How deep nets escape the curse of dimensionality
00:52:19 Why predict latents instead of tokens
01:02:49 The sample-efficiency case for latent prediction
01:08:31 Diffusion, scaling laws and text entropy
01:16:40 The scientists we learn from and the mistakes we make
REFERENCES:
person:
[00:00:43] Noam Chomsky
https://linguistics.mit.edu/user/chomsky/
tool:
[00:02:08] Notion Developer Platform
https://www.notion.com/en-gb/blog/introducing-developer-platform
paper:
[00:04:43] Mastering the game of Go with deep neural networks and tree search
https://www.nature.com/articles/nature16961
[00:05:52] Reconciling modern machine-learning practice and the bias-variance trade-off
https://arxiv.org/abs/1812.11118
[00:25:54] How Deep Neural Networks Learn Compositional Data: The Random Hierarchy Model
https://arxiv.org/abs/2307.02129
[00:42:12] Efficient Estimation of Word Representations in Vector Space
https://arxiv.org/abs/1301.3781
[00:52:46] Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
https://arxiv.org/abs/2301.08243
[00:52:54] Learn from your own latents and not from tokens: A sample-complexity theory
https://arxiv.org/abs/2605.27734
[01:08:31] A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Data
https://arxiv.org/abs/2402.16991
[01:11:39] Scaling Laws for Neural Language Models
https://arxiv.org/abs/2001.08361
[01:12:17] Deriving Neural Scaling Laws from the statistics of natural language
https://arxiv.org/abs/2602.07488
[01:13:34] Prediction and Entropy of Printed English
https://ieeexplore.ieee.org/document/6773263
LINKS:
Download PDF transcript: https://app.rescript.info/share/f7644cdaa86c5cc1e41e484e290f2bd4 AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart](https://i.ytimg.com/vi/revreN8LZ_M/mqdefault.jpg)

![Why High Benchmark Scores Don’t Mean Better AI [SPONSORED]
Is a car that wins a Formula 1 race the best choice for your morning commute? Probably not. In this sponsored deep dive with Prolific, we explore why the same logic applies to Artificial Intelligence. While models are currently shattering records on technical exams, they often fail the most important test of all: *the human experience.*
Why High Benchmark Scores Don’t Mean Better AI
Joining us are *Andrew Gordon* (Staff Researcher in Behavioral Science) and *Nora Petrova* (AI Researcher) from *Prolific* . They reveal the hidden flaws in how we currently rank AI and introduce a more rigorous, humane way to measure whether these models are actually helpful, safe, and relatable for real people.
Key Insights in This Episode:
* *The F1 Car Analogy:* Andrew explains why a model that excels at the Humanities Last Exam might be a nightmare for daily use. Technical benchmarks often ignore the nuances of human communication and adaptability.
* *The Wild West of AI Safety:* As users turn to AI for sensitive topics like mental health, Nora highlights the alarming lack of oversight and the thin veneer of safety training—citing recent controversial incidents like Grok-3’s Mecha Hitler.
* *Fixing the Leaderboard Illusion:* The team critiques current popular rankings like Chatbot Arena, discussing how anonymous, unstratified voting can lead to biased results and how companies can game the system.
* *The Xbox Secret to AI Ranking:* Discover how Prolific uses *TrueSkill* —the same algorithm Microsoft developed for Xbox Live matchmaking—to create a fairer, more statistically sound leaderboard for LLMs.
* *The Personality Gap:* Early data from the *Humane Leaderboard* suggests that while AI is getting smarter, it is actually performing *worse* on metrics like personality, culture, and sycophancy (the tendency for models to become annoying people-pleasers).
About the HUMAINE Leaderboard
Moving beyond simple A vs. B testing, the researchers discuss their new framework that samples participants based on *census data* (Age, Ethnicity, Political Alignment). By using a representative sample of the general public rather than just tech enthusiasts, they are building a standard that reflects the values of the real world.
*Are we building models for benchmarks, or are we building them for humans? It’s time to change the scoreboard.*
Rescript link:
https://app.rescript.info/public/share/IDqwjY9Q43S22qSgL5EkWGFymJwZ3SVxvrfpgHZLXQc
TIMESTAMPS:
00:00:00 Introduction & The Benchmarking Problem
00:01:58 The Fractured State of AI Evaluation
00:03:54 AI Safety & Interpretability
00:05:45 Bias in Chatbot Arena
00:06:45 Prolifics Three Pillars Approach
00:09:01 TrueSkill Ranking & Efficient Sampling
00:12:04 Census-Based Representative Sampling
00:13:00 Key Findings: Culture, Personality & Sycophancy
REFERENCES:
Paper:
[00:00:15] MMLU
https://arxiv.org/abs/2009.03300
[00:05:10] Constitutional AI
https://arxiv.org/abs/2212.08073
[00:06:45] The Leaderboard Illusion
https://arxiv.org/abs/2504.20879
[00:09:41] HUMAINE Framework Paper
https://huggingface.co/blog/ProlificAI/humaine-framework
Company:
[00:00:30] Prolific
https://www.prolific.com
[00:01:45] Chatbot Arena
https://lmarena.ai/
Person:
[00:00:35] Andrew Gordon
https://www.linkedin.com/in/andrew-gordon-03879919a/
[00:00:45] Nora Petrova
https://www.linkedin.com/in/nora-petrova/
Event:
Algorithm:
[00:09:01] Microsoft TrueSkill
https://www.microsoft.com/en-us/research/project/trueskill-ranking-system/
Leaderboard:
[00:09:21] Prolific HUMAINE Leaderboard
https://www.prolific.com/humaine
[00:09:31] HUMAINE HuggingFace Space
https://huggingface.co/spaces/ProlificAI/humaine-leaderboard
[00:10:21] Prolific AI Leaderboard Portal
https://www.prolific.com/leaderboard
Dataset:
[00:09:51] Prolific Social Reasoning RLHF Dataset
https://huggingface.co/datasets/ProlificAI/social-reasoning-rlhf
Organization:
[00:10:31] MLCommons
https://mlcommons.org/ Why High Benchmark Scores Don’t Mean Better AI [SPONSORED]](https://i.ytimg.com/vi/rqiC9a2z8Io/mqdefault.jpg)


![Its Not About Scale, Its About Abstraction
MLST is sponsored by Tufa Labs:
Are you interested in working on ARC and cutting-edge AI research with the MindsAI team (current ARC winners)?
Focus: ARC, LLMs, test-time-compute, active inference, system2 reasoning, and more.
Future plans: Expanding to complex environments like Warcraft 2 and Starcraft 2.
Interested? Apply for an ML research position: benjamin@tufa.ai
Francois Chollet, creator of Keras and the ARC-AGI benchmark, delivers his AGI-24 keynote on why scaling LLMs will not get us to AGI. He walks through concrete failure modes LLMs that break on trivial rephrasing of memorized problems, that pattern-match the Monty Hall problem without parsing the actual numbers, that solve Caesar ciphers only for key sizes found in online examples. The failures all point the same way: LLM performance tracks task familiarity, not task complexity.
Chollet introduces his Kaleidoscope Hypothesis: the world looks infinitely complex on the surface, but it is built from a small set of repeating atoms of meaning. Intelligence, in his framing, is the process of mining experience to extract those atoms and recombining them to handle genuinely novel situations. This is what the ARC benchmark is designed to test abstraction and reasoning that cannot be memorized.
The talk closes with a proposal: combine deep learning (good at perception and pattern recognition) with discrete program synthesis (good at precise, compositional reasoning). Neither approach alone gets there, but the hybrid might. Chollet points to early results on ARC from Ryan Greenblatt and others as evidence that the research community outside big labs may be where the next breakthrough comes from.
TIMESTAMPS:
00:00:00 LLM Limitations and Composition
00:12:05 Intelligence as Process vs. Skill
00:17:15 Generalization as Key to AI Progress
00:19:59 Introduction to ARC-AGI Benchmark
00:26:10 The Kaleidoscope Hypothesis and Abstraction Spectrum
00:34:05 Limitations of Transformers and Program Synthesis
00:39:59 Applying Combined Approaches to ARC Tasks
00:44:20 State-of-the-Art Solutions and Future Directions
REFERENCES:
paper:
[00:01:15] On the Measure of Intelligence
https://arxiv.org/abs/1911.01547
[00:03:30] Embers of Autoregression
https://arxiv.org/abs/2309.13638
[00:05:30] Monty Hall problem
https://www.tandfonline.com/doi/abs/10.1080/00031305.1975.10479121
[00:06:20] LLM Training Dynamics Analysis
https://arxiv.org/abs/2205.10770
[00:07:33] GPT-4 Technical Report
https://cdn.openai.com/papers/gpt-4.pdf
[00:10:20] Faith and Fate: Limits of Transformers on Compositionality
https://arxiv.org/abs/2305.18654
[00:10:25] The Reversal Curse in LLMs
https://arxiv.org/abs/2309.12288
[00:10:52] LM-Polygraph: Uncertainty Estimation for LLMs
https://arxiv.org/abs/2311.07383
[00:20:34] Baldur: Whole-Proof Generation
https://arxiv.org/abs/2303.04910
[00:34:00] Core Knowledge in Infants
https://www.harvardlds.org/wp-content/uploads/2017/01/SpelkeKinzler07-1.pdf
[00:44:20] Hypothesis Search with LLMs for ARC
https://arxiv.org/abs/2309.05660
tool:
[00:20:10] ARC-AGI GitHub Repository
https://github.com/fchollet/ARC-AGI
[00:22:15] ARC Prize
https://arcprize.org/
book:
[00:33:30] Thinking, Fast and Slow
https://www.amazon.com/Thinking-Fast-Slow-Daniel-Kahneman/dp/0374533555
LINKS:
Full Transcript: https://app.rescript.info/share/c8b5bacdf1ffefab4f65060edc295d4a
Download PDF transcript: https://app.rescript.info/api/public/sessions/b537d0b92ae48338/pdf
[0:20:10] ARC-AGI: GitHub repository (François Chollet)
https://github.com/fchollet/ARC-AGI Its Not About Scale, Its About Abstraction](https://i.ytimg.com/vi/s7_NlkBwdj8/mqdefault.jpg)