Uploaded December 2025 | Updated September 2026, 1 week ago
What if everything we think we know about AI understanding is wrong? Is compression the key to intelligence? Or is there something more—a leap from memorization to true abstraction?
In this fascinating conversation, we sit down with *Professor Yi Ma*—world-renowned expert in deep learning, IEEE/ACM Fellow, and author of the groundbreaking new book *Learning Deep Representations of Data Distributions*. Professor Ma challenges our assumptions about what large language models actually do, reveals why 3D reconstruction isn't the same as understanding, and presents a unified mathematical theory of intelligence built on just two principles: *parsimony* and *self-consistency*.
**SPONSOR MESSAGES START**
—
Prolific - Quality data. From real people. For faster breakthroughs.
prolific.com/?utm_source=mlst
—
cyber•Fund https://cyber.fund/?utm_source=mlst is a founder-led investment firm accelerating the cybernetic economy
Hiring a SF VC Principal: https://talent.cyber.fund/companies/cyber-fund-2/jobs/57674170-ai-investment-principal#content?utm_source=mlst
Submit investment deck: https://cyber.fund/contact?utm_source=mlst
—
**END**
NOTE: Based on lower retention we have EDITED out 33:45 of the original interview from near the start - philosophical stuff, an extended analogy comparing DNA evolution to AI development, speculation about where theories come from (Platonism, deductive trees), Norbert Wiener's cybernetics history. You can watch the original version on the rescript link or Spotify version.
Key Insights:
*LLMs Don't Understand—They Memorize*
Language models process text (*already* compressed human knowledge) using the same mechanism we use to learn from raw data.
*The Illusion of 3D Vision*
Sora and NeRFs etc that can reconstruct 3D scenes still fail miserably at basic spatial reasoning
*"All Roads Lead to Rome"*
Why adding noise is *necessary* for discovering structure.
*Why Gradient Descent Actually Works*
Natural optimization landscapes are surprisingly smooth—a "blessing of dimensionality"
*Transformers from First Principles*
Transformer architectures can be mathematically derived from compression principles
—
INTERACTIVE AI TRANSCRIPT PLAYER w/REFS (ReScript):
app.rescript.info/public/share/Z-dMPiUhXaeMEcdeU6Bz84GOVsvdcfxU_8Ptu6CTKMQ
About Professor Yi Ma
Yi Ma is the inaugural director of the School of Computing and Data Science at Hong Kong University and a visiting professor at UC Berkeley.
https://people.eecs.berkeley.edu/~yima/
scholar.google.com/citations?user=XqLiBQMAAAAJ&hl=en
https://x.com/YiMaTweets
*Slides from this conversation:*
dropbox.com/scl/fi/sbhbyievw7idup8j06mlr/slides.pdf?rlkey=7ptovemezo8bj8tkhfi393fh9&dl=0
*Related Talks by Professor Ma:*
- Pursuing the Nature of Intelligence (ICLR): youtube.com/watch?v=LT-F0xSNSjo
- Earlier talk at Berkeley: youtube.com/watch?v=TihaCUjyRLM
--
TIMESTAMPS:
00:00:00 Introduction
00:02:08 The First Principles Book & Research Vision
00:05:21 Two Pillars: Parsimony & Consistency
00:09:50 Evolution vs. Learning: The Compression Mechanism
00:14:37 The Illusion of 3D Understanding: Sora & NeRF
00:20:41 All Roads Lead to Rome: The Role of Noise
00:26:11 All Roads Lead to Rome: The Role of Noise
00:26:29 Benign Non-Convexity: Why Optimization Works
00:32:50 Double Descent & The Myth of Overfitting
00:40:41 Self-Consistency: Closed-Loop Learning
00:47:18 Deriving Transformers from First Principles
00:56:26 Verification & The Kevin Murphy Question
01:00:26 CRATE vs. ViT: White-Box AI & Conclusion
---
REFERENCES:
Book:
[00:03:04] Learning Deep Representations of Data Distributions
ma-lab-berkeley.github.io/deep-representation-learning-book
Book (Yi Ma):
[00:03:14] An Invitation to 3-D Vision
link.springer.com/book/10.1007/978-0-387-21779-6
[00:03:24] Generalized Principal Component Analysis
link.springer.com/book/10.1007/978-0-387-87811-9
[00:03:34] High-Dimensional Data Analysis with Low-Dimensional Models
book-wright-ma.github.io
Slide:
[00:44:11] Slide 26: Neuroscience Evidence
arxiv.org/abs/2207.04630)
Person:
[00:08:24] Albert Einstein
quoteinvestigator.com/2011/05/13/einstein-simple
[00:56:41] Kevin Murphy
probml.github.io/pml-book/book1.html
Paper:
[00:18:09] Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
arxiv.org/abs/2401.06209
[00:26:13] A Global Geometric Analysis of Maximal Coding Rate Reduction
arxiv.org/pdf/2406.01909
[00:47:26] CRATE: White-Box Transformers via Sparse Rate Reduction
arxiv.org/abs/2306.01129
[00:55:05] DINOv2: Learning Robust Visual Features without Supervision
arxiv.org/abs/2304.07193
[01:00:36] An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (ViT)
arxiv.org/abs/2010.11929
What if everything we think we know about AI understanding is wrong? Is compression the key to intelligence? Or is there something more—a leap from memorization to true abstraction?
In this fascinating conversation, we sit down with *Professor Yi Ma*—world-renowned expert in deep learning, IEEE/ACM Fellow, and author of the groundbreaking new book *Learning Deep Representations of Data Distributions*. Professor Ma challenges our assumptions about what large language models actually do, reveals why 3D reconstruction isn't the same as understanding, and presents a unified mathematical theory of intelligence built on just two principles: *parsimony* and *self-consistency*.
**SPONSOR MESSAGES START**
—
Prolific - Quality data. From real people. For faster breakthroughs.
prolific.com/?utm_source=mlst
—
cyber•Fund https://cyber.fund/?utm_source=mlst is a founder-led investment firm accelerating the cybernetic economy
Hiring a SF VC Principal: https://talent.cyber.fund/companies/cyber-fund-2/jobs/57674170-ai-investment-principal#content?utm_source=mlst
Submit investment deck: https://cyber.fund/contact?utm_source=mlst
—
**END**
NOTE: Based on lower retention we have EDITED out 33:45 of the original interview from near the start - philosophical stuff, an extended analogy comparing DNA evolution to AI development, speculation about where theories come from (Platonism, deductive trees), Norbert Wiener's cybernetics history. You can watch the original version on the rescript link or Spotify version.
Key Insights:
*LLMs Don't Understand—They Memorize*
Language models process text (*already* compressed human knowledge) using the same mechanism we use to learn from raw data.
*The Illusion of 3D Vision*
Sora and NeRFs etc that can reconstruct 3D scenes still fail miserably at basic spatial reasoning
*"All Roads Lead to Rome"*
Why adding noise is *necessary* for discovering structure.
*Why Gradient Descent Actually Works*
Natural optimization landscapes are surprisingly smooth—a "blessing of dimensionality"
*Transformers from First Principles*
Transformer architectures can be mathematically derived from compression principles
—
INTERACTIVE AI TRANSCRIPT PLAYER w/REFS (ReScript):
app.rescript.info/public/share/Z-dMPiUhXaeMEcdeU6Bz84GOVsvdcfxU_8Ptu6CTKMQ
About Professor Yi Ma
Yi Ma is the inaugural director of the School of Computing and Data Science at Hong Kong University and a visiting professor at UC Berkeley.
https://people.eecs.berkeley.edu/~yima/
scholar.google.com/citations?user=XqLiBQMAAAAJ&hl=en
https://x.com/YiMaTweets
*Slides from this conversation:*
dropbox.com/scl/fi/sbhbyievw7idup8j06mlr/slides.pdf?rlkey=7ptovemezo8bj8tkhfi393fh9&dl=0
*Related Talks by Professor Ma:*
- Pursuing the Nature of Intelligence (ICLR): youtube.com/watch?v=LT-F0xSNSjo
- Earlier talk at Berkeley: youtube.com/watch?v=TihaCUjyRLM
--
TIMESTAMPS:
00:00:00 Introduction
00:02:08 The First Principles Book & Research Vision
00:05:21 Two Pillars: Parsimony & Consistency
00:09:50 Evolution vs. Learning: The Compression Mechanism
00:14:37 The Illusion of 3D Understanding: Sora & NeRF
00:20:41 All Roads Lead to Rome: The Role of Noise
00:26:11 All Roads Lead to Rome: The Role of Noise
00:26:29 Benign Non-Convexity: Why Optimization Works
00:32:50 Double Descent & The Myth of Overfitting
00:40:41 Self-Consistency: Closed-Loop Learning
00:47:18 Deriving Transformers from First Principles
00:56:26 Verification & The Kevin Murphy Question
01:00:26 CRATE vs. ViT: White-Box AI & Conclusion
---
REFERENCES:
Book:
[00:03:04] Learning Deep Representations of Data Distributions
ma-lab-berkeley.github.io/deep-representation-learning-book
Book (Yi Ma):
[00:03:14] An Invitation to 3-D Vision
link.springer.com/book/10.1007/978-0-387-21779-6
[00:03:24] Generalized Principal Component Analysis
link.springer.com/book/10.1007/978-0-387-87811-9
[00:03:34] High-Dimensional Data Analysis with Low-Dimensional Models
book-wright-ma.github.io
Slide:
[00:44:11] Slide 26: Neuroscience Evidence
arxiv.org/abs/2207.04630)
Person:
[00:08:24] Albert Einstein
quoteinvestigator.com/2011/05/13/einstein-simple
[00:56:41] Kevin Murphy
probml.github.io/pml-book/book1.html
Paper:
[00:18:09] Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
arxiv.org/abs/2401.06209
[00:26:13] A Global Geometric Analysis of Maximal Coding Rate Reduction
arxiv.org/pdf/2406.01909
[00:47:26] CRATE: White-Box Transformers via Sparse Rate Reduction
arxiv.org/abs/2306.01129
[00:55:05] DINOv2: Learning Robust Visual Features without Supervision
arxiv.org/abs/2304.07193
[01:00:36] An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (ViT)
arxiv.org/abs/2010.11929
![Why Humans Are Still Powering AI [Sponsored] - Phelim Bradley
Ever wonder where AI models actually get their intelligence? We reveal the dirty secret of Silicon Valley: behind every impressive AI system are thousands of real humans providing crucial data, feedback, and expertise.
Guest: Phelim Bradley, CEO and Co-founder of Prolific
Phelim Bradley runs Prolific, a platform that connects AI companies with verified human experts who help train and evaluate their models. Think of it as a sophisticated marketplace matching the right human expertise to the right AI task - whether thats doctors evaluating medical chatbots or coders reviewing AI-generated software.
Prolific: https://prolific.com/?utm_source=mlst
https://uk.linkedin.com/in/phelim-bradley-84300826
INTERACTIVE TRANSCRIPT:
https://app.rescript.info/public/share/8wsXTkYiczP2rZyc2_I2aylvaxJ-6QJQ2VTu_9woJyQ
The discussion dives into:
**The human data pipeline**: How AI companies rely on human intelligence to train, refine, and validate their models - something rarely discussed openly
**Quality over quantity**: Why paying humans well and treating them as partners (not commodities) produces better AI training data
**The matching challenge**: How Prolific solves the complex problem of finding the right expert for each specific task, similar to matching Uber drivers to riders but with deep expertise requirements
**Future of work**: What it means when human expertise becomes an on-demand service, and why this might actually create more opportunities rather than fewer
**Geopolitical implications**: Why the centralization of AI development in US tech companies should concern Europe and the UK Why Humans Are Still Powering AI [Sponsored] - Phelim Bradley](https://i.ytimg.com/vi/R11ESdfVX64/mqdefault.jpg)

![Optimize GPU performance for AI - Prof. Gennady Pekhimenko
Prof. Gennady Pekhimenko - CEO of CentML joins us in this *sponsored episode* about AI system optimization and enterprise implementation of AI. From NVIDIAs technical leadership model to the rise of open-source AI, Pekhimenko bridges the gap between academic research and industrial applications. Learn about dark silicon, GPU utilization challenges in ML workloads, and how modern enterprises can optimize their AI infrastructure. The conversation explores why some companies achieve only ~10% GPU efficiency and practical solutions for improving AI system performance. A must-watch for anyone interested in the technical foundations of enterprise AI and GPU hardware optimization.
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments. Cheaper, faster, no commitments, pay as you go, scale massively, simple to setup. Check it out!
https://centml.ai/pricing/
SPONSOR MESSAGES:
MLST is also sponsored by Tufa AI Labs - https://tufalabs.ai/
They are hiring cracked ML engineers/researchers to work on ARC and build AGI!
SHOWNOTES (diarised transcript, TOC, references, summary, best quotes etc)
https://www.dropbox.com/scl/fi/w9kbpso7fawtm286kkp6j/Gennady.pdf?rlkey=aqjqmncx3kjnatk2il1gbgknk&st=2a9mccj8&dl=0
TOC:
1. AI Strategy and Leadership
[00:00:00] 1.1 Technical Leadership and Corporate Structure
[00:09:55] 1.2 Open Source vs Proprietary AI Models
[00:16:04] 1.3 Hardware and System Architecture Challenges
[00:23:37] 1.4 Enterprise AI Implementation and Optimization
[00:35:30] 1.5 AI Reasoning Capabilities and Limitations
2. AI System Development
[00:38:45] 2.1 Computational and Cognitive Limitations of AI Systems
[00:42:40] 2.2 Human-LLM Communication Adaptation and Patterns
[00:46:18] 2.3 AI-Assisted Software Development Challenges
[00:47:55] 2.4 Future of Software Engineering Careers in AI Era
[00:49:49] 2.5 Enterprise AI Adoption Challenges and Implementation
3. ML Infrastructure Optimization
[00:54:41] 3.1 MLOps Evolution and Platform Centralization
[00:55:43] 3.2 Hardware Optimization and Performance Constraints
[01:05:24] 3.3 ML Compiler Optimization and Python Performance
[01:15:57] 3.4 Enterprise ML Deployment and Cloud Provider Partnerships
4. Distributed AI Architecture
[01:27:05] 4.1 Multi-Cloud ML Infrastructure and Optimization
[01:29:45] 4.2 AI Agent Systems and Production Readiness
[01:32:00] 4.3 RAG Implementation and Fine-Tuning Considerations
[01:33:45] 4.4 Distributed AI Systems Architecture and Ray Framework
5. AI Industry Standards and Research
[01:37:55] 5.1 Origins and Evolution of MLPerf Benchmarking
[01:43:15] 5.2 MLPerf Methodology and Industry Impact
[01:50:17] 5.3 Academic Research vs Industry Implementation in AI
[01:58:59] 5.4 AI Research History and Safety Concerns Optimize GPU performance for AI - Prof. Gennady Pekhimenko](https://i.ytimg.com/vi/RvVCyCmsCjg/mqdefault.jpg)
![Your Brain Doesnt Command Your Body. It Predicts It. [Max Bennett]
Tim sits down with Max Bennett to explore how our brains evolved over 600 million years—and what that means for understanding both human intelligence and AI.
Max isnt a neuroscientist by training. Hes a tech entrepreneur who got curious, started reading, and ended up weaving together three fields that rarely talk to each other: comparative psychology (what different animals can actually do), evolutionary neuroscience (how brains changed over time), and AI (what actually works in practice).
*Your Brain Is a Guessing Machine*
You dont actually see the world. Your brain builds a simulation of what it *thinks* is out there and just uses your eyes to check if its right. Thats why optical illusions work—your brain is filling in a triangle that isnt there, or cant decide if its looking at a duck or a rabbit.
*Rats Have Regrets*
In a fascinating experiment called Restaurant Row, rats make choices about waiting for food. When they skip a short wait for something they like and end up stuck with a long wait for something they dont—you can literally watch their brain imagine eating the food they passed up. They regret their choice and make different decisions next time.
*Chimps Are Machiavellian*
The most gripping story is about two chimps, Rock and Belle. Belle learns where food is hidden. Rock figures out he can just follow her and steal it. So Belle starts hiding the food when she finds it. Then Rock starts *pretending* not to watch her, then sprinting to grab the food once she moves. This escalates into an arms race of deception and counter-deception—proof that apes can think about what others are thinking.
*Language Is the Human Superpower*
Other animals learn by watching each others actions. Humans can share whats happening *inside our minds*. You can describe a dream, plan a hunt with five other people, or warn someone about a snake you saw yesterday. This ability to share mental simulations is what lets knowledge accumulate across generations—and its arguably the singularity that already happened.
*Does ChatGPT Think?*
ChatGPT clearly has *a model* (it wouldnt work otherwise), but it doesnt have a *world model* in the way brains do. A real world model means you can form a hypothesis, test it, and update your beliefs based on what happens. GPT learns only from its training data—it cant run experiments or reject information it knows to be false.
Understanding how the brain evolved isnt just about the past. It gives us clues about:
- Whats actually different between human intelligence and AI
- Why were so easily fooled by status games and tribal thinking
- What features we might want to build into—or leave out of—future AI systems
Get Maxs book:
https://www.amazon.com/Brief-History-Intelligence-Humans-Breakthroughs/dp/0063286343
Rescript: https://app.rescript.info/public/share/R234b7AXyDXZusqQ_43KMGsUSvJ2TpSz2I3emnI6j9A
TIMESTAMPS:
00:00:00 Introduction: Outsiders Advantage & Neocortex Theories
00:11:34 Perception as Inference: The Filling-In Machine
00:19:11 Understanding, Recognition & Generative Models
00:36:39 How Mice Plan: Vicarious Trial & Error
00:46:15 Evolution of Self: The Layer 4 Mystery
00:58:31 Ancient Minds & The Social Brain: Machiavellian Apes
01:19:36 AI Alignment, Instrumental Convergence & Status Games
01:33:07 Metacognition & The IQ Paradox
01:48:40 Does GPT Have Theory of Mind?
02:00:40 Memes, Language Singularity & Brain Size Myths
02:16:44 Communication, Language & The Cyborg Future
02:44:25 Shared Fictions, World Models & The Reality Gap
REFERENCES:Person:
[00:00:05] Karl Friston (UCL)
https://www.youtube.com/watch?v=PNYWi996Beg
[00:00:06] Jeff Hawkins
https://www.youtube.com/watch?v=6VQILbDqaI4
[00:12:19] Hermann von Helmholtz
https://plato.stanford.edu/entries/hermann-helmholtz/
[00:38:34] David Redish (U. Minnesota)
https://redishlab.umn.edu/
[01:10:19] Robin Dunbar
https://www.psy.ox.ac.uk/people/robin-dunbar
[01:15:04] Emil Menzel
https://www.sciencedirect.com/bookseries/behavior-of-nonhuman-primates/vol/5/suppl/C
[01:19:49] Nick Bostrom
https://nickbostrom.com/
[02:28:25] Noam Chomsky
https://linguistics.mit.edu/user/chomsky/
[03:01:22] Judea Pearl
https://samueli.ucla.edu/people/judea-pearl/
Concept/Framework:
[00:05:04] Active Inference
https://www.youtube.com/watch?v=KkR24ieh5Ow
Paper:
[00:35:59] Predictions not commands [Rick A Adams]
https://pubmed.ncbi.nlm.nih.gov/23129312/
Book:
[01:25:42] The Elephant in the Brain
https://www.amazon.com/Elephant-Brain-Hidden-Motives-Everyday/dp/0190495995
[01:28:27] The Status Game
https://www.goodreads.com/book/show/58642436-the-status-game
[02:00:40] The Selfish Gene
https://amazon.com/dp/0198788606
[02:14:25] The Language Game
https://www.amazon.com/Language-Game-Improvisation-Created-Changed/dp/1541674987
[02:54:40] The Evolution of Language
https://www.amazon.com/Evolution-Language-Approaches/dp/052167736X
[03:09:37] The Three-Body Problem
https://amazon.com/dp/0765377063 Your Brain Doesnt Command Your Body. It Predicts It. [Max Bennett]](https://i.ytimg.com/vi/RvYSsi6rd4g/mqdefault.jpg)


![Build Specialist LLMs Like It’s 2019 (Randall Balestriero)
Randall Balestriero (Meta AI) shares three recent results that each push back on conventional wisdom in ML.
First, the headline finding: if you take a 7-billion-parameter language model, initialize it randomly, and train it from scratch on just 20,000 labeled examples for a classification task like sentiment analysis, it works. Stable training curves, minimal overfitting, performance that matches LoRA-finetuned pre-trained models. The obvious question is months of expensive pre-training on internet-scale data actually worth it? gets a surprisingly qualified answer. For narrow discriminative tasks, random initialization is competitive. Pre-training still wins for generation and open-ended reasoning, but there is a whole spectrum between the two extremes that nobody is really exploring yet.
Second, a theoretical result with Yann LeCun proving that self-supervised and supervised learning objectives are mathematically equivalent up to how you define the label structure. SSL does not learn better representations because of its loss function; it learns them because it uses finer-grained pairwise relationships instead of collapsing all cars into car. This equivalence lets you port decades of supervised learning theory class imbalance corrections, neural collapse results, semi-supervised weighting directly into SSL, and Randall walks through how VICReg falls out naturally from a least-squares supervised objective under this framework.
Third, a fairness audit of implicit neural representations used for earth/climate data. Models that look accurate on average turn out to be nearly random around islands and coastlines exactly the places where policy decisions about climate adaptation matter most. The culprit is partly architectural: Fourier bases assume stationarity, and switching to wavelets recovers some of the lost localization. But the deeper problem is data bias, including the same geographic skew Mark Ibrahim documented in ImageNet, where most training images come from North America.
SPONSOR MESSAGES:
***
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich.
Goto https://tufalabs.ai/
***
TIMESTAMPS:
00:00:00 Random Initialization Rivals Pre-Training
00:01:29 Is Next-Token Prediction Worth the Cost?
00:04:44 What Do These Models Actually Learn?
00:07:59 Build Specialist LLMs Like It Is 2019
00:10:31 The Fair Language Model Paradox
00:13:38 Benchmarks, Generation, and Understanding
00:16:04 The Birth of Self-Supervised Learning
00:19:14 Class Balance, VICReg, and Unifying Representation Learning
00:25:18 No Location Left Behind: Fairness in Earth Models
00:30:24 Policy, Accountability, and Crowdsourced Data Bias
REFERENCES:
[00:00:00] Is LLM Pre-Training by Next Token Prediction Worth the Cost? https://sslneurips24.github.io/
[00:05:46] Lottery Ticket Hypothesis https://arxiv.org/abs/1803.03635
[00:10:31] The Fair Language Model Paradox https://arxiv.org/abs/2410.11985
[00:16:04] The Birth of Self-Supervised Learning https://openreview.net/forum?id=NhYAjAAdQT
[00:19:14] VICReg https://arxiv.org/abs/2105.04906
[00:25:18] No Location Left Behind https://arxiv.org/abs/2502.06831
[00:33:14] Geographic bias in large visual models https://arxiv.org/abs/2304.12210
LINKS:
Full Transcript: https://app.rescript.info/share/1fecffe43479a465c6b19622356faf8f
Download PDF transcript: https://app.rescript.info/api/public/sessions/8f1ca777a45ad475/pdf Build Specialist LLMs Like It’s 2019 (Randall Balestriero)](https://i.ytimg.com/vi/SP-kORMUZns/mqdefault.jpg)


![Math vs AI: Who Decides Whats True? [Dr. Paul Lessard]
In this episode, hosts Tim Scarfe and Keith Duggar welcome guest Paul Lessard, a mathematician who has transitioned into the world of machine learning, for a deep dive into the philosophy behind AI, mathematics, and the quest for true understanding.
INTERACTIVE TRANSCRIPT PLAYER:
https://app.rescript.info/public/share/_TAiM5iOOePzOGIqCGf69tASjP9bAHNf0tIUKanYIpY
They start by exploring a classic philosophical question: Is the universe built on fundamental, unchanging truths that we discover (a Platonic view), or is it more like were constantly building and creating structure as we go (a constructivist view)? Paul suggests a middle ground, arguing that while the world may be fundamentally constructive, we create the illusions of Platonism as a powerful problem-solving strategy.
This leads to a discussion about the nature of modern AI models. Tim introduces a powerful metaphor, describing deep learning models as sandcastles—structures that look impressive but lack a solid foundation and collapse easily when prodded. Paul challenges this, suggesting there is an emerging science to it, pointing to how benchmarks have historically been used to judge progress, though this method is now showing its limits.
So, how can we build more robust models? Keith asks how the highly abstract field of category theory can help. Paul explains it not as a specific tool, but as a powerful algebra for constructing systems. It provides a formal language to design and experiment with different model architectures in a principled way. He also draws an analogy between transformers and RNNs, framing a transformer as a parallelized, finite-depth version of an RNN.
The conversation then shifts to the human side of science and learning.
Culture Shock in Academia: Paul humorously contrasts the incredibly cautious and boring titles of pure math papers with the bombastic and authoritative titles common in machine learning.
The Walled Garden of Education: Keith shares a relatable story about the shock of discovering that, unlike school textbook problems, most real-world scientific problems dont have a neat, clean solution. Paul explains this is by design—education creates a walled garden to build a students confidence before they face the messy, unpredictable nature of true research.
The episode concludes with Paul sharing his current, overarching view of his work. He sees machine learning as the task of designing a fake physics. The goal is to build a system where the training process acts like a natural physical process, causing the model to settle into a low-energy state that effectively represents the data it was shown.
Position: Categorical Deep Learning is an Algebraic Theory of All Architectures
https://arxiv.org/abs/2402.15332
Bruno Gavranović, Paul Lessard, Andrew Dudzik, Tamara von Glehn, João G. M. Araújo, Petar Veličković
Paul Lessard:
https://www.linkedin.com/in/paul-roy-lessard/?originalSubdomain=au
TOC:
[00:00:00] Truth, Benchmarks, and Sandcastles
[00:00:45] Platonism vs. Constructivism
[00:05:00] The Role of Category Theory
[00:08:00] The Anything Goes Science
[00:12:50] Explaining Why Things Work
[00:16:56] Bombastic Academic Paper Titles
[00:18:18] Automatically Discovering Constraints
[00:29:17] The Walled Garden of Education
[00:35:26] From Math to Machine Learning
[00:43:47] Machine Learning as Fake Physics Math vs AI: Who Decides Whats True? [Dr. Paul Lessard]](https://i.ytimg.com/vi/TBjCvB_4mdo/mqdefault.jpg)
![Exploring Program Synthesis: Francois Chollet, Kevin Ellis, Zenna Tavares
Panel discussion with Francois Chollet, Kevin Ellis, and Zenna Tavares on why program synthesis matters and where deep learning falls short. Chollet recounts how his early work on theorem proving with Christian Szegedy at Google made him realise gradient descent cannot learn discrete algorithms, even when the correct solution is representable by the network. Ellis, whose PhD with Armando Solar-Lezama helped shape the modern program synthesis field, asks how much of the bottleneck is the learning mechanism versus the representation. Tavares considers a deeper integration of neural networks into programming language semantics, where neural operators implement the interpreter rather than sitting outside it.
The group discusses the limits of transformers at function composition, the failure of Cyc-style hand-built ontologies, and what ARC has revealed about strong generalisation. Chollet explains how test-time training and O1-style iterative program writing let static models adapt to novelty, then previews ARC 2, which will include human difficulty data and push harder on compositional complexity. Ellis and Tavares describe MARA, their new project that extends ARC-style tasks toward active experimentation where the agent must choose what questions to ask.
Recorded as part of a broader discussion at the intersection of program synthesis, neural-symbolic integration, and abstract reasoning. Published March 2025.
SPONSOR MESSAGES:
***
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich.
REFERENCES:
website:
[00:00:01] Basis Research Institute
https://www.basis.ai/
[00:14:30] Keras
https://keras.io/
[00:14:50] Armando Solar-Lezama
https://www.csail.mit.edu/news/solar-lezama-wins-robin-milner-young-researcher-award
[00:15:05] Kevin Ellis
https://www.cs.cornell.edu/~ellisk/
[00:18:50] Cyc Project
https://en.wikipedia.org/wiki/Cyc
[00:28:10] ARC Prize
https://arcprize.org/
paper:
[00:01:00] HolStep Dataset
https://openreview.net/pdf?id=ryuxYmvel
[00:05:20] On the Measure of Intelligence
https://arxiv.org/abs/1911.01547
[00:07:00] Neural Turing Machines
https://arxiv.org/pdf/1410.5401
[00:07:20] Manifold Hypothesis
https://arxiv.org/abs/2208.05314
[00:21:50] Test-Time Training on Nearest Neighbors
https://ekinakyurek.github.io/papers/ttt.pdf
[00:26:20] AlphaZero-style Program Synthesis
https://arxiv.org/abs/2205.14229
LINKS:
Full Transcript: https://app.rescript.info/share/b8e9612724c01a88ef103804be1e79d5
Download PDF transcript: https://app.rescript.info/api/public/sessions/3380bd2c998bf22d/pdf
Francois Chollet:
https://x.com/fchollet
https://ndea.com/
https://arcprize.org/
[00:21:55] Test-Time Training, Akyurek et al.
https://ekinakyurek.github.io/papers/ttt.pdf Exploring Program Synthesis: Francois Chollet, Kevin Ellis, Zenna Tavares](https://i.ytimg.com/vi/TQDCsyuuwsg/mqdefault.jpg)