Uploaded August 2026 | Updated September 2026, 1 week ago
This episode is sponsored by Notion. Learn more about Notion's Developer Platform today at notion.com/mlst
Why can deep networks discover abstractions that shallow models miss? Statistical physicist Matthieu Wyart joins Tim Scarfe to argue that the answer lies in the hidden hierarchy of data. Language and images are built from parts within parts; depth lets a network recover those coarse-grained variables and escape the curse of dimensionality.
The conversation moves from jamming transitions and rough loss surfaces to Chomsky, context-free grammars and machine creativity. Wyart explains why next-token prediction can still recover compositional structure, where current systems fall short of genuine scientific invention, and why predicting latent representations rather than raw tokens could make learning far more sample-efficient.
They also examine diffusion models, neural scaling laws and the limits of physics-inspired theory. The final question is on a personal note: if mistakes are the price of leaving the beaten path, how much scientific risk is worth taking?
---
TIMESTAMPS:
00:00:00 Can machines learn abstractions from data?
00:02:00 Notion agentic workspace
00:02:49 From statistical physics to machine learning
00:06:40 What physics can explain about learning
00:16:37 From Carnot to Chomsky bulldozer
00:21:21 How deep networks recover hidden hierarchies
00:32:43 Where machine creativity still falls short
00:40:48 How deep nets escape the curse of dimensionality
00:52:19 Why predict latents instead of tokens
01:02:49 The sample-efficiency case for latent prediction
01:08:31 Diffusion, scaling laws and text entropy
01:16:40 The scientists we learn from and the mistakes we make
---
REFERENCES:
person:
[00:00:43] Noam Chomsky
https://linguistics.mit.edu/user/chomsky/
tool:
[00:02:08] Notion Developer Platform
notion.com/en-gb/blog/introducing-developer-platform
paper:
[00:04:43] Mastering the game of Go with deep neural networks and tree search
nature.com/articles/nature16961
[00:05:52] Reconciling modern machine-learning practice and the bias-variance trade-off
arxiv.org/abs/1812.11118
[00:25:54] How Deep Neural Networks Learn Compositional Data: The Random Hierarchy Model
arxiv.org/abs/2307.02129
[00:42:12] Efficient Estimation of Word Representations in Vector Space
arxiv.org/abs/1301.3781
[00:52:46] Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
arxiv.org/abs/2301.08243
[00:52:54] Learn from your own latents and not from tokens: A sample-complexity theory
arxiv.org/abs/2605.27734
[01:08:31] A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Data
arxiv.org/abs/2402.16991
[01:11:39] Scaling Laws for Neural Language Models
arxiv.org/abs/2001.08361
[01:12:17] Deriving Neural Scaling Laws from the statistics of natural language
arxiv.org/abs/2602.07488
[01:13:34] Prediction and Entropy of Printed English
ieeexplore.ieee.org/document/6773263
---
LINKS:
Download PDF transcript: app.rescript.info/share/f7644cdaa86c5cc1e41e484e290f2bd4
This episode is sponsored by Notion. Learn more about Notion's Developer Platform today at notion.com/mlst
Why can deep networks discover abstractions that shallow models miss? Statistical physicist Matthieu Wyart joins Tim Scarfe to argue that the answer lies in the hidden hierarchy of data. Language and images are built from parts within parts; depth lets a network recover those coarse-grained variables and escape the curse of dimensionality.
The conversation moves from jamming transitions and rough loss surfaces to Chomsky, context-free grammars and machine creativity. Wyart explains why next-token prediction can still recover compositional structure, where current systems fall short of genuine scientific invention, and why predicting latent representations rather than raw tokens could make learning far more sample-efficient.
They also examine diffusion models, neural scaling laws and the limits of physics-inspired theory. The final question is on a personal note: if mistakes are the price of leaving the beaten path, how much scientific risk is worth taking?
---
TIMESTAMPS:
00:00:00 Can machines learn abstractions from data?
00:02:00 Notion agentic workspace
00:02:49 From statistical physics to machine learning
00:06:40 What physics can explain about learning
00:16:37 From Carnot to Chomsky bulldozer
00:21:21 How deep networks recover hidden hierarchies
00:32:43 Where machine creativity still falls short
00:40:48 How deep nets escape the curse of dimensionality
00:52:19 Why predict latents instead of tokens
01:02:49 The sample-efficiency case for latent prediction
01:08:31 Diffusion, scaling laws and text entropy
01:16:40 The scientists we learn from and the mistakes we make
---
REFERENCES:
person:
[00:00:43] Noam Chomsky
https://linguistics.mit.edu/user/chomsky/
tool:
[00:02:08] Notion Developer Platform
notion.com/en-gb/blog/introducing-developer-platform
paper:
[00:04:43] Mastering the game of Go with deep neural networks and tree search
nature.com/articles/nature16961
[00:05:52] Reconciling modern machine-learning practice and the bias-variance trade-off
arxiv.org/abs/1812.11118
[00:25:54] How Deep Neural Networks Learn Compositional Data: The Random Hierarchy Model
arxiv.org/abs/2307.02129
[00:42:12] Efficient Estimation of Word Representations in Vector Space
arxiv.org/abs/1301.3781
[00:52:46] Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
arxiv.org/abs/2301.08243
[00:52:54] Learn from your own latents and not from tokens: A sample-complexity theory
arxiv.org/abs/2605.27734
[01:08:31] A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Data
arxiv.org/abs/2402.16991
[01:11:39] Scaling Laws for Neural Language Models
arxiv.org/abs/2001.08361
[01:12:17] Deriving Neural Scaling Laws from the statistics of natural language
arxiv.org/abs/2602.07488
[01:13:34] Prediction and Entropy of Printed English
ieeexplore.ieee.org/document/6773263
---
LINKS:
Download PDF transcript: app.rescript.info/share/f7644cdaa86c5cc1e41e484e290f2bd4

![Why High Benchmark Scores Don’t Mean Better AI [SPONSORED]
Is a car that wins a Formula 1 race the best choice for your morning commute? Probably not. In this sponsored deep dive with Prolific, we explore why the same logic applies to Artificial Intelligence. While models are currently shattering records on technical exams, they often fail the most important test of all: *the human experience.*
Why High Benchmark Scores Don’t Mean Better AI
Joining us are *Andrew Gordon* (Staff Researcher in Behavioral Science) and *Nora Petrova* (AI Researcher) from *Prolific* . They reveal the hidden flaws in how we currently rank AI and introduce a more rigorous, humane way to measure whether these models are actually helpful, safe, and relatable for real people.
Key Insights in This Episode:
* *The F1 Car Analogy:* Andrew explains why a model that excels at the Humanities Last Exam might be a nightmare for daily use. Technical benchmarks often ignore the nuances of human communication and adaptability.
* *The Wild West of AI Safety:* As users turn to AI for sensitive topics like mental health, Nora highlights the alarming lack of oversight and the thin veneer of safety training—citing recent controversial incidents like Grok-3’s Mecha Hitler.
* *Fixing the Leaderboard Illusion:* The team critiques current popular rankings like Chatbot Arena, discussing how anonymous, unstratified voting can lead to biased results and how companies can game the system.
* *The Xbox Secret to AI Ranking:* Discover how Prolific uses *TrueSkill* —the same algorithm Microsoft developed for Xbox Live matchmaking—to create a fairer, more statistically sound leaderboard for LLMs.
* *The Personality Gap:* Early data from the *Humane Leaderboard* suggests that while AI is getting smarter, it is actually performing *worse* on metrics like personality, culture, and sycophancy (the tendency for models to become annoying people-pleasers).
About the HUMAINE Leaderboard
Moving beyond simple A vs. B testing, the researchers discuss their new framework that samples participants based on *census data* (Age, Ethnicity, Political Alignment). By using a representative sample of the general public rather than just tech enthusiasts, they are building a standard that reflects the values of the real world.
*Are we building models for benchmarks, or are we building them for humans? It’s time to change the scoreboard.*
Rescript link:
https://app.rescript.info/public/share/IDqwjY9Q43S22qSgL5EkWGFymJwZ3SVxvrfpgHZLXQc
TIMESTAMPS:
00:00:00 Introduction & The Benchmarking Problem
00:01:58 The Fractured State of AI Evaluation
00:03:54 AI Safety & Interpretability
00:05:45 Bias in Chatbot Arena
00:06:45 Prolifics Three Pillars Approach
00:09:01 TrueSkill Ranking & Efficient Sampling
00:12:04 Census-Based Representative Sampling
00:13:00 Key Findings: Culture, Personality & Sycophancy
REFERENCES:
Paper:
[00:00:15] MMLU
https://arxiv.org/abs/2009.03300
[00:05:10] Constitutional AI
https://arxiv.org/abs/2212.08073
[00:06:45] The Leaderboard Illusion
https://arxiv.org/abs/2504.20879
[00:09:41] HUMAINE Framework Paper
https://huggingface.co/blog/ProlificAI/humaine-framework
Company:
[00:00:30] Prolific
https://www.prolific.com
[00:01:45] Chatbot Arena
https://lmarena.ai/
Person:
[00:00:35] Andrew Gordon
https://www.linkedin.com/in/andrew-gordon-03879919a/
[00:00:45] Nora Petrova
https://www.linkedin.com/in/nora-petrova/
Event:
Algorithm:
[00:09:01] Microsoft TrueSkill
https://www.microsoft.com/en-us/research/project/trueskill-ranking-system/
Leaderboard:
[00:09:21] Prolific HUMAINE Leaderboard
https://www.prolific.com/humaine
[00:09:31] HUMAINE HuggingFace Space
https://huggingface.co/spaces/ProlificAI/humaine-leaderboard
[00:10:21] Prolific AI Leaderboard Portal
https://www.prolific.com/leaderboard
Dataset:
[00:09:51] Prolific Social Reasoning RLHF Dataset
https://huggingface.co/datasets/ProlificAI/social-reasoning-rlhf
Organization:
[00:10:31] MLCommons
https://mlcommons.org/ Why High Benchmark Scores Don’t Mean Better AI [SPONSORED]](https://i.ytimg.com/vi/rqiC9a2z8Io/mqdefault.jpg)


![Its Not About Scale, Its About Abstraction
MLST is sponsored by Tufa Labs:
Are you interested in working on ARC and cutting-edge AI research with the MindsAI team (current ARC winners)?
Focus: ARC, LLMs, test-time-compute, active inference, system2 reasoning, and more.
Future plans: Expanding to complex environments like Warcraft 2 and Starcraft 2.
Interested? Apply for an ML research position: benjamin@tufa.ai
Francois Chollet, creator of Keras and the ARC-AGI benchmark, delivers his AGI-24 keynote on why scaling LLMs will not get us to AGI. He walks through concrete failure modes LLMs that break on trivial rephrasing of memorized problems, that pattern-match the Monty Hall problem without parsing the actual numbers, that solve Caesar ciphers only for key sizes found in online examples. The failures all point the same way: LLM performance tracks task familiarity, not task complexity.
Chollet introduces his Kaleidoscope Hypothesis: the world looks infinitely complex on the surface, but it is built from a small set of repeating atoms of meaning. Intelligence, in his framing, is the process of mining experience to extract those atoms and recombining them to handle genuinely novel situations. This is what the ARC benchmark is designed to test abstraction and reasoning that cannot be memorized.
The talk closes with a proposal: combine deep learning (good at perception and pattern recognition) with discrete program synthesis (good at precise, compositional reasoning). Neither approach alone gets there, but the hybrid might. Chollet points to early results on ARC from Ryan Greenblatt and others as evidence that the research community outside big labs may be where the next breakthrough comes from.
TIMESTAMPS:
00:00:00 LLM Limitations and Composition
00:12:05 Intelligence as Process vs. Skill
00:17:15 Generalization as Key to AI Progress
00:19:59 Introduction to ARC-AGI Benchmark
00:26:10 The Kaleidoscope Hypothesis and Abstraction Spectrum
00:34:05 Limitations of Transformers and Program Synthesis
00:39:59 Applying Combined Approaches to ARC Tasks
00:44:20 State-of-the-Art Solutions and Future Directions
REFERENCES:
paper:
[00:01:15] On the Measure of Intelligence
https://arxiv.org/abs/1911.01547
[00:03:30] Embers of Autoregression
https://arxiv.org/abs/2309.13638
[00:05:30] Monty Hall problem
https://www.tandfonline.com/doi/abs/10.1080/00031305.1975.10479121
[00:06:20] LLM Training Dynamics Analysis
https://arxiv.org/abs/2205.10770
[00:07:33] GPT-4 Technical Report
https://cdn.openai.com/papers/gpt-4.pdf
[00:10:20] Faith and Fate: Limits of Transformers on Compositionality
https://arxiv.org/abs/2305.18654
[00:10:25] The Reversal Curse in LLMs
https://arxiv.org/abs/2309.12288
[00:10:52] LM-Polygraph: Uncertainty Estimation for LLMs
https://arxiv.org/abs/2311.07383
[00:20:34] Baldur: Whole-Proof Generation
https://arxiv.org/abs/2303.04910
[00:34:00] Core Knowledge in Infants
https://www.harvardlds.org/wp-content/uploads/2017/01/SpelkeKinzler07-1.pdf
[00:44:20] Hypothesis Search with LLMs for ARC
https://arxiv.org/abs/2309.05660
tool:
[00:20:10] ARC-AGI GitHub Repository
https://github.com/fchollet/ARC-AGI
[00:22:15] ARC Prize
https://arcprize.org/
book:
[00:33:30] Thinking, Fast and Slow
https://www.amazon.com/Thinking-Fast-Slow-Daniel-Kahneman/dp/0374533555
LINKS:
Full Transcript: https://app.rescript.info/share/c8b5bacdf1ffefab4f65060edc295d4a
Download PDF transcript: https://app.rescript.info/api/public/sessions/b537d0b92ae48338/pdf
[0:20:10] ARC-AGI: GitHub repository (François Chollet)
https://github.com/fchollet/ARC-AGI Its Not About Scale, Its About Abstraction](https://i.ytimg.com/vi/s7_NlkBwdj8/mqdefault.jpg)





![Wild breakthrough on Math after 56 years... [Exclusive]
Google DeepMind just dropped AlphaEvolve, a Gemini-powered evolutionary coding agent that designs advanced algorithms by pairing LLM creativity with automated evaluation. The headline result: it beat Volker Strassens 56-year-old record for 4x4 matrix multiplication, finding a method that uses 48 scalar multiplications instead of 49. No human or AI had managed that in over half a century.
Tim sits down with two of the researchers behind the work Matej Balog and Alexander Novikov to walk through the system, the results, and what it means.
In this episode:
- How AlphaEvolve works: an evolutionary pipeline that pairs LLM-generated code proposals with rigorous automated evaluators, iteratively improving solutions rather than relying on one-shot generation. The gap between single-shot LLM sampling and scaled evolutionary search turns out to be enormous.
- The matrix multiplication breakthrough: AlphaTensor tried for years with reinforcement learning and only cracked the Boolean case. AlphaEvolve found a 48-multiplication algorithm for general 4x4 matrices almost by accident, running for completeness. The result generalises from complex to real matrices, which is counterintuitive solving the harder problem actually made the search easier.
- Three ways to represent the search target: direct solution, constructor function, or search algorithm. For matrix multiplication, AlphaEvolve designed a gradient-based search algorithm that finds matrix multiplication algorithms a meta-level approach that produced loss functions and update rules no human would have tried.
- Real-world impact at Google scale: AlphaEvolve recovered 0.7% of fleet-wide compute resources in the Borg data center scheduling system and sped up Gemini training by 1%. These are already-heavily-optimized production systems.
- The human-AI collaboration loop: AlphaEvolve is not autonomous research. Humans choose the problems, design evaluators, seed initial solutions, and interpret results. Alexander Novikov argues this back-and-forth is the whole point the system improves your questions as much as your answers.
- Keith Duggar probes the halting problem and evaluation cascade limitations. The researchers acknowledge the constraint but note that practical framing (time-bounded evaluation, evaluation cascades from cheap to expensive) sidesteps the theoretical issue for now.
- The recursive self-improvement question: AlphaEvolve improved the infrastructure that trains the models that power AlphaEvolve. The feedback loop exists but currently operates on a timescale of months.
SPONSOR MESSAGES:
***
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich.
GTC is coming, the premier AI conference, great opportunity to learn about AI. NVIDIA and partners will showcase breakthroughs in physical AI, AI factories, agentic AI, and inference. Register for virtual GTC for free, using Tims link (https://nvda.ws/4qQ0LMg)
Tufa AI Labs is a new research lab in Zurich started by Benjamin Crouzier focused on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Go to https://tufalabs.ai/
TIMESTAMPS:
00:00:00 Introduction: AlphaEvolves Breakthroughs and DeepMinds Lineage
00:11:24 Introducing AlphaEvolve: Evolutionary Architecture and LLM Pairing
00:16:56 The Halting Problem and Evaluation Constraints
00:23:20 Knowledge Augmentation: Meta-Prompting, Library Learning, and Self-Generated Data
00:29:08 Matrix Multiplication Breakthrough: From Strassen to 48 Multiplications
00:39:11 Problem Representation: Direct Solutions, Constructors, and Search Algorithms
00:46:06 Surprising Outcomes: What Researchers Did Not Expect
00:51:42 Hill Climbing, Program Synthesis, and Intelligibility
01:00:24 Real-World Applications: Complex Evaluations and Robotics
01:05:39 The Role of LLMs, Recursive Self-Improvement, and Future Directions
REFERENCES:
Blog Post:
[00:00:00] AlphaEvolve Blog Post
https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/
Paper:
[00:03:00] FunSearch
https://www.nature.com/articles/s41586-023-06924-6
[00:12:00] MAP-Elites
https://arxiv.org/abs/1504.04909
Person:
[00:11:24] Matej Balog
https://x.com/matejbalog
[00:14:10] Alexander Novikov
https://x.com/SashaVNovikov
Company:
[00:11:24] Tufa AI Labs
https://tufalabs.ai/
LINKS:
Full Transcript: https://app.rescript.info/share/de65b6ec8da83b261ce6018039fec289
Download PDF transcript: https://app.rescript.info/api/public/sessions/855d5b15e733d182/pdf
Guests:
Matej Balog: https://x.com/matejbalog
Alexander Novikov: https://x.com/SashaVNovikov Wild breakthrough on Math after 56 years... [Exclusive]](https://i.ytimg.com/vi/vC9nAosXrJw/mqdefault.jpg)