Moving Beyond Surface Statistics (Apple researcher) [Iman Mirzadeh] @MachineLearningStreetTalk
Moving Beyond Surface Statistics (Apple researcher) [Iman Mirzadeh]  @MachineLearningStreetTalk
Uploaded March 2025 | Updated September 2026, 1 week ago
Iman Mirzadeh is a machine learning research engineer at Apple and the lead author of the GSM-Symbolic paper, which exposed deep fragility in how large language models handle mathematical reasoning. In this conversation, he draws a sharp line between intelligence and achievement -- between what a system can score on a benchmark and what it actually understands.

The discussion starts with chess. Mirzadeh explains how grandmasters don't use engines to memorize moves; they use them to develop theory. AlphaZero discovered unprecedented strategies, but that knowledge stays trapped in the game. Humans, by contrast, extract abstract principles like 'control the center' and transfer them to entirely different domains. That capacity for abstraction is what he thinks current AI architecturally lacks.

His critique of LLMs is structural. These systems are trained to minimize cross-entropy loss over a distribution, and by construction they cannot reason beyond what that distribution contains. Change the surface form of a problem -- swap names, add irrelevant clauses -- and performance varies wildly, even on grade-school math.

That's the core finding of GSM-Symbolic. By generating templated variants of math word problems, Mirzadeh's team showed that even frontier models exhibit large performance variance from changes that should be semantically irrelevant. The implication: what looks like reasoning is closer to sophisticated pattern matching across memorized distributions.

Mirzadeh proposes that intelligence should be measured by the slope of a system's scaling -- how fast it can learn novel things -- rather than its current benchmark position. The conversation also covers the connectionism-symbolism divide, active engagement and agency as necessary conditions for learning, and why we might need a fundamentally different vessel to reach genuine reasoning.

SPONSOR MESSAGES:
***
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich.

---
TIMESTAMPS:
00:00:00 Intelligence vs Achievement in AI
00:03:27 AlphaZero and Abstract Understanding in Chess
00:10:10 Language Models as Distribution Learners
00:14:47 The State of AI Research Methodology
00:24:24 Interpolation vs True Reasoning in LLMs
00:29:00 Measuring Intelligence: From Chollet to the Iman Moon Test
00:35:35 Agency, Active Learning, and World Models
00:47:15 Scaling Laws and the Connectionism-Symbolism Debate
00:58:09 GSM-Symbolic: Exposing LLM Reasoning Fragility

---
REFERENCES:
paper:
[00:00:55] Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
arxiv.org/abs/1712.01815
[00:17:15] GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
arxiv.org/abs/2410.05229
[00:21:20] Connectionism and Cognitive Architecture: A Critical Analysis
sciencedirect.com/science/article/pii/001002779090014B
[00:29:35] On the Measure of Intelligence
arxiv.org/abs/1911.01547
[00:33:25] On definition of intelligence
sciencedirect.com/science/article/pii/S0160289624000266
[00:35:25] Defining Intelligence
https://cis.temple.edu/~wangp/papers.html
[00:43:10] Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
arxiv.org/abs/2201.11903
[00:47:45] Scaling Laws for Neural Language Models
arxiv.org/abs/2001.08361
[00:55:10] Tensor Product Variable Binding and the Representation of Symbolic Structures in Connectionist Systems
sciencedirect.com/science/article/abs/pii/000437029090007M
book:
[00:07:05] Game Changer: AlphaZero's Groundbreaking Chess Strategies
amazon.com/Game-Changer-AlphaZeros-Groundbreaking-Strategies/dp/9056918184
[00:37:35] How We Learn: Why Brains Learn Better Than Any Machine... for Now
amazon.com/How-We-Learn-Brains-Machine/dp/0525559884
[00:39:30] Surfaces and Essences: Analogy as the Fuel and Fire of Thinking
amazon.com/Surfaces-Essences-Analogy-Fuel-Thinking/dp/0465018475
reference:
[00:11:30] NLP Course: Language Modeling
lena-voita.github.io/nlp_course/language_modeling.html
[01:08:40] GSM8K: Training Verifiers to Solve Math Word Problems
huggingface.co/datasets/openai/gsm8k

---
LINKS:
Full Transcript: app.rescript.info/share/72689965572c7fd460954f70e26b2eaa
Download PDF transcript: app.rescript.info/api/public/sessions/abf4f57b59fcbbd5/pdf
Moving Beyond Surface Statistics (Apple researcher) [Iman Mirzadeh]Pioneer Yoshua Bengio on AI agency #aiWhy AI Has a Plato Problem — Mazviita ChirimuutaThis Tiny Code Made Artificial Life! (Blaise Agüera y Arcas)Solving Chollets ARC-AGI with GPT4oA little bit of love from the main man ♥️The Free Energy Principle approach to AgencyThe AI Progress Chart Everyone Is Misreading — Beth Barnes & David ReinGenuine understanding is the worlds most valuable commodity right nowThere are monsters in your LLM. (Murray Shanahan)We need AIs with PHYSICAL experience (Jeff Beck)
Machine Learning Street Talk |

Moving Beyond Surface Statistics (Apple researcher) [Iman Mirzadeh]

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER