Uploaded March 2025 | Updated September 2026, 1 week ago
SPONSOR MESSAGES:
***
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich.
Dr. Max Bartolo from Cohere discusses the gap between model capabilities and genuine robustness: why next-token prediction can produce impressive results yet still fail on slightly reformulated…
---
TIMESTAMPS:
00:00:00 Model Reasoning and Consistency Verification
00:03:25 Influence Functions and Distributed Knowledge Analysis
00:10:28 AI Application Development and Model Deployment
00:14:24 AI Alignment and Human Feedback Limitations
00:20:15 Human Evaluation Challenges and Factuality Assessment
00:27:15 Cultural and Demographic Influences on Model Behavior
00:32:43 Adversarial Examples and Model Robustness
00:41:54 DynaBench and Dynamic Benchmarking Approaches
00:50:02 Benchmarking Challenges and Data-Centric Evaluation
00:55:15 Cohere Command A Development Process
01:00:26 Model Quantization and Performance Evaluation
01:05:18 Reasoning Capabilities and Training Progression
01:13:48 Context Windows and Enterprise Applications
---
REFERENCES:
person:
[00:00:00] Max Bartolo Website
maxbartolo.com
company:
[00:00:00] Cohere
cohere.com/command
[00:12:10] Command A Model
huggingface.co/CohereForAI/c4ai-command-a-03-2025
paper:
[00:03:25] Procedural Knowledge in Pretraining Drives Reasoning in LLMs
cohere.com/research/papers/procedural-knowledge-in-pretraining-drives-reasoning-in-large-language-models-2024-11-20
[00:04:15] Influence Functions in Machine Learning
arxiv.org/abs/1703.04730
[00:08:05] Studying Large Language Model Generalization with Influence Functions
arxiv.org/abs/2308.03296
[00:16:15] Human Feedback is not Gold Standard
arxiv.org/abs/2309.16349
[00:27:15] The PRISM Alignment Dataset
arxiv.org/abs/2404.16019
[00:32:50] Adversarial Examples Are Not Bugs, They Are Features
arxiv.org/abs/1905.02175
[00:43:00] DynaBench: Rethinking Benchmarking in NLP
aclanthology.org/2021.naacl-main.324.pdf
[00:50:15] Sara Hooker on Compute Limitations
arxiv.org/html/2407.05694v1
[00:53:25] DataPerf: Benchmarks for Data-Centric AI
arxiv.org/abs/2207.10062
[01:04:35] DROP: A Reading Comprehension Benchmark
arxiv.org/abs/1903.00161
[01:07:05] GSM8k
paperswithcode.com/sota/arithmetic-reasoning-on-gsm8k
[01:09:30] ARC-AGI Challenge
github.com/fchollet/ARC-AGI
---
LINKS:
Full Transcript: app.rescript.info/share/163bf5e7338685f635fc0b8d6920005c
Download PDF transcript: app.rescript.info/api/public/sessions/7167a679366f97f8/pdf
REFS:
[00:03:10] Research at Cohere with Laura Ruis et al., Max Bartolo, Laura Ruis et al.
cohere.com/research/papers/procedural-knowledge-in-pretraining-drives-reasoning-in-large-language-models-2024-11-20
[00:04:15] Influence functions in machine learning, Koh & Liang
arxiv.org/abs/1703.04730
[00:08:05] Studying Large Language Model Generalization with Influence Functions, Roger Grosse et al.
storage.prod.researchhub.com/uploads/papers/2023/08/08/2308.03296.pdf
[00:11:10] The LLM ARChitect: Solving ARC-AGI Is A Matter of Perspective, Daniel Franzen, Jan Disselhoff, and David Hartmann
github.com/da-fr/arc-prize-2024/blob/main/the_architects.pdf
[00:12:10] Hugging Face model repo for C4AI Command A, Cohere and Cohere For AI
huggingface.co/CohereForAI/c4ai-command-a-03-2025
[00:13:30] OpenInterpreter
github.com/KillianLucas/open-interpreter
[00:16:15] Human Feedback is not Gold Standard, Tom Hosking, Max Bartolo, Phil Blunsom
arxiv.org/abs/2309.16349
[00:27:15] The PRISM Alignment Dataset, Hannah Kirk et al.
arxiv.org/abs/2404.16019
[00:32:50] How adversarial examples arise, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, Aleksander Madry
arxiv.org/abs/1905.02175
[00:43:00] DynaBench platform paper, Douwe Kiela et al.
aclanthology.org/2021.naacl-main.324.pdf
[00:50:15] Sara Hooker's work on compute limitations, Sara Hooker
arxiv.org/html/2407.05694v1
[00:53:25] DataPerf: Community-led benchmark suite, Mazumder et al.
arxiv.org/abs/2207.10062
[01:04:35] DROP, Dheeru Dua et al.
arxiv.org/abs/1903.00161
[01:07:05] GSM8k, Cobbe et al.
paperswithcode.com/sota/arithmetic-reasoning-on-gsm8k
[01:09:30] ARC, François Chollet
github.com/fchollet/ARC-AGI
[01:15:50] Command A, Cohere
cohere.com/blog/command-a
[01:22:55] Enterprise search using LLMs, Cohere
cohere.com/blog/commonly-asked-questions-about-search-from-coheres-enterprise-customers
SPONSOR MESSAGES:
***
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich.
Dr. Max Bartolo from Cohere discusses the gap between model capabilities and genuine robustness: why next-token prediction can produce impressive results yet still fail on slightly reformulated…
---
TIMESTAMPS:
00:00:00 Model Reasoning and Consistency Verification
00:03:25 Influence Functions and Distributed Knowledge Analysis
00:10:28 AI Application Development and Model Deployment
00:14:24 AI Alignment and Human Feedback Limitations
00:20:15 Human Evaluation Challenges and Factuality Assessment
00:27:15 Cultural and Demographic Influences on Model Behavior
00:32:43 Adversarial Examples and Model Robustness
00:41:54 DynaBench and Dynamic Benchmarking Approaches
00:50:02 Benchmarking Challenges and Data-Centric Evaluation
00:55:15 Cohere Command A Development Process
01:00:26 Model Quantization and Performance Evaluation
01:05:18 Reasoning Capabilities and Training Progression
01:13:48 Context Windows and Enterprise Applications
---
REFERENCES:
person:
[00:00:00] Max Bartolo Website
maxbartolo.com
company:
[00:00:00] Cohere
cohere.com/command
[00:12:10] Command A Model
huggingface.co/CohereForAI/c4ai-command-a-03-2025
paper:
[00:03:25] Procedural Knowledge in Pretraining Drives Reasoning in LLMs
cohere.com/research/papers/procedural-knowledge-in-pretraining-drives-reasoning-in-large-language-models-2024-11-20
[00:04:15] Influence Functions in Machine Learning
arxiv.org/abs/1703.04730
[00:08:05] Studying Large Language Model Generalization with Influence Functions
arxiv.org/abs/2308.03296
[00:16:15] Human Feedback is not Gold Standard
arxiv.org/abs/2309.16349
[00:27:15] The PRISM Alignment Dataset
arxiv.org/abs/2404.16019
[00:32:50] Adversarial Examples Are Not Bugs, They Are Features
arxiv.org/abs/1905.02175
[00:43:00] DynaBench: Rethinking Benchmarking in NLP
aclanthology.org/2021.naacl-main.324.pdf
[00:50:15] Sara Hooker on Compute Limitations
arxiv.org/html/2407.05694v1
[00:53:25] DataPerf: Benchmarks for Data-Centric AI
arxiv.org/abs/2207.10062
[01:04:35] DROP: A Reading Comprehension Benchmark
arxiv.org/abs/1903.00161
[01:07:05] GSM8k
paperswithcode.com/sota/arithmetic-reasoning-on-gsm8k
[01:09:30] ARC-AGI Challenge
github.com/fchollet/ARC-AGI
---
LINKS:
Full Transcript: app.rescript.info/share/163bf5e7338685f635fc0b8d6920005c
Download PDF transcript: app.rescript.info/api/public/sessions/7167a679366f97f8/pdf
REFS:
[00:03:10] Research at Cohere with Laura Ruis et al., Max Bartolo, Laura Ruis et al.
cohere.com/research/papers/procedural-knowledge-in-pretraining-drives-reasoning-in-large-language-models-2024-11-20
[00:04:15] Influence functions in machine learning, Koh & Liang
arxiv.org/abs/1703.04730
[00:08:05] Studying Large Language Model Generalization with Influence Functions, Roger Grosse et al.
storage.prod.researchhub.com/uploads/papers/2023/08/08/2308.03296.pdf
[00:11:10] The LLM ARChitect: Solving ARC-AGI Is A Matter of Perspective, Daniel Franzen, Jan Disselhoff, and David Hartmann
github.com/da-fr/arc-prize-2024/blob/main/the_architects.pdf
[00:12:10] Hugging Face model repo for C4AI Command A, Cohere and Cohere For AI
huggingface.co/CohereForAI/c4ai-command-a-03-2025
[00:13:30] OpenInterpreter
github.com/KillianLucas/open-interpreter
[00:16:15] Human Feedback is not Gold Standard, Tom Hosking, Max Bartolo, Phil Blunsom
arxiv.org/abs/2309.16349
[00:27:15] The PRISM Alignment Dataset, Hannah Kirk et al.
arxiv.org/abs/2404.16019
[00:32:50] How adversarial examples arise, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, Aleksander Madry
arxiv.org/abs/1905.02175
[00:43:00] DynaBench platform paper, Douwe Kiela et al.
aclanthology.org/2021.naacl-main.324.pdf
[00:50:15] Sara Hooker's work on compute limitations, Sara Hooker
arxiv.org/html/2407.05694v1
[00:53:25] DataPerf: Community-led benchmark suite, Mazumder et al.
arxiv.org/abs/2207.10062
[01:04:35] DROP, Dheeru Dua et al.
arxiv.org/abs/1903.00161
[01:07:05] GSM8k, Cobbe et al.
paperswithcode.com/sota/arithmetic-reasoning-on-gsm8k
[01:09:30] ARC, François Chollet
github.com/fchollet/ARC-AGI
[01:15:50] Command A, Cohere
cohere.com/blog/command-a
[01:22:55] Enterprise search using LLMs, Cohere
cohere.com/blog/commonly-asked-questions-about-search-from-coheres-enterprise-customers


![Manhattan Project for AI Safety [Connor Leahy]
Connor Leahy and Gabriel Alfour, AI researchers from Conjecture and authors of The Compendium, join us for a critical discussion centered on Artificial Superintelligence (ASI) safety and governance. Drawing from their comprehensive analysis in The Compendium, they articulate a stark warning about the existential risks inherent in uncontrolled AI development, framing it through the lens of intelligence domination—where a sufficiently advanced AI could subordinate humanity, much like humans dominate less intelligent species.
SPONSOR MESSAGES:
***
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich.
Goto https://tufalabs.ai/
***
TRANSCRIPT + REFS + NOTES:
https://www.dropbox.com/scl/fi/p86l75y4o2ii40df5t7no/Compendium.pdf?rlkey=tukczgf3flw133sr9rgss0pnj&dl=0
Transcript (new system): https://app.rescript.info/public/share/FObT5GvZYfnixs-aSAD8px4Je_4_aC1kHYG76bvBBRU
https://www.thecompendium.ai/
https://en.wikipedia.org/wiki/Connor_Leahy
https://www.conjecture.dev/about
https://substack.com/@gabecc
TOC:
1. AI Intelligence and Safety Fundamentals
[00:00:00] 1.1 Understanding Intelligence and AI Capabilities
[00:06:20] 1.2 Emergence of Intelligence and Regulatory Challenges
[00:10:18] 1.3 Human vs Animal Intelligence Debate
[00:18:00] 1.4 AI Regulation and Risk Assessment Approaches
[00:26:14] 1.5 Competing AI Development Ideologies
2. Economic and Social Impact
[00:29:10] 2.1 Labor Market Disruption and Post-Scarcity Scenarios
[00:32:40] 2.2 Institutional Frameworks and Tech Power Dynamics
[00:37:40] 2.3 Ethical Frameworks and AI Governance Debates
[00:40:52] 2.4 AI Alignment Evolution and Technical Challenges
3. Technical Governance Framework
[00:55:07] 3.1 Three Levels of AI Safety: Alignment, Corrigibility, and Boundedness
[00:55:30] 3.2 Challenges of AI System Corrigibility and Constitutional Models
[00:57:35] 3.3 Limitations of Current Boundedness Approaches
[00:59:11] 3.4 Abstract Governance Concepts and Policy Solutions
4. Democratic Implementation and Coordination
[00:59:20] 4.1 Governance Design and Measurement Challenges
[01:00:10] 4.2 Democratic Institutions and Experimental Governance
[01:14:10] 4.3 Political Engagement and AI Safety Advocacy
[01:25:30] 4.4 Practical AI Safety Measures and International Coordination
CORE REFS:
[00:01:45] The Compendium (2023), Leahy et al.
https://pdf.thecompendium.ai/the_compendium.pdf
[00:06:50] Geoffrey Hinton Leaves Google, BBC News
https://www.bbc.com/news/world-us-canada-65452940
[00:10:00] ARC-AGI, Chollet
https://arcprize.org/arc-agi
[00:13:25] A Brief History of Intelligence, Bennett
https://www.amazon.com/Brief-History-Intelligence-Humans-Breakthroughs/dp/0063286343
[00:25:35] Statement on AI Risk, Center for AI Safety
https://www.safe.ai/work/statement-on-ai-risk
[00:26:15] Machines of Love and Grace, Amodei
https://darioamodei.com/machines-of-loving-grace
[00:26:35] The Techno-Optimist Manifesto, Andreessen
https://a16z.com/the-techno-optimist-manifesto/
[00:31:55] Techno-Feudalism, Varoufakis
https://www.amazon.co.uk/Technofeudalism-Killed-Capitalism-Yanis-Varoufakis/dp/1847927270
[00:42:40] Introducing Superalignment, OpenAI
https://openai.com/index/introducing-superalignment/
[00:47:20] Three Laws of Robotics, Asimov
https://www.britannica.com/topic/Three-Laws-of-Robotics
[00:50:00] Symbolic AI (GOFAI), Haugeland
https://en.wikipedia.org/wiki/Symbolic_artificial_intelligence
[00:52:30] Intent Alignment, Christiano
https://www.alignmentforum.org/posts/HEZgGBZTpT4Bov7nH/mapping-the-conceptual-territory-in-ai-existential-safety
[00:55:10] Large Language Model Alignment: A Survey, Jiang et al.
http://arxiv.org/pdf/2309.15025
[00:55:40] Constitutional Checks and Balances, Bok
https://plato.stanford.edu/entries/montesquieu/
[00:59:25] Goodharts Law, Manheim & Garrabrant
https://arxiv.org/pdf/1803.04585
[01:00:25] French Constitution Article 37-1, French National Assembly
https://www2.assemblee-nationale.fr/langues/welcome-to-the-english-website-of-the-french-national-assembly
[01:03:00] Intelligence Enhancement and AI Alignment, Yudkowsky
https://intelligence.org/files/AIPosNegFactor.pdf
[01:14:10] UK Government AI Regulation Response, UK Government
https://www.gov.uk/government/consultations/ai-regulation-a-pro-innovation-approach-policy-proposals/outcome/a-pro-innovation-approach-to-ai-regulation-government-response
[01:20:20] French AI Action Summit, Caroli
https://www.csis.org/analysis/frances-ai-action-summit
[01:20:45] Holden Karnofsky Joins Anthropic, Fortune Magazine
https://fortune.com/2025/02/13/anthropic-hired-president-daniela-amodei-husband-ai-safety-responsible-scaling/
[01:25:35] Kill Switch Mechanisms, IEEE/ISO
https://arxiv.org/html/2410.22151v1 Manhattan Project for AI Safety [Connor Leahy]](https://i.ytimg.com/vi/Dt1ySXYTGuA/mqdefault.jpg)
![He Co-Invented the Transformer. Now: Continuous Thought Machines [Llion Jones / Luke Darlow]
The Transformer architecture (which powers ChatGPT and nearly all modern AI) might be trapping the industry in a localized rut, preventing us from finding true intelligent reasoning, according to the person who co-invented it. Llion Jones and Luke Darlow, key figures at the research lab Sakana AI, join the show to make this provocative argument, and also introduce new research (CTM) which might lead the way forwards.
We speak about Inventors Remorse & The Trap of Success Despite being one of the original authors of the famous Attention Is All You Need paper that gave birth to the Transformer, Llion explains why he has largely stopped working on them. He argues that the industry is suffering from success capture—because Transformers work so well, everyone is focused on making small tweaks to the same architecture rather than discovering the next big leap.
**SPONSOR MESSAGES START**
—
Build your ideas with AI Studio from Google - http://ai.studio/build
—
Tufa AI Labs is hiring ML Research Engineers https://tufalabs.ai/
—
cyber•Fund https://cyber.fund/?utm_source=mlst is a founder-led investment firm accelerating the cybernetic economy
Hiring a SF VC Principal: https://talent.cyber.fund/companies/cyber-fund-2/jobs/57674170-ai-investment-principal#content?utm_source=mlst
Submit investment deck: https://cyber.fund/contact?utm_source=mlst
—
**END**
The Spiral Problem – Llion uses a striking visual analogy to explain what current AI is missing. If you ask a standard neural network to understand a spiral shape, it solves it by drawing tiny straight lines that just happen to look like a spiral. It fakes the shape without understanding the concept of spiraling. They argue that todays AI models are similar—they are incredible at mimicking intelligent answers without having an internal process of thinking.
Introducing the Continuous Thought Machine (CTM) Luke Darlow deep dives into their solution: a biology-inspired model that fundamentally changes how AI processes information.
The Maze Analogy: Luke explains that standard AI tries to solve a maze by staring at the whole image and guessing the entire path instantly. Their new machine walks through the maze step-by-step.
Thinking Time: This allows the AI to ponder. If a problem is hard, the model can naturally spend more time thinking about it before answering, effectively allowing it to correct its own mistakes and backtrack—something current Language Models struggle to do genuinely.
The pair discuss the culture of Sakana AI, which is modeled after the early days of Google Brain/DeepMind. Llion nostalgically recalls that the Transformer wasnt born from a corporate mandate, but from random people talking over lunch about interesting problems.
https://sakana.ai/
https://x.com/YesThisIsLion
https://x.com/LearningLukeD
TRANSCRIPT:
https://app.rescript.info/public/share/crjzQ-Jo2FQsJc97xsBdfzfOIeMONpg0TFBuCgV2Fu8
TOC:
00:00:00 - Stepping Back from Transformers
00:00:43 - Introduction to Continuous Thought Machines (CTM)
00:01:09 - The Changing Atmosphere of AI Research
00:04:13 - Sakana’s Philosophy: Research Freedom
00:07:45 - The Local Minimum of Large Language Models
00:18:30 - Representation Problems: The Spiral Example
00:29:12 - Technical Deep Dive: CTM Architecture
00:36:00 - Adaptive Computation & Maze Solving
00:47:15 - Model Calibration & Uncertainty
01:00:43 - Sudoku Bench: Measuring True Reasoning
REFS:
Why Greatness Cannot be planned [Kenneth Stanley]
https://www.amazon.co.uk/Why-Greatness-Cannot-Planned-Objective/dp/3319155237
https://www.youtube.com/watch?v=lhYGXYeMq_E
The Hardware Lottery [Sara Hooker]
https://arxiv.org/abs/2009.06489
https://www.youtube.com/watch?v=sQFxbQ7ade0
Continuous Thought Machines [Luke Darlow et al / Sakana]
https://arxiv.org/abs/2505.05522
https://sakana.ai/ctm/
https://youtu.be/5X9cjGLggv0 great walkthrough of algo by Yacine Mahdid
LSTM: The Comeback Story? [Prof. Sepp Hochreiter]
https://www.youtube.com/watch?v=8u2pW2zZLCs
Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis [Kumar/Stanley]
https://arxiv.org/pdf/2505.11581
Intelligent Matrix Exponentiation [Thomas Fischbacher] (Spiral reference)
https://arxiv.org/abs/2008.03936
A Spline Theory of Deep Networks [Randall Balestriero]
https://proceedings.mlr.press/v80/balestriero18b/balestriero18b.pdf
https://www.youtube.com/watch?v=86ib0sfdFtw
https://www.youtube.com/watch?v=l3O2J3LMxqI
On the Biology of a Large Language Model [Anthropic, Jack Lindsey et al]
https://transformer-circuits.pub/2025/attribution-graphs/biology.html
The ARC Prize 2024 Winning Algorithm [Daniel Franzen and Jan Disselhoff] “The ARChitects”
https://www.youtube.com/watch?v=mTX_sAq zY
Neural Turing Machine [Graves]
https://arxiv.org/pdf/1410.5401
Adaptive Computation Time for Recurrent Neural Networks [Graves]
https://arxiv.org/abs/1603.08983
Sudoko Bench [Sakana]
https://pub.sakana.ai/sudoku/ He Co-Invented the Transformer. Now: Continuous Thought Machines [Llion Jones / Luke Darlow]](https://i.ytimg.com/vi/DtePicx_kFY/mqdefault.jpg)


![When AI Discovers the Next Transformer — Robert Lange
Robert Lange, founding researcher at Sakana AI, joins Tim to discuss *Shinka Evolve* — a framework that combines LLMs with evolutionary algorithms to do open-ended program search. The core claim: systems like AlphaEvolve can optimize solutions to fixed problems, but real scientific progress requires co-evolving the problems themselves.
GTC is coming, the premier AI conference, great opportunity to learn about AI. NVIDIA and partners will showcase breakthroughs in physical AI, AI factories, agentic AI, and inference, exploring the next wave of AI innovation for developers and researchers. Register for virtual GTC for free, using my link and win NVIDIA DGX Spark (https://nvda.ws/4qQ0LMg)
In this episode:
• Why AlphaEvolve gets stuck — it needs a human to hand it the right problem. Shinka tries to invent new problems automatically, drawing on ideas from POET, PowerPlay, and MAP-Elites quality-diversity search.
• The *architecture* of Shinka: an archive of programs organized as islands, LLMs used as mutation operators, and a UCB bandit that adaptively selects between frontier models (GPT-5, Sonnet 4.5, Gemini) mid-run. The credit-assignment problem across models turns out to be genuinely hard.
• Concrete results — state-of-the-art circle packing with dramatically fewer evaluations, second place in an AtCoder competitive programming challenge, evolved load-balancing loss functions for mixture-of-experts models, and agent scaffolds for AIME math benchmarks.
• Are these systems actually thinking outside the box, or are they parasitic on their starting conditions? When LLMs run autonomously, nothing interesting happens. Robert pushes back with the stepping-stone argument — evolution doesnt need to extrapolate, just recombine usefully.
• The AI Scientist question: can automated research pipelines produce real science, or just workshop-level slop that passes surface-level review? Robert is honest that the current version is more co-pilot than autonomous researcher.
• Where this lands in 5-20 years — Roberts prediction that scientific research will be fundamentally transformed, and Tims thought experiment about alien mathematical artifacts that no human could have conceived.
Robert Lange: https://roberttlange.com/
TIMESTAMPS:
00:00:00 Introduction: Robert Lange, Sakana AI and Shinka Evolve
00:04:15 AlphaEvolves Blind Spot: Co-Evolving Problems with Solutions
00:09:05 Unknown Unknowns, POET, and Auto-Curricula for AI Science
00:14:20 MAP-Elites and Quality-Diversity: Shinkas Evolutionary Architecture
00:28:00 UCB Bandits, Mutations and the Vibe Research Vision
00:40:00 Scaling Shinka: Meta-Evolution, Democratisation and the Three-Axis Model
00:47:10 Applications, ARC-AGI and the Future of Work
00:57:00 The AI Scientist and the Human Co-Pilot: Who Steers the Search?
01:06:00 AI Scientist v2, Slop Critique and the Future of Scientific Publishing
REFERENCES:
paper:
[00:03:30] ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution
https://arxiv.org/abs/2509.19349
[00:04:15] AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery
https://arxiv.org/abs/2506.13131
[00:06:30] Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
https://arxiv.org/abs/2505.22954
[00:09:05] Paired Open-Ended Trailblazer (POET)
https://arxiv.org/abs/1901.01753
[00:10:00] PowerPlay: Training an Increasingly General Problem Solver by Continually Searching for the Simplest Still Unsolvable Problem
https://arxiv.org/abs/1112.5309
[00:10:40] Automated Capability Discovery via Foundation Model Self-Exploration
https://arxiv.org/abs/2502.07577
[00:15:30] Illuminating Search Spaces by Mapping Elites (MAP-Elites)
https://arxiv.org/abs/1504.04909
[00:47:10] Automated Design of Agentic Systems (ADAS)
https://arxiv.org/abs/2408.08435
[00:49:50] Discovering Preference Optimization Algorithms with and for Large Language Models (DiscoPOP)
https://arxiv.org/abs/2406.08414
[00:57:00] The AI Scientist v2: Automating the Full Research Pipeline
https://arxiv.org/abs/2504.08066
book:
[00:06:48] Why Greatness Cannot Be Planned
https://link.springer.com/book/10.1007/978-3-319-15524-1
benchmark:
[00:47:10] ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering
https://arxiv.org/abs/2506.09050
[00:50:50] On the Measure of Intelligence (ARC-AGI)
https://arxiv.org/abs/1911.01547
LINKS:
Download PDF transcript: https://app.rescript.info/api/sessions/b8a9dcf60623657c/pdf/download
Full Transcript: https://app.rescript.info/public/share/SDOD_3oXOcli3zTqcAtR8eibT5U3gam84oo4KRtI-Vk When AI Discovers the Next Transformer — Robert Lange](https://i.ytimg.com/vi/EInEmGaMRLc/mqdefault.jpg)
![Transformers Need Glasses! [Federico Barbero]
SPONSOR MESSAGES:
***
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments. Check out their super fast DeepSeek R1 hosting!
https://centml.ai/pricing/
Federico Barbero (DeepMind/Oxford) is the lead author of Transformers Need Glasses!, a paper revealing fundamental architectural limitations in transformer-based language models. The conversation explores why LLMs fail at seemingly trivial tasks like copying the last token of a sequence or counting repeated elements, tracing these failures to representation collapse — where internal representations of distinct inputs converge below machine precision as context length grows. Federico connects these findings to information propagation theory from graph neural networks, showing how causal attention creates an inherent bias toward the start of a sequence, while training dynamics push models to attend to the most recent tokens, leaving the middle of long contexts as a blind spot. The discussion covers connections to spectral graph theory, heat equations on graphs, Petar Veličkovićs work on graph attention networks, the role of softmax in limiting sharp attention, and practical glasses — architectural tweaks that can help transformers see more clearly.
REFERENCES:
paper:
[00:01:05] Transformers Need Glasses!
https://proceedings.neurips.cc/paper_files/paper/2024/file/b1d35561c4a4a0e0b6012b2af531e149-Paper-Conference.pdf
[00:05:30] Softmax is Not Enough
https://arxiv.org/abs/2410.01104
[00:15:05] Graph Attention Networks
https://arxiv.org/abs/1710.10903
[00:43:23] Neural Networks and the Chomsky Hierarchy
https://arxiv.org/abs/2207.02098
[00:51:04] On the Measure of Intelligence
https://arxiv.org/abs/1911.01547
[01:00:35] Epistemic Foraging
https://www.frontiersin.org/journals/computational-neuroscience/articles/10.3389/fncom.2016.00056/full
LINKS:
Full Transcript: https://app.rescript.info/share/d43b3918dc7cb8c8b822ebef40b8f66f
Download PDF transcript: https://app.rescript.info/api/public/sessions/53f7d10ff5c72eca/pdf Transformers Need Glasses! [Federico Barbero]](https://i.ytimg.com/vi/FAspMnu4Rt0/mqdefault.jpg)


![29.4% ARC-AGI-2 🤯 (TOP SCORE!) - Jeremy Berman
Jeremy Berman took the top spot on the ARC-AGI v2 public leaderboard with a score of about 30% using an approach that trades code for natural language. Where his first attempt evolved Python programs to solve abstract reasoning puzzles, this version evolves plain English descriptions of transformation rules, then uses a strong thinking model as a checker agent to verify them against training examples. The shift to natural language is the interesting part: English can describe ARC v2 tasks in five bullet points where Python takes dozens of lines, and the higher expressiveness lets the system explore solution spaces that rigid code simply cannot reach.
The conversation goes deep on why this matters for intelligence research. Berman and Tim work through the relationship between reinforcement learning and genuine reasoning whether RL can replace the messy pretrained knowledge web with a clean deductive tree, what catastrophic forgetting really blocks, and why the meta-skill of reasoning (the ability to create new skills) is the actual target for AGI. Berman makes a sharp distinction between knowledge that is memorized and knowledge that is deduced, arguing that pretraining treats everything as an interconnected web when what we actually need is causal structure.
They also dig into composability (freezing expert layers, Docker-for-models), whether neural networks can ever run Turing-complete algorithms the way biological brains seem to, and what it would take to build an invention circuit the machinery for genuine creative synthesis rather than pattern recombination. The discussion lands on a shared framework where intelligence is the efficiency of building epistemic trees, reasoning is constructing them, and understanding is possessing them.
**SPONSOR MESSAGES**
—
Take the Prolific human data survey - https://www.prolific.com/humandatasurvey?utm_source=mlst and be the first to see the results and benchmark their practices against the wider community!
—
cyber•Fund https://cyber.fund/?utm_source=mlst is a founder-led investment firm accelerating the cybernetic economy
Oct SF conference - https://dagihouse.com/?utm_source=mlst - Joscha Bach keynoting(!) + OAI, Anthropic, NVDA,++
Hiring a SF VC Principal: https://talent.cyber.fund/companies/cyber-fund-2/jobs/57674170-ai-investment-principal#content?utm_source=mlst
Submit investment deck: https://cyber.fund/contact?utm_source=mlst
—
REFERENCES:
Blog Post:
[00:03:51] Jeremy Bermans ARC-AGI v1 Blog Post
https://jeremyberman.substack.com/p/how-i-got-a-record-536-on-arc-agi
[00:07:20] Getting 50% on ARC-AGI with GPT-4o
https://blog.redwoodresearch.org/p/getting-50-sota-on-arc-agi-with-gpt
Book:
[00:04:30] A Thousand Brains
https://www.amazon.com/Thousand-Brains-New-Theory-Intelligence/dp/1541675819
[00:48:07] Deep Learning with Python Rev 3
https://deeplearningwithpython.io/
Company:
[00:05:35] NDEA
https://ndea.com/
Paper:
[00:13:27] On the Biology of a Large Language Model
https://transformer-circuits.pub/2025/attribution-graphs/biology.html
[00:24:09] Connectionism and Cognitive Architecture
https://uh.edu/~garson/F&P1.PDF
[00:29:50] Fractured Entangled Representation Hypothesis
https://arxiv.org/pdf/2505.11581
[00:44:00] Shinka Evolve
https://sakana.ai/shinka-evolve/
[00:46:22] On the Measure of Intelligence
https://arxiv.org/abs/1911.01547
Video:
[00:19:12] The ARChitects
https://www.youtube.com/watch?v=mTX_sAq zY
[00:44:00] AlphaEvolve
https://www.youtube.com/watch?v=vC9nAosXrJw
LINKS:
Full Transcript: https://app.rescript.info/share/nMcpyRPCWh652R0DoQbHAS90BWzs4yTGrwr994YHQRM
Download PDF transcript: https://app.rescript.info/api/public/sessions/9f3d5e5741358a24/pdf
Jeremy Berman:
https://x.com/jerber888
REFS:
Jeremys 2024 article on winning ARCAGI1-pub
https://jeremyberman.substack.com/p/how-i-got-a-record-536-on-arc-agi 29.4% ARC-AGI-2 🤯 (TOP SCORE!) - Jeremy Berman](https://i.ytimg.com/vi/FcnLiPyfRZM/mqdefault.jpg)