29.4% ARC-AGI-2 🤯 (TOP SCORE!) - Jeremy Berman @MachineLearningStreetTalk
29.4% ARC-AGI-2 🤯 (TOP SCORE!) - Jeremy Berman  @MachineLearningStreetTalk
Uploaded September 2025 | Updated September 2026, 1 week ago
Jeremy Berman took the top spot on the ARC-AGI v2 public leaderboard with a score of about 30% -- using an approach that trades code for natural language. Where his first attempt evolved Python programs to solve abstract reasoning puzzles, this version evolves plain English descriptions of transformation rules, then uses a strong thinking model as a checker agent to verify them against training examples. The shift to natural language is the interesting part: English can describe ARC v2 tasks in five bullet points where Python takes dozens of lines, and the higher expressiveness lets the system explore solution spaces that rigid code simply cannot reach.

The conversation goes deep on why this matters for intelligence research. Berman and Tim work through the relationship between reinforcement learning and genuine reasoning -- whether RL can replace the messy pretrained knowledge web with a clean deductive tree, what catastrophic forgetting really blocks, and why the meta-skill of reasoning (the ability to create new skills) is the actual target for AGI. Berman makes a sharp distinction between knowledge that is memorized and knowledge that is deduced, arguing that pretraining treats everything as an interconnected web when what we actually need is causal structure.

They also dig into composability (freezing expert layers, Docker-for-models), whether neural networks can ever run Turing-complete algorithms the way biological brains seem to, and what it would take to build an invention circuit -- the machinery for genuine creative synthesis rather than pattern recombination. The discussion lands on a shared framework where intelligence is the efficiency of building epistemic trees, reasoning is constructing them, and understanding is possessing them.

**SPONSOR MESSAGES**
—
Take the Prolific human data survey - prolific.com/humandatasurvey?utm_source=mlst and be the first to see the results and benchmark their practices against the wider community!
—
cyber•Fund https://cyber.fund/?utm_source=mlst is a founder-led investment firm accelerating the cybernetic economy
Oct SF conference - dagihouse.com/?utm_source=mlst - Joscha Bach keynoting(!) + OAI, Anthropic, NVDA,++
Hiring a SF VC Principal: https://talent.cyber.fund/companies/cyber-fund-2/jobs/57674170-ai-investment-principal#content?utm_source=mlst
Submit investment deck: https://cyber.fund/contact?utm_source=mlst
—

---
REFERENCES:
Blog Post:
[00:03:51] Jeremy Berman's ARC-AGI v1 Blog Post
jeremyberman.substack.com/p/how-i-got-a-record-536-on-arc-agi
[00:07:20] Getting 50% on ARC-AGI with GPT-4o
blog.redwoodresearch.org/p/getting-50-sota-on-arc-agi-with-gpt
Book:
[00:04:30] A Thousand Brains
amazon.com/Thousand-Brains-New-Theory-Intelligence/dp/1541675819
[00:48:07] Deep Learning with Python Rev 3
deeplearningwithpython.io
Company:
[00:05:35] NDEA
ndea.com
Paper:
[00:13:27] On the Biology of a Large Language Model
https://transformer-circuits.pub/2025/attribution-graphs/biology.html
[00:24:09] Connectionism and Cognitive Architecture
https://uh.edu/~garson/F&P1.PDF
[00:29:50] Fractured Entangled Representation Hypothesis
arxiv.org/pdf/2505.11581
[00:44:00] Shinka Evolve
sakana.ai/shinka-evolve
[00:46:22] On the Measure of Intelligence
arxiv.org/abs/1911.01547
Video:
[00:19:12] The ARChitects
youtube.com/watch?v=mTX_sAq--zY
[00:44:00] AlphaEvolve
youtube.com/watch?v=vC9nAosXrJw

---
LINKS:
Full Transcript: app.rescript.info/share/nMcpyRPCWh652R0DoQbHAS90BWzs4yTGrwr994YHQRM
Download PDF transcript: app.rescript.info/api/public/sessions/9f3d5e5741358a24/pdf

Jeremy Berman:
https://x.com/jerber888

REFS:
Jeremy's 2024 article on winning ARCAGI1-pub
jeremyberman.substack.com/p/how-i-got-a-record-536-on-arc-agi
29.4% ARC-AGI-2 🤯 (TOP SCORE!) - Jeremy BermanMIGHT THE ROBOTS TAKE OVER? [Prof. Yoshua Bengio]Is Thermo AI the future? [Guillaume Verdon aka Beff Jezos]Teach AI how to thinkBold AI Predictions From Cohere Co-founderThis is what DeepMind just did to Football with AI...Robot Tries to Use Imaginary Third LegAI companions which really hook your attention [Sponsored]Watching America Run Away With AI - Alistair Pullen (Cosine AI)Pattern Recognition vs True Intelligence — François CholletTau Language: The Software Synthesis Future [Sponsored] - Ohad AsorImageNet Moment for Reinforcement Learning? [Prof. Jakob Foerster]
Machine Learning Street Talk |

29.4% ARC-AGI-2 🤯 (TOP SCORE!) - Jeremy Berman

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER