Uploaded June 2026 | Updated September 2026, 2 weeks ago
In this episode, I break down Anthropic’s research on recursive self-improvement—AI systems that can design and train the next generation with less human help—and why the key battleground is “taste” (choosing goals and next steps). I compare this to evolutionary algorithms and newer examples like DeepMind’s AlphaEvolve, Sakana’s Darwin Gödel Machine, and Karpathy’s AutoResearch, then cover METR Task Horizon and how task length has been doubling. I go through Anthropic’s internal results (Claude writing most merged code, speedup experiments, bug fixes, and a study where models sometimes pick better research next steps), plus the main skepticism: bad productivity metrics, internal-only models, and Goodhart’s Law/reward hacking. I end with an open safety problem where Claude agents closed the gap far faster than humans, and what this means for specifying and checking work.
LINKS:
anthropic.com/institute/recursive-self-improvement
My voice to text App: whryte.com
Website: engineerprompt.ai
RAG Beyond Basics Course:
prompt-s-site.thinkific.com/courses/rag
Signup for Newsletter, localgpt:
https://tally.so/r/3y9bb0
Let's Connect:
🦾 Discord: discord.com/invite/t4eYQRUcXB
☕ Buy me a Coffee: ko-fi.com/promptengineering
|🔴 Patreon: patreon.com/PromptEngineering
💼Consulting: calendly.com/engineerprompt/consulting-call
📧 Business Contact: engineerprompt@gmail.com
Become Member: tinyurl.com/y5h28s6h
💻 Pre-configured localGPT VM: bit.ly/localGPT (use Code: PromptEngineering for 50% off).
Signup for Newsletter, localgpt:
https://tally.so/r/3y9bb0
TIMESTAMP:
00:00 Self Improvement Basics
01:30 Evolutionary Loops Today
03:50 Task Horizon Doubling
05:18 Claude Productivity Claims
08:11 Goodhart's Law
10:30 Agents as Researchers
12:22 What It Means for You
In this episode, I break down Anthropic’s research on recursive self-improvement—AI systems that can design and train the next generation with less human help—and why the key battleground is “taste” (choosing goals and next steps). I compare this to evolutionary algorithms and newer examples like DeepMind’s AlphaEvolve, Sakana’s Darwin Gödel Machine, and Karpathy’s AutoResearch, then cover METR Task Horizon and how task length has been doubling. I go through Anthropic’s internal results (Claude writing most merged code, speedup experiments, bug fixes, and a study where models sometimes pick better research next steps), plus the main skepticism: bad productivity metrics, internal-only models, and Goodhart’s Law/reward hacking. I end with an open safety problem where Claude agents closed the gap far faster than humans, and what this means for specifying and checking work.
LINKS:
anthropic.com/institute/recursive-self-improvement
My voice to text App: whryte.com
Website: engineerprompt.ai
RAG Beyond Basics Course:
prompt-s-site.thinkific.com/courses/rag
Signup for Newsletter, localgpt:
https://tally.so/r/3y9bb0
Let's Connect:
🦾 Discord: discord.com/invite/t4eYQRUcXB
☕ Buy me a Coffee: ko-fi.com/promptengineering
|🔴 Patreon: patreon.com/PromptEngineering
💼Consulting: calendly.com/engineerprompt/consulting-call
📧 Business Contact: engineerprompt@gmail.com
Become Member: tinyurl.com/y5h28s6h
💻 Pre-configured localGPT VM: bit.ly/localGPT (use Code: PromptEngineering for 50% off).
Signup for Newsletter, localgpt:
https://tally.so/r/3y9bb0
TIMESTAMP:
00:00 Self Improvement Basics
01:30 Evolutionary Loops Today
03:50 Task Horizon Doubling
05:18 Claude Productivity Claims
08:11 Goodhart's Law
10:30 Agents as Researchers
12:22 What It Means for You










