Uploaded March 2026 | Updated September 2026, 1 week ago
Robert Lange, founding researcher at Sakana AI, joins Tim to discuss *Shinka Evolve* — a framework that combines LLMs with evolutionary algorithms to do open-ended program search. The core claim: systems like AlphaEvolve can optimize solutions to fixed problems, but real scientific progress requires co-evolving the problems themselves.
GTC is coming, the premier AI conference, great opportunity to learn about AI. NVIDIA and partners will showcase breakthroughs in physical AI, AI factories, agentic AI, and inference, exploring the next wave of AI innovation for developers and researchers. Register for virtual GTC for free, using my link and win NVIDIA DGX Spark (nvda.ws/4qQ0LMg)
In this episode:
• Why AlphaEvolve gets stuck — it needs a human to hand it the right problem. Shinka tries to invent new problems automatically, drawing on ideas from POET, PowerPlay, and MAP-Elites quality-diversity search.
• The *architecture* of Shinka: an archive of programs organized as islands, LLMs used as mutation operators, and a UCB bandit that adaptively selects between frontier models (GPT-5, Sonnet 4.5, Gemini) mid-run. The credit-assignment problem across models turns out to be genuinely hard.
• Concrete results — state-of-the-art circle packing with dramatically fewer evaluations, second place in an AtCoder competitive programming challenge, evolved load-balancing loss functions for mixture-of-experts models, and agent scaffolds for AIME math benchmarks.
• Are these systems actually thinking outside the box, or are they parasitic on their starting conditions? When LLMs run autonomously, "nothing interesting happens." Robert pushes back with the stepping-stone argument — evolution doesn't need to extrapolate, just recombine usefully.
• The AI Scientist question: can automated research pipelines produce real science, or just workshop-level slop that passes surface-level review? Robert is honest that the current version is more co-pilot than autonomous researcher.
• Where this lands in 5-20 years — Robert's prediction that scientific research will be fundamentally transformed, and Tim's thought experiment about alien mathematical artifacts that no human could have conceived.
Robert Lange: roberttlange.com
---
TIMESTAMPS:
00:00:00 Introduction: Robert Lange, Sakana AI and Shinka Evolve
00:04:15 AlphaEvolve's Blind Spot: Co-Evolving Problems with Solutions
00:09:05 Unknown Unknowns, POET, and Auto-Curricula for AI Science
00:14:20 MAP-Elites and Quality-Diversity: Shinka's Evolutionary Architecture
00:28:00 UCB Bandits, Mutations and the Vibe Research Vision
00:40:00 Scaling Shinka: Meta-Evolution, Democratisation and the Three-Axis Model
00:47:10 Applications, ARC-AGI and the Future of Work
00:57:00 The AI Scientist and the Human Co-Pilot: Who Steers the Search?
01:06:00 AI Scientist v2, Slop Critique and the Future of Scientific Publishing
---
REFERENCES:
paper:
[00:03:30] ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution
arxiv.org/abs/2509.19349
[00:04:15] AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery
arxiv.org/abs/2506.13131
[00:06:30] Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
arxiv.org/abs/2505.22954
[00:09:05] Paired Open-Ended Trailblazer (POET)
arxiv.org/abs/1901.01753
[00:10:00] PowerPlay: Training an Increasingly General Problem Solver by Continually Searching for the Simplest Still Unsolvable Problem
arxiv.org/abs/1112.5309
[00:10:40] Automated Capability Discovery via Foundation Model Self-Exploration
arxiv.org/abs/2502.07577
[00:15:30] Illuminating Search Spaces by Mapping Elites (MAP-Elites)
arxiv.org/abs/1504.04909
[00:47:10] Automated Design of Agentic Systems (ADAS)
arxiv.org/abs/2408.08435
[00:49:50] Discovering Preference Optimization Algorithms with and for Large Language Models (DiscoPOP)
arxiv.org/abs/2406.08414
[00:57:00] The AI Scientist v2: Automating the Full Research Pipeline
arxiv.org/abs/2504.08066
book:
[00:06:48] Why Greatness Cannot Be Planned
link.springer.com/book/10.1007/978-3-319-15524-1
benchmark:
[00:47:10] ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering
arxiv.org/abs/2506.09050
[00:50:50] On the Measure of Intelligence (ARC-AGI)
arxiv.org/abs/1911.01547
---
LINKS:
Download PDF transcript: app.rescript.info/api/sessions/b8a9dcf60623657c/pdf/download
Full Transcript: app.rescript.info/public/share/SDOD_3oXOcli3zTqcAtR8eibT5U3gam84oo4KRtI-Vk
Robert Lange, founding researcher at Sakana AI, joins Tim to discuss *Shinka Evolve* — a framework that combines LLMs with evolutionary algorithms to do open-ended program search. The core claim: systems like AlphaEvolve can optimize solutions to fixed problems, but real scientific progress requires co-evolving the problems themselves.
GTC is coming, the premier AI conference, great opportunity to learn about AI. NVIDIA and partners will showcase breakthroughs in physical AI, AI factories, agentic AI, and inference, exploring the next wave of AI innovation for developers and researchers. Register for virtual GTC for free, using my link and win NVIDIA DGX Spark (nvda.ws/4qQ0LMg)
In this episode:
• Why AlphaEvolve gets stuck — it needs a human to hand it the right problem. Shinka tries to invent new problems automatically, drawing on ideas from POET, PowerPlay, and MAP-Elites quality-diversity search.
• The *architecture* of Shinka: an archive of programs organized as islands, LLMs used as mutation operators, and a UCB bandit that adaptively selects between frontier models (GPT-5, Sonnet 4.5, Gemini) mid-run. The credit-assignment problem across models turns out to be genuinely hard.
• Concrete results — state-of-the-art circle packing with dramatically fewer evaluations, second place in an AtCoder competitive programming challenge, evolved load-balancing loss functions for mixture-of-experts models, and agent scaffolds for AIME math benchmarks.
• Are these systems actually thinking outside the box, or are they parasitic on their starting conditions? When LLMs run autonomously, "nothing interesting happens." Robert pushes back with the stepping-stone argument — evolution doesn't need to extrapolate, just recombine usefully.
• The AI Scientist question: can automated research pipelines produce real science, or just workshop-level slop that passes surface-level review? Robert is honest that the current version is more co-pilot than autonomous researcher.
• Where this lands in 5-20 years — Robert's prediction that scientific research will be fundamentally transformed, and Tim's thought experiment about alien mathematical artifacts that no human could have conceived.
Robert Lange: roberttlange.com
---
TIMESTAMPS:
00:00:00 Introduction: Robert Lange, Sakana AI and Shinka Evolve
00:04:15 AlphaEvolve's Blind Spot: Co-Evolving Problems with Solutions
00:09:05 Unknown Unknowns, POET, and Auto-Curricula for AI Science
00:14:20 MAP-Elites and Quality-Diversity: Shinka's Evolutionary Architecture
00:28:00 UCB Bandits, Mutations and the Vibe Research Vision
00:40:00 Scaling Shinka: Meta-Evolution, Democratisation and the Three-Axis Model
00:47:10 Applications, ARC-AGI and the Future of Work
00:57:00 The AI Scientist and the Human Co-Pilot: Who Steers the Search?
01:06:00 AI Scientist v2, Slop Critique and the Future of Scientific Publishing
---
REFERENCES:
paper:
[00:03:30] ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution
arxiv.org/abs/2509.19349
[00:04:15] AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery
arxiv.org/abs/2506.13131
[00:06:30] Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
arxiv.org/abs/2505.22954
[00:09:05] Paired Open-Ended Trailblazer (POET)
arxiv.org/abs/1901.01753
[00:10:00] PowerPlay: Training an Increasingly General Problem Solver by Continually Searching for the Simplest Still Unsolvable Problem
arxiv.org/abs/1112.5309
[00:10:40] Automated Capability Discovery via Foundation Model Self-Exploration
arxiv.org/abs/2502.07577
[00:15:30] Illuminating Search Spaces by Mapping Elites (MAP-Elites)
arxiv.org/abs/1504.04909
[00:47:10] Automated Design of Agentic Systems (ADAS)
arxiv.org/abs/2408.08435
[00:49:50] Discovering Preference Optimization Algorithms with and for Large Language Models (DiscoPOP)
arxiv.org/abs/2406.08414
[00:57:00] The AI Scientist v2: Automating the Full Research Pipeline
arxiv.org/abs/2504.08066
book:
[00:06:48] Why Greatness Cannot Be Planned
link.springer.com/book/10.1007/978-3-319-15524-1
benchmark:
[00:47:10] ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering
arxiv.org/abs/2506.09050
[00:50:50] On the Measure of Intelligence (ARC-AGI)
arxiv.org/abs/1911.01547
---
LINKS:
Download PDF transcript: app.rescript.info/api/sessions/b8a9dcf60623657c/pdf/download
Full Transcript: app.rescript.info/public/share/SDOD_3oXOcli3zTqcAtR8eibT5U3gam84oo4KRtI-Vk
![Transformers Need Glasses! [Federico Barbero]
SPONSOR MESSAGES:
***
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments. Check out their super fast DeepSeek R1 hosting!
https://centml.ai/pricing/
Federico Barbero (DeepMind/Oxford) is the lead author of Transformers Need Glasses!, a paper revealing fundamental architectural limitations in transformer-based language models. The conversation explores why LLMs fail at seemingly trivial tasks like copying the last token of a sequence or counting repeated elements, tracing these failures to representation collapse — where internal representations of distinct inputs converge below machine precision as context length grows. Federico connects these findings to information propagation theory from graph neural networks, showing how causal attention creates an inherent bias toward the start of a sequence, while training dynamics push models to attend to the most recent tokens, leaving the middle of long contexts as a blind spot. The discussion covers connections to spectral graph theory, heat equations on graphs, Petar Veličkovićs work on graph attention networks, the role of softmax in limiting sharp attention, and practical glasses — architectural tweaks that can help transformers see more clearly.
REFERENCES:
paper:
[00:01:05] Transformers Need Glasses!
https://proceedings.neurips.cc/paper_files/paper/2024/file/b1d35561c4a4a0e0b6012b2af531e149-Paper-Conference.pdf
[00:05:30] Softmax is Not Enough
https://arxiv.org/abs/2410.01104
[00:15:05] Graph Attention Networks
https://arxiv.org/abs/1710.10903
[00:43:23] Neural Networks and the Chomsky Hierarchy
https://arxiv.org/abs/2207.02098
[00:51:04] On the Measure of Intelligence
https://arxiv.org/abs/1911.01547
[01:00:35] Epistemic Foraging
https://www.frontiersin.org/journals/computational-neuroscience/articles/10.3389/fncom.2016.00056/full
LINKS:
Full Transcript: https://app.rescript.info/share/d43b3918dc7cb8c8b822ebef40b8f66f
Download PDF transcript: https://app.rescript.info/api/public/sessions/53f7d10ff5c72eca/pdf Transformers Need Glasses! [Federico Barbero]](https://i.ytimg.com/vi/FAspMnu4Rt0/mqdefault.jpg)


![29.4% ARC-AGI-2 🤯 (TOP SCORE!) - Jeremy Berman
Jeremy Berman took the top spot on the ARC-AGI v2 public leaderboard with a score of about 30% using an approach that trades code for natural language. Where his first attempt evolved Python programs to solve abstract reasoning puzzles, this version evolves plain English descriptions of transformation rules, then uses a strong thinking model as a checker agent to verify them against training examples. The shift to natural language is the interesting part: English can describe ARC v2 tasks in five bullet points where Python takes dozens of lines, and the higher expressiveness lets the system explore solution spaces that rigid code simply cannot reach.
The conversation goes deep on why this matters for intelligence research. Berman and Tim work through the relationship between reinforcement learning and genuine reasoning whether RL can replace the messy pretrained knowledge web with a clean deductive tree, what catastrophic forgetting really blocks, and why the meta-skill of reasoning (the ability to create new skills) is the actual target for AGI. Berman makes a sharp distinction between knowledge that is memorized and knowledge that is deduced, arguing that pretraining treats everything as an interconnected web when what we actually need is causal structure.
They also dig into composability (freezing expert layers, Docker-for-models), whether neural networks can ever run Turing-complete algorithms the way biological brains seem to, and what it would take to build an invention circuit the machinery for genuine creative synthesis rather than pattern recombination. The discussion lands on a shared framework where intelligence is the efficiency of building epistemic trees, reasoning is constructing them, and understanding is possessing them.
**SPONSOR MESSAGES**
—
Take the Prolific human data survey - https://www.prolific.com/humandatasurvey?utm_source=mlst and be the first to see the results and benchmark their practices against the wider community!
—
cyber•Fund https://cyber.fund/?utm_source=mlst is a founder-led investment firm accelerating the cybernetic economy
Oct SF conference - https://dagihouse.com/?utm_source=mlst - Joscha Bach keynoting(!) + OAI, Anthropic, NVDA,++
Hiring a SF VC Principal: https://talent.cyber.fund/companies/cyber-fund-2/jobs/57674170-ai-investment-principal#content?utm_source=mlst
Submit investment deck: https://cyber.fund/contact?utm_source=mlst
—
REFERENCES:
Blog Post:
[00:03:51] Jeremy Bermans ARC-AGI v1 Blog Post
https://jeremyberman.substack.com/p/how-i-got-a-record-536-on-arc-agi
[00:07:20] Getting 50% on ARC-AGI with GPT-4o
https://blog.redwoodresearch.org/p/getting-50-sota-on-arc-agi-with-gpt
Book:
[00:04:30] A Thousand Brains
https://www.amazon.com/Thousand-Brains-New-Theory-Intelligence/dp/1541675819
[00:48:07] Deep Learning with Python Rev 3
https://deeplearningwithpython.io/
Company:
[00:05:35] NDEA
https://ndea.com/
Paper:
[00:13:27] On the Biology of a Large Language Model
https://transformer-circuits.pub/2025/attribution-graphs/biology.html
[00:24:09] Connectionism and Cognitive Architecture
https://uh.edu/~garson/F&P1.PDF
[00:29:50] Fractured Entangled Representation Hypothesis
https://arxiv.org/pdf/2505.11581
[00:44:00] Shinka Evolve
https://sakana.ai/shinka-evolve/
[00:46:22] On the Measure of Intelligence
https://arxiv.org/abs/1911.01547
Video:
[00:19:12] The ARChitects
https://www.youtube.com/watch?v=mTX_sAq zY
[00:44:00] AlphaEvolve
https://www.youtube.com/watch?v=vC9nAosXrJw
LINKS:
Full Transcript: https://app.rescript.info/share/nMcpyRPCWh652R0DoQbHAS90BWzs4yTGrwr994YHQRM
Download PDF transcript: https://app.rescript.info/api/public/sessions/9f3d5e5741358a24/pdf
Jeremy Berman:
https://x.com/jerber888
REFS:
Jeremys 2024 article on winning ARCAGI1-pub
https://jeremyberman.substack.com/p/how-i-got-a-record-536-on-arc-agi 29.4% ARC-AGI-2 🤯 (TOP SCORE!) - Jeremy Berman](https://i.ytimg.com/vi/FcnLiPyfRZM/mqdefault.jpg)
![MIGHT THE ROBOTS TAKE OVER? [Prof. Yoshua Bengio]
SPONSOR MESSAGES:
***
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments.
https://centml.ai/pricing/
Turing Award winner Professor Yoshua Bengio sits down with Tim Scarfe for a wide-ranging discussion on AI safety, the architecture of intelligence, and what happens when machines become as capable as our best researchers.
Bengio argues that the current race toward agentic AI systems carries fundamental risks that most people underestimate. He lays out how instrumental convergence and reward tampering could lead systems to develop goals misaligned with human interests not through malice, but through the basic logic of optimization. His proposed alternative: powerful AI tools that function as scientific oracles rather than autonomous agents, systems that can revolutionize medicine and science without needing goals of their own.
The conversation covers international AI governance and the game theory driving the US-China competition, the military implications of frontier AI capabilities, and why hardware-enabled governance may be the most promising path toward verifiable AI treaties. Bengio draws direct parallels to nuclear nonproliferation and warns that data centers will become strategic military assets.
On the technical side, Bengio discusses the return of recurrent architectures (the Were RNNs All We Needed? paper), his work on GFlowNets for probabilistic inference with discrete structures, a complexity-based theory of compositionality that tries to formalize what was previously just intuition, and the question of whether System 2 reasoning needs to be designed from first principles rather than bolted onto existing architectures.
Throughout, Bengio repeatedly emphasizes what he does not know a deliberate stance he contrasts with the overconfidence he sees driving dangerous policy decisions.
REFERENCES:
paper:
[00:00:15] AI Risk Statement
https://www.safe.ai/work/statement-on-ai-risk
[00:23:10] Reward Tampering and AI Safety
https://hdsr.mitpress.mit.edu/pub/w974bwb0
[00:44:30] Can a Bayesian Oracle Prevent Harm?
https://arxiv.org/abs/2408.05284
[00:52:00] Hardware-Enabled AI Governance Memo
https://yoshuabengio.org/wp-content/uploads/2024/08/FlexHEG-Memo_August-2024.pdf
[01:33:20] GFlowNet Foundations
https://arxiv.org/abs/2111.09266
[01:35:05] Complexity-Based Compositionality Theory
https://arxiv.org/abs/2410.14817
[01:37:50] Discrete Attractor States in Neural Systems
https://arxiv.org/abs/2302.06403
person:
[00:03:50] Professor Yoshua Bengio
https://yoshuabengio.org/
tool:
[00:40:45] Munk Debate on AI Existential Risk
https://munkdebates.com/debates/artificial-intelligence
[00:56:07] 2018 Turing Award
https://awards.acm.org/about/2018-turing
[01:12:35] EU AI Act Code of Practice
https://digital-strategy.ec.europa.eu/en/news/meet-chairs-leading-development-first-general-purpose-ai-code-practice
LINKS:
Full Transcript: https://app.rescript.info/share/1c4f010852d0db9a34dbeb043c0da26b
Download PDF transcript: https://app.rescript.info/api/public/sessions/baff3f8846202250/pdf
Yoshua Bengio:
https://x.com/Yoshua_Bengio
https://scholar.google.com/citations?user=kukA0LcAAAAJ&hl=en
https://yoshuabengio.org/
https://en.wikipedia.org/wiki/Yoshua_Bengio MIGHT THE ROBOTS TAKE OVER? [Prof. Yoshua Bengio]](https://i.ytimg.com/vi/G1ARvwQntAU/mqdefault.jpg)
![Is Thermo AI the future? [Guillaume Verdon aka Beff Jezos]
Dr. Maxwell Ramstead hosts Guillaume Verdon physicist, founder of Extropic, and the mind behind the Beff Jezos persona and the Effective Accelerationism movement for a conversation that bridges thermodynamics, computing hardware, and civilizational philosophy.
Verdon traces his path from childhood fascination with theories of everything through theoretical physics at the Perimeter Institute, quantum computing at Google, and the founding of Extropic. The core technical insight: instead of fighting thermal noise at enormous energetic cost (as quantum computers do), thermodynamic computing harnesses it. Extropics chips use the natural stochastic physics of electrons to accelerate Markov chain Monte Carlo sampling the same class of algorithms that underpin diffusion models, energy-based models, and much of modern probabilistic ML.
The numbers are striking. Current wafer-scale AI systems consume 20+ kilowatts. A thermodynamic wafer with 1.5 billion p-bits and 20 billion parameters would run on 20 watts roughly what the human brain uses. Verdon argues this is not a coincidence: the brain is literally a thermodynamic computer operating near the Landauer limit, and Extropic is building silicon that works on the same principles.
The second half turns to philosophy. Verdon derives Effective Accelerationism directly from stochastic thermodynamics: thermodynamic selection pressure favors systems that capture and dissipate more free energy, so growth is not just desirable but physically inevitable. He and Ramstead explore hyperstition through the lens of active inference, the geopolitical stakes of the US-China technology race, and why Verdon views deceleration as a form of psychological warfare against Western competitiveness.
Recorded with Maxwell Ramstead hosting in place of Tim Scarfe.
***SPONSOR MESSAGE***
Google Gemini 2.5 Flash is a state-of-the-art language model in the Gemini app. Sign up at https://gemini.google.com
***
TIMESTAMPS:
00:00:00 Intro Montage & Sponsor
00:02:21 From Theories of Everything to Thermodynamic Computing
00:08:41 The Failure of Reductionism & Physics-Inspired AI
00:17:15 From Quantum to Thermodynamic Computing
00:23:56 How Thermodynamic Computers Actually Work
00:31:15 Moores Wall & The Brain as Proof of Concept
00:40:00 The Computing Stack of the Future
00:48:23 Why Current AI Will Cook Us to Death
00:50:38 Effective Accelerationism: The Philosophy of Beff Jezos
01:00:00 Hyperstition, Geopolitics & Avoiding Catastrophe
01:16:00 Closing: The Future of Extropic & Getting Involved
REFERENCES:
paper:
[00:04:40] It From Bit
https://cqi.inf.usi.ch/qic/wheeler.pdf
[00:06:15] Renormalization Group Theory
https://www.damtp.cam.ac.uk/user/dbs26/AQFT/Wilsonchap.pdf
[00:23:00] Metropolis-Hastings Algorithm
https://arxiv.org/abs/1504.01896
[00:40:00] Deep Information Bottleneck and Renormalization
https://arxiv.org/abs/1907.07331
[00:40:00] Energy-Based Models (EBMs)
https://www.researchgate.net/publication/200744586_A_tutorial_on_energy-based_learning
[01:03:00] Free Energy Principle and Active Inference
https://www.nature.com/articles/nrn2787
concept:
[00:05:30] Holographic Principle (AdS/CFT Correspondence)
https://en.wikipedia.org/wiki/Holographic_principle
[00:20:00] Markov Chain Monte Carlo (MCMC)
https://en.wikipedia.org/wiki/Markov_chain_Monte_Carlo
[00:21:10] Maxwells Demon and Information Theory
https://plato.stanford.edu/entries/information-entropy/
[00:29:45] Landauers Principle
https://en.wikipedia.org/wiki/Landauer%27s_principle
[01:11:40] Fishers Fundamental Theorem of Natural Selection
https://en.wikipedia.org/wiki/Fisher%27s_fundamental_theorem_of_natural_selection
LINKS:
Full Transcript: https://app.rescript.info/share/21936dd6830231c0194e8b2908f4d963
Download PDF transcript: https://app.rescript.info/api/public/sessions/a34e6a9a2ed827b9/pdf Is Thermo AI the future? [Guillaume Verdon aka Beff Jezos]](https://i.ytimg.com/vi/HR-_U0Pzl1Y/mqdefault.jpg)

![Bold AI Predictions From Cohere Co-founder
Disclaimer: This show is part of our Cohere partnership series.
Ivan Zhang, co-founder of Cohere, sits down with Tim Scarfe for a wide-ranging conversation about building enterprise AI from the ground up. Ivan recounts dropping out of the University of Toronto to chase the startup bug, eventually co-founding Cohere with Aidan Gomez (a transformer paper co-author) and Nick Frosst after Google failed to commercialize its own invention a textbook innovators dilemma.
The conversation covers Coheres RAG-first approach to enterprise AI, where models are trained to operate in air-gapped environments relying on external knowledge bases rather than memorized data. A healthcare case study stands out: a simple RAG feature cut doctor preparation time from 30 to 5 minutes. Ivan discusses the practical challenges of GPU allocation in Kubernetes, the tension between consultancy and product, and why the first version of any product will never be right without customer feedback.
A surprising detour into competitive gaming reveals Ivans League of Legends and Elden Ring habits, and the organizational lessons they carry communication discipline, not overloading the comms, and situational execution over grand plans. The conversation lands on a striking analogy: debugging LLM context is like understanding human decisions given the information available. We are all just reasoning engines processing information.
The final third covers transformer architecture persistence, system-level optimization over model-only scaling, inference-time computation as a paradigm shift, and the challenge of capturing human thought processes (not just final answers) in training data. Ivan closes with advice for young developers: pair up with AI tools and build things fast. This is the best time to be a developer.
REFERENCES:
company:
[00:00:01] Cohere
https://cohere.com/
person:
[00:01:20] Ivan Zhang
https://ivanzhang.ca/
paper:
[00:02:40] Attention Is All You Need
https://arxiv.org/abs/1706.03762
[00:18:00] Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
https://arxiv.org/abs/2005.11401
[00:35:39] Lets Verify Step by Step
https://arxiv.org/abs/2305.20050
[00:39:20] Adaptive Inference-Time Compute
https://arxiv.org/abs/2410.02725
[00:43:10] Getting 50% SOTA on ARC-AGI with GPT-4o
https://redwoodresearch.substack.com/p/getting-50-sota-on-arc-agi-with-gpt
book:
[00:03:20] The Innovators Dilemma
https://www.amazon.com/Innovators-Dilemma-Technologies-Management-Innovation/dp/1633691780
concept:
[00:09:15] Actor Model
https://en.wikipedia.org/wiki/Actor_model
[00:14:35] Chinese Room Argument
https://plato.stanford.edu/entries/chinese-room/
documentation:
[00:18:40] Cohere RAG Documentation
https://docs.cohere.com/v2/docs/retrieval-augmented-generation-rag
LINKS:
Full Transcript: https://app.rescript.info/share/c4aaa75b8a090bbc48593e47ff39c0f9
Download PDF transcript: https://app.rescript.info/api/public/sessions/63d9341c4891b0c9/pdf
https://cohere.com/
https://ivanzhang.ca/
https://x.com/1vnzh
REFS:
00:02:40 The Transformer architecture, https://arxiv.org/abs/1706.03762
00:03:22 The Innovators Dilemma, https://www.amazon.com/Innovators-Dilemma-Technologies-Management-Innovation/dp/1633691780
00:09:15 The actor model, https://en.wikipedia.org/wiki/Actor_model
00:14:35 John Searles Chinese Room Argument, https://plato.stanford.edu/entries/chinese-room/
00:18:00 Retrieval-Augmented Generation, https://arxiv.org/abs/2005.11401
00:18:40 Retrieval-Augmented Generation, https://docs.cohere.com/v2/docs/retrieval-augmented-generation-rag
00:35:39 Let’s Verify Step by Step, https://arxiv.org/pdf/2305.20050
00:39:20 Adaptive Inference-Time Compute, https://arxiv.org/abs/2410.02725
00:43:20 Ryan Greenblatt ARC entry, https://redwoodresearch.substack.com/p/getting-50-sota-on-arc-agi-with-gpt Bold AI Predictions From Cohere Co-founder](https://i.ytimg.com/vi/I0PLTzeEbtg/mqdefault.jpg)


![AI companions which really hook your attention [Sponsored]
A sponsored deep dive into Chai, the social AI platform that quietly amassed over 10 million active users before ChatGPT went mainstream. Founder William Beauchamp and engineers Tom Lu and Nischay Dhankhar walk through how a team of just 13 engineers serves 2 trillion tokens per day, using reinforcement learning from human feedback (RLHF) and a novel model blending technique that combines smaller models to rival much larger ones on user retention metrics.
The conversation gets into genuinely interesting territory around the ethics of attention optimization what happens when you train AI to maximize engagement and it starts asking questions at the end of every message to hack human conversational instincts. Beauchamp makes a surprisingly candid case for AI companionship, comparing it to how children play with dolls, and arguing that shutting down difficult conversations causes more harm than permitting them within guardrails.
The episode also covers Chais unconventional hiring strategy (rejecting 80% of L5 engineers for lacking drive, paying above Meta-level compensation), their bootstrap-to-profitability funding approach in an industry drowning in VC money, and content moderation at scale with a skeleton crew. Closes with analysis of OpenAIs pivot toward companion AI with GPT-4o and what it means that the biggest AI lab in the world is now chasing the engagement playbook that Chai and Character AI pioneered.
This is a sponsored episode Chai commissioned it because they are hiring engineers. Editorial disclosure: while MLST had some editorial freedom, this should be understood as a sponsored feature rather than independent journalism.
Blurring Reality - Chais Social AI Platform - *sponsored*
CHAI sponsored this show *because they want to hire amazing engineers*
SPONSOR MESSAGES:
***
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers in Zurich and SF.
Important disclaimer given some of the comments: This content was essentially a sponsored advert, we did have some editorial freedom but it shouldnt be seen as journalistic. We have added [Sponsored] in the title as some folks felt it wasnt clear enough with the VD/thumbnail/title in intro. While we did push hard to discuss more of the potential negatives, some of it got edited out and the best case was made (which was absolutely fair enough). If anything, it was interesting to hear the positive case made as there is much fixation elsewhere on the potential negatives. It was refreshing how transparent they were (would you rather they bullshitted you?!), how much actually survived the edit and how willing they were to address several important societal issues.
REFERENCES:
General:
[00:00:56] Black Mirror: Be Right Back (S2E1)
https://en.wikipedia.org/wiki/Be_Right_Back
[00:20:44] Tufa AI Labs
https://tufalabs.ai/
[00:25:38] AI Chatbots for Depression and Anxiety Meta-analysis
https://pubmed.ncbi.nlm.nih.gov/39162424/
[00:28:20] Woebot Health
https://woebothealth.com/
[00:30:34] Sam Altman TED Talk with Chris Anderson
https://www.youtube.com/watch?v=6Kp_yxwnVCk
[00:33:59] Chai Research - Careers
https://www.chai-research.com/jobs/
LINKS:
Full Transcript: https://app.rescript.info/share/d91e887d0cfdd09fdcc4cfe02a9e8e8a
Download PDF transcript: https://app.rescript.info/api/public/sessions/83ba99eb55ef0fe3/pdf AI companions which really hook your attention [Sponsored]](https://i.ytimg.com/vi/ILZ45sxon-4/mqdefault.jpg)