Uploaded July 2025 | Updated September 2026, 2 weeks ago
Traditional reinforcement learning for LLMs hits a ceiling—curated math and coding problems can only take you so far. What if models could discover or invent reasoning challenges on their own?
In this talk, Nick Haber (Stanford University) explores two research breakthroughs from the Autonomous Agents Lab:
Quiet-STaR: A method for continued pretraining that inserts and rewards self-generated “thoughts” in natural text, improving downstream reasoning performance.
Minimo: A fully self-supervised theorem discovery and proving system, where a transformer trained from scratch invents its own math conjectures and learns to prove them—no human data required.
These projects redefine how we train LLMs for general reasoning, pushing beyond the limitations of human-curated benchmarks and toward truly open-ended cognitive capabilities.
🔬 Quiet-STaR Paper: Language Models Can Teach Themselves to Think Before Speaking – arxiv.org/abs/2403.09629
📘 Minimo Paper: Learning Formal Mathematics from Intrinsic Motivation – arxiv.org/abs/2407.00695
📍 Talk recorded at Arize:Observe 2025
🌐 Join Arize AI community events: arize.com/community
Full talk descripition from the event:
Traditional reinforcement learning for LLM reasoning relies on curated problem sets—mathematics benchmarks, coding challenges, and other structured tasks. This approach faces fundamental scalability limits: human-designed problems are finite, expensive to create, and may not capture the full breadth of reasoning required in real applications. We examine two of our works that expand beyond this paradigm: Quiet-STaR discovers reasoning opportunities in ordinary pretraining text, while Minimo generates its own mathematical conjectures for exploration. Both approaches transform the training environment itself—from solving predefined problems to finding and creating reasoning challenges autonomously. This shift suggests a path toward LLMs that can continuously discover new reasoning domains rather than being confined to human-curated problem distributions.
Traditional reinforcement learning for LLMs hits a ceiling—curated math and coding problems can only take you so far. What if models could discover or invent reasoning challenges on their own?
In this talk, Nick Haber (Stanford University) explores two research breakthroughs from the Autonomous Agents Lab:
Quiet-STaR: A method for continued pretraining that inserts and rewards self-generated “thoughts” in natural text, improving downstream reasoning performance.
Minimo: A fully self-supervised theorem discovery and proving system, where a transformer trained from scratch invents its own math conjectures and learns to prove them—no human data required.
These projects redefine how we train LLMs for general reasoning, pushing beyond the limitations of human-curated benchmarks and toward truly open-ended cognitive capabilities.
🔬 Quiet-STaR Paper: Language Models Can Teach Themselves to Think Before Speaking – arxiv.org/abs/2403.09629
📘 Minimo Paper: Learning Formal Mathematics from Intrinsic Motivation – arxiv.org/abs/2407.00695
📍 Talk recorded at Arize:Observe 2025
🌐 Join Arize AI community events: arize.com/community
Full talk descripition from the event:
Traditional reinforcement learning for LLM reasoning relies on curated problem sets—mathematics benchmarks, coding challenges, and other structured tasks. This approach faces fundamental scalability limits: human-designed problems are finite, expensive to create, and may not capture the full breadth of reasoning required in real applications. We examine two of our works that expand beyond this paradigm: Quiet-STaR discovers reasoning opportunities in ordinary pretraining text, while Minimo generates its own mathematical conjectures for exploration. Both approaches transform the training environment itself—from solving predefined problems to finding and creating reasoning challenges autonomously. This shift suggests a path toward LLMs that can continuously discover new reasoning domains rather than being confined to human-curated problem distributions.










