Uploaded August 2026 | Updated September 2026, 2 weeks ago
Want to dive deeper? This curriculum is covered in the following online courses:
- Agentic AI professional education course: stanford.io/4zOyjPN
- XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents
A similar curriculum is covered in XCS329z https://online.stanford.edu/courses/cs329z-engineering-ai-agents
Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/
View the course playlist: youtube.com/playlist?list=PLangBM27OtEA
Video Summary:
This lecture video from Stanford's CS329A, Self-Improving AI Agents, taught by Aakanksha Chowdhery on October 10, 2025, covers train-time scaling and scaling reinforcement learning through three papers. STaR, the Self-Taught Reasoner, bootstraps reasoning chains through rationalization and filtering by answer correctness. DeepSeekMath introduces Group Relative Policy Optimization, a memory-efficient alternative to PPO, and shows gains from training on curated math data. DAPO addresses entropy collapse and training instability in reinforcement learning on long chain-of-thought reasoning through techniques including asymmetric clipping and dynamic sampling. Using the AIME math benchmark, the lecture traces how these methods let smaller models match the accuracy of much larger systems, and it closes with open questions on why majority-at-K accuracy improves while pass-at-K does not.
Speaker Bio:
Aakanksha Chowdhery
Adjunct Professor of Computer Science, Stanford University
Dr. Aakanksha Chowdhery is pushing the frontier of agentic LLMs, focusing on recursive self-improvement and long-horizon agents that learn and deploy in the real world. She is one of the few researchers globally who has led frontier model training end-to-end, across both dense and mixture-of-experts (MoE) architectures. At Google, she led the 540B PaLM model, the largest densely trained language model in the world at the time. She subsequently drove pre-training and scaling of Gemini's MoE models across multiple generations, and contributed key components to PaLM-E, Med-PaLM, and the Pathways infrastructure underpinning Google's large-model efforts. She went on to build and lead pretraining teams for open intelligence efforts at Reflection and Meta. Earlier, she held research roles at Microsoft Research and Princeton. At Stanford, where she earned her PhD, she teaches CS329A (Self-Improving AI Agents) and serves as Program Chair for MLSys 2026.
Want to dive deeper? This curriculum is covered in the following online courses:
- Agentic AI professional education course: stanford.io/4zOyjPN
- XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents
A similar curriculum is covered in XCS329z https://online.stanford.edu/courses/cs329z-engineering-ai-agents
Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/
View the course playlist: youtube.com/playlist?list=PLangBM27OtEA
Video Summary:
This lecture video from Stanford's CS329A, Self-Improving AI Agents, taught by Aakanksha Chowdhery on October 10, 2025, covers train-time scaling and scaling reinforcement learning through three papers. STaR, the Self-Taught Reasoner, bootstraps reasoning chains through rationalization and filtering by answer correctness. DeepSeekMath introduces Group Relative Policy Optimization, a memory-efficient alternative to PPO, and shows gains from training on curated math data. DAPO addresses entropy collapse and training instability in reinforcement learning on long chain-of-thought reasoning through techniques including asymmetric clipping and dynamic sampling. Using the AIME math benchmark, the lecture traces how these methods let smaller models match the accuracy of much larger systems, and it closes with open questions on why majority-at-K accuracy improves while pass-at-K does not.
Speaker Bio:
Aakanksha Chowdhery
Adjunct Professor of Computer Science, Stanford University
Dr. Aakanksha Chowdhery is pushing the frontier of agentic LLMs, focusing on recursive self-improvement and long-horizon agents that learn and deploy in the real world. She is one of the few researchers globally who has led frontier model training end-to-end, across both dense and mixture-of-experts (MoE) architectures. At Google, she led the 540B PaLM model, the largest densely trained language model in the world at the time. She subsequently drove pre-training and scaling of Gemini's MoE models across multiple generations, and contributed key components to PaLM-E, Med-PaLM, and the Pathways infrastructure underpinning Google's large-model efforts. She went on to build and lead pretraining teams for open intelligence efforts at Reflection and Meta. Earlier, she held research roles at Microsoft Research and Princeton. At Stanford, where she earned her PhD, she teaches CS329A (Self-Improving AI Agents) and serves as Program Chair for MLSys 2026.






