Stanford CS329A Self-Improving AI Agents | Part 6 | Train Time Scaling/Scaling RL @stanfordonline
Stanford CS329A Self-Improving AI Agents | Part 6 | Train Time Scaling/Scaling RL  @stanfordonline
Uploaded August 2026 | Updated September 2026, 2 weeks ago
Want to dive deeper? This curriculum is covered in the following online courses:
- Agentic AI professional education course: stanford.io/4zOyjPN
- XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents

A similar curriculum is covered in XCS329z https://online.stanford.edu/courses/cs329z-engineering-ai-agents

Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/

View the course playlist: youtube.com/playlist?list=PLangBM27OtEA

Video Summary:
This lecture video from Stanford's CS329A, Self-Improving AI Agents, taught by Aakanksha Chowdhery on October 10, 2025, covers train-time scaling and scaling reinforcement learning through three papers. STaR, the Self-Taught Reasoner, bootstraps reasoning chains through rationalization and filtering by answer correctness. DeepSeekMath introduces Group Relative Policy Optimization, a memory-efficient alternative to PPO, and shows gains from training on curated math data. DAPO addresses entropy collapse and training instability in reinforcement learning on long chain-of-thought reasoning through techniques including asymmetric clipping and dynamic sampling. Using the AIME math benchmark, the lecture traces how these methods let smaller models match the accuracy of much larger systems, and it closes with open questions on why majority-at-K accuracy improves while pass-at-K does not.

Speaker Bio:
Aakanksha Chowdhery
Adjunct Professor of Computer Science, Stanford University

Dr. Aakanksha Chowdhery is pushing the frontier of agentic LLMs, focusing on recursive self-improvement and long-horizon agents that learn and deploy in the real world. She is one of the few researchers globally who has led frontier model training end-to-end, across both dense and mixture-of-experts (MoE) architectures. At Google, she led the 540B PaLM model, the largest densely trained language model in the world at the time. She subsequently drove pre-training and scaling of Gemini's MoE models across multiple generations, and contributed key components to PaLM-E, Med-PaLM, and the Pathways infrastructure underpinning Google's large-model efforts. She went on to build and lead pretraining teams for open intelligence efforts at Reflection and Meta. Earlier, she held research roles at Microsoft Research and Princeton. At Stanford, where she earned her PhD, she teaches CS329A (Self-Improving AI Agents) and serves as Program Chair for MLSys 2026.
Stanford CS329A Self-Improving AI Agents | Part 6 | Train Time Scaling/Scaling RLStanford CS221 | Autumn 2025 | Lecture 1: Course Overview and AI FoundationsStanford CS547 HCI Seminar | Winter 2026 | Computational EcosystemsStanford CS547 HCI Seminar | Spring 2026 | Just-in-Time Objectives for Specialized AI InteractionsStanford CS221 | Autumn 2025 | Lecture 2: Learning IStanford CS229 Machine Learning | Spring 2026 | Lecture 5: Gaussian Discriminant AnalysisStanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 6: Direct MethodsStanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 17: RL Value-Based Methods
Stanford Online |

Stanford CS329A Self-Improving AI Agents | Part 6 | Train Time Scaling/Scaling RL

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER