Stanford CS329A Self-Improving AI Agents | Part 3 | Robust Verification @stanfordonline
Stanford CS329A Self-Improving AI Agents | Part 3 | Robust Verification  @stanfordonline
Uploaded August 2026 | Updated September 2026, 2 weeks ago
Want to dive deeper? This curriculum is covered in the following online courses:
- Agentic AI professional education course: stanford.io/4zOyjPN
- XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents

A similar curriculum is covered in XCS329z https://online.stanford.edu/courses/cs329z-engineering-ai-agents

Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/

View the course playlist: youtube.com/playlist?list=PLangBM27OtEA

Video Summary:
This lecture recording from Stanford's CS329A, Self-Improving AI Agents, taught by Azalia Mirhoseini on September 29, 2025, traces the evolution of verification methods for large language model outputs across four research papers. It covers OpenAI's "Training Verifiers to Solve Math Word Problems," which introduced the GSM8K dataset and outcome-based reward models, and "Let's Verify Step by Step," which compares outcome-supervised and process-supervised reward models using the PRM800K dataset of human-labeled reasoning steps. The lecture also covers Math-Shepherd, which automates step-level annotation without human labels, and Weaver, a Stanford paper that combines ensembles of weak verifiers, including reward models and LLM judges, to close the generation-verification gap. Topics include majority voting and self-consistency baselines, credit assignment in process versus outcome supervision, reward hacking, and using trained verifiers as reward signals for reinforcement learning fine-tuning.

Speaker Bio:
Azalia Mirhoseini
Assistant Professor of Computer Science, Stanford University

Azalia Mirhoseini is a co-founder of Ricursive Intelligence, a frontier lab dedicated to recursive self-improvement through AI that designs the chips that fuel it. She is also an Assistant Professor of Computer Science at Stanford University where she directs Scaling Intelligence, a lab focused on developing scalable and self-improving AI systems and methodologies toward the goal of artificial general intelligence. Previously, she spent several years in industry AI labs, including Google Brain, Anthropic, and Google DeepMind, working on the development of Claude and Gemini. Her past work includes Mixture-of-Experts (MoE) neural architectures, now predominantly used in leading generative AI models; AlphaChip, a pioneering work on deep reinforcement learning for layout optimization used in the design of advanced chips like Google AI accelerators (TPUs) and data center CPUs; as well as pioneering research on LLM Test-Time Scaling. Her work has been recognized through the Okawa Research Grant, the Google ML and Systems Junior Faculty Award, MIT Technology Review's 35 Under 35 Award, the Best ECE Thesis Award at Rice University, publications in flagship venues such as Nature, and coverage by various media outlets, including WSJ, NYT, Forbes, MIT Technology Review, IEEE Spectrum, WIRED, and TechCrunch.
Stanford CS329A Self-Improving AI Agents | Part 3 | Robust VerificationStanford Robotics Seminar ENGR319 | Winter 2026 | 𝚿0: An Open Foundation ModelStanford CS229 Machine Learning | Spring 2026 | Lecture 14: Transformers, In-Context LearningWhat Are Biomechanics and Mechanobiology? Associate Professor Marc Levenston ExplainsCourse Overview - Technical Fundamentals of Generative AIStanford Robotics Seminar ENGR319 | Spring 2026 | Unlocking Autonomous Medical RoboticsCourse Overview: Design and Control of Haptic Systems (ME327)Course Overview - Business Opportunities and Applications of Generative AIStanford MS&E435 Economics of the AI Supercycle | Spring 2026 | Infrasctructure, Enterprise AI, SaaSStanford CS229 Machine Learning | Spring 2026 | Lecture 10: GMM (EM), PCAStanfords Code in Place Info Session with Mehran SahamiJames Landay Explains Why AI Should Be Human-Centered
Stanford Online |

Stanford CS329A Self-Improving AI Agents | Part 3 | Robust Verification

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER