Proximal Policy Optimization (PPO) - How to train Large Language Models @SerranoAcademy
Proximal Policy Optimization (PPO) - How to train Large Language Models  @SerranoAcademy
Uploaded January 2024 | Updated September 2026, 2 weeks ago
Reinforcement Learning with Human Feedback (RLHF) is a method used for training Large Language Models (LLMs). In the heart of RLHF lies a very powerful reinforcement learning method called Proximal Policy Optimization. Learn about it in this simple video!

This is the first one in a series of 3 videos dedicated to the reinforcement learning methods used for training LLMs.

Full Playlist: youtube.com/playlist?list=PLs8w1Cdi-zvYviYYw_V3qe6SINReGF5M-

Video 0 (Optional): Introduction to deep reinforcement learning youtube.com/watch?v=SgC6AZss478
Video 1 (This one): Proximal Policy Optimization
Video 2: Reinforcement Learning with Human Feedback youtube.com/watch?v=Z_JUqJBpVOk
Video 3 (Coming soon!): Deterministic Policy Optimization

00:00 Introduction
01:25 Gridworld
03:10 States and Action
04:01 Values
07:30 Policy
09:39 Neural Networks
16:14 Training the value neural network (Gain)
22:50 Training the policy neural network (Surrogate Objective Function)
33:38 Clipping the surrogate objective function
36:49 Summary

Get the Grokking Machine Learning book!
manning.com/books/grokking-machine-learning
Discount code (40%): serranoyt
(Use the discount code on checkout)
Proximal Policy Optimization (PPO) - How to train Large Language ModelsMath and OCD - My story with the Thue-Morse sequenceThank you for 100K subscribers! I’m planning tons of new content coming soon, so excited!A friendly introduction to Recurrent Neural NetworksThe math behind Attention: Keys, Queries, and Values matricesWill AI help us, or make us dependent? - A Tale of Two CitiesThe covariance matrixGRPO - Group Relative Policy Optimization  - How DeepSeek trains reasoning modelsReinforcement Learning with Human Feedback (RLHF) - How to train and fine-tune Transformer ModelsHow does Netflix recommend movies? Matrix FactorizationWhy do we divide by n-1 to estimate the variance? A visual tour through Bessel correctionMachine Learning: Testing and Error Metrics
Luis Serrano Academy |

Proximal Policy Optimization (PPO) - How to train Large Language Models

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER