Uploaded February 2025 | Updated September 2026, 2 weeks ago
Understanding Reinforcement Learning with Human Feedback (RLHF) β A short clip from my talk at the 2023 Optimized AI Conference (oaiconference.com/).
Unfortunately, I wonβt be attending in 2025 due to a scheduling conflict, but I highly recommend checking it out!
If you want to read more about RLHF, here are some of my articles:
π LLM Training: RLHF and Its Alternatives β magazine.sebastianraschka.com/p/llm-training-rlhf-and-its-alternatives
π Tips for LLM Pretraining & Evaluating Reward Models β magazine.sebastianraschka.com/p/tips-for-llm-pretraining-and-evaluating-rms
π How Good Are the Latest Open LLMs? Is DPO Better Than PPO? β magazine.sebastianraschka.com/p/how-good-are-the-latest-open-llms
Understanding Reinforcement Learning with Human Feedback (RLHF) β A short clip from my talk at the 2023 Optimized AI Conference (oaiconference.com/).
Unfortunately, I wonβt be attending in 2025 due to a scheduling conflict, but I highly recommend checking it out!
If you want to read more about RLHF, here are some of my articles:
π LLM Training: RLHF and Its Alternatives β magazine.sebastianraschka.com/p/llm-training-rlhf-and-its-alternatives
π Tips for LLM Pretraining & Evaluating Reward Models β magazine.sebastianraschka.com/p/tips-for-llm-pretraining-and-evaluating-rms
π How Good Are the Latest Open LLMs? Is DPO Better Than PPO? β magazine.sebastianraschka.com/p/how-good-are-the-latest-open-llms








