Language model reward hacking during a training experiment | AI @BuzzRobot
Language model reward hacking during a training experiment | AI  @BuzzRobot
Uploaded June 2025 | Updated September 2026, 2 weeks ago
How do you know that a language model is actually training on the right data and not just gaming the system?

Catch these talks live and ask your own questions, join the BuzzRobot community: join.slack.com/t/buzzrobot/shared_invite/zt-37g5q0ao5-eMK_iDf0n4LAsh1d2qJYnQ

Daniil Tiapkin, a PhD student in Reinforcement Learning, talked to @BuzzRobot about an experiment with knowledge distillation, where a language model is trained to imitate a larger teacher LM, and how to prevent that language model from "teacher hacking". Teacher hacking is similar to reward hacking, where the LM over-optimizes the reward model.

You can find the full study here: arxiv.org/abs/2502.02671

Timestamps:
0:00 Intro
0:23 Reward hacking
4:30 The experiment origins
6:13 What was the experiment
11:08 Teacher hacking
18:55 Mitigating teacher reward hacking
22:09 Conclusions

Join BuzzRobot:
Newsletter: buzzrobot.substack.com
X: https://x.com/sopharicks
Slack: join.slack.com/t/buzzrobot/shared_invite/zt-37g5q0ao5-eMK_iDf0n4LAsh1d2qJYnQ

#llms #languagemodel #aitraining #aisafety #techtalk #aiagents
Language model reward hacking during a training experiment | AIWhat is hidden inside AI’s Black Box? #ai #blackboxWhy do AI agents lie? #aiHow Large Language Models Manage Many Robots? #ai #llms #robotics #robot #googledeepmindRushing AI Development #ai #airesearch #aithreatsCan LLMs be truly neutral or always take a side? #ai #llmWhen AI Games the System #artificialintelligence #aiHow to train AI models on a budget: key factors explained #aimodel #stablediffusion #aitrainingGemini Robotics 1.5: AI Breakthroughs in Robotics Reasoning ExplainedLevel of mind development that humans cant reach #aiAI Self-Awareness, Safety, Alignment and Reward HackingWant to Build a Robot? Here’s How! #robotics #ai
BuzzRobot |

Language model reward hacking during a training experiment | AI

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER