Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking @allenai
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking  @allenai
Uploaded February 2024 | Updated September 2026, 4 days ago
Abstract: Reward models are commonly used in the process of large language model alignment but are prone to reward hacking, where the true reward diverges from the estimated reward as the language model drifts out-of-distribution. In this talk, I will discuss a recent study on the use of reward ensembles to mitigate reward hacking. The study demonstrates that reward models that originate from different pretrain seeds are effective at mitigating reward hacking, but when errors of ensemble members correlate, the ensemble is likely to inherit the same behaviour, in which case reward hacking persists. I hope to spend some time discussing alternative approaches for uncertainty estimation and to discuss the role of exploration in language model alignment.

Bio: Jonathan Berant is an associate professor at the School of Computer Science at Tel Aviv University, currently on sabbatical as a visiting faculty researcher at Google DeepMind. Jonathan earned a Ph.D. in Computer Science at Tel-Aviv University, and was a post-doctoral fellow at Stanford University, and subsequently a post-doctoral fellow at Google Research, Mountain View. Jonathan Received several awards and fellowships including The Rothschild fellowship, The ACL 2011 best student paper award, EMNLP 2014 best paper award, NAACL 2019 best resource paper award, and several honorable mentions. Jonathan has won the Kadar prize for outstanding research and is currently an ERC grantee.
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward HackingTowards the Age of Computation and AI for High Performance ClimateWhen Not to Trust Language Models: Investigating Effectiveness of Parametric&Non-Parametric MemoriesOpen-Ended Learning Leads to Generally Capable Agents | Embodied AI Lecture Series at AI2From LLMs to Agents: Generalizability from the Inside OutData-Centric Approaches to Adapting Foundation ModelsThe University of Washington eScience Institute: a Home for Data-Intensive DiscoveryGeneralization for Robot Learning In The Wild | Embodied AI Lecture series at AI2Towards Generalist Agents for Accelerating Scientific Discovery171Modular Language ModelsTowards robust long-form text generation systemsDigital Socrates: Evaluating LLMs through Explanation Critiques
Ai2 |

Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER