Uploaded February 2024 | Updated September 2026, 4 days ago
Abstract: Reward models are commonly used in the process of large language model alignment but are prone to reward hacking, where the true reward diverges from the estimated reward as the language model drifts out-of-distribution. In this talk, I will discuss a recent study on the use of reward ensembles to mitigate reward hacking. The study demonstrates that reward models that originate from different pretrain seeds are effective at mitigating reward hacking, but when errors of ensemble members correlate, the ensemble is likely to inherit the same behaviour, in which case reward hacking persists. I hope to spend some time discussing alternative approaches for uncertainty estimation and to discuss the role of exploration in language model alignment.
Bio: Jonathan Berant is an associate professor at the School of Computer Science at Tel Aviv University, currently on sabbatical as a visiting faculty researcher at Google DeepMind. Jonathan earned a Ph.D. in Computer Science at Tel-Aviv University, and was a post-doctoral fellow at Stanford University, and subsequently a post-doctoral fellow at Google Research, Mountain View. Jonathan Received several awards and fellowships including The Rothschild fellowship, The ACL 2011 best student paper award, EMNLP 2014 best paper award, NAACL 2019 best resource paper award, and several honorable mentions. Jonathan has won the Kadar prize for outstanding research and is currently an ERC grantee.
Abstract: Reward models are commonly used in the process of large language model alignment but are prone to reward hacking, where the true reward diverges from the estimated reward as the language model drifts out-of-distribution. In this talk, I will discuss a recent study on the use of reward ensembles to mitigate reward hacking. The study demonstrates that reward models that originate from different pretrain seeds are effective at mitigating reward hacking, but when errors of ensemble members correlate, the ensemble is likely to inherit the same behaviour, in which case reward hacking persists. I hope to spend some time discussing alternative approaches for uncertainty estimation and to discuss the role of exploration in language model alignment.
Bio: Jonathan Berant is an associate professor at the School of Computer Science at Tel Aviv University, currently on sabbatical as a visiting faculty researcher at Google DeepMind. Jonathan earned a Ph.D. in Computer Science at Tel-Aviv University, and was a post-doctoral fellow at Stanford University, and subsequently a post-doctoral fellow at Google Research, Mountain View. Jonathan Received several awards and fellowships including The Rothschild fellowship, The ACL 2011 best student paper award, EMNLP 2014 best paper award, NAACL 2019 best resource paper award, and several honorable mentions. Jonathan has won the Kadar prize for outstanding research and is currently an ERC grantee.










