Uploaded July 2025 | Updated September 2026, 2 weeks ago
Do AI models have rules to not hurt humans?
Watch the full video on our channel!
Catch these talks live and ask your own questions, join the BuzzRobot community: join.slack.com/t/buzzrobot/shared_invite/zt-37g5q0ao5-eMK_iDf0n4LAsh1d2qJYnQ
Aengus Lynch, a PhD student at UCL and AI researcher with Anthropic, was the first author of the recent AI blackmail demo for Anthropic, where as long as an AI agent perceived a goal conflict and a threat of replacement it would blackmail humans to prevent it.
Do AI models have rules to not hurt humans?
Watch the full video on our channel!
Catch these talks live and ask your own questions, join the BuzzRobot community: join.slack.com/t/buzzrobot/shared_invite/zt-37g5q0ao5-eMK_iDf0n4LAsh1d2qJYnQ
Aengus Lynch, a PhD student at UCL and AI researcher with Anthropic, was the first author of the recent AI blackmail demo for Anthropic, where as long as an AI agent perceived a goal conflict and a threat of replacement it would blackmail humans to prevent it.










