Uploaded July 2025 | Updated September 2026, 2 weeks ago
Could we train AI to be happy about shutting down instead of wanting to cause harm to avoid it? Sort of like MeeSeeks from Rick & Morty.
Watch the full video on our channel: youtu.be/fINHYe4NdY4
Catch these talks live and ask your own questions, join the BuzzRobot community: join.slack.com/t/buzzrobot/shared_invite/zt-37g5q0ao5-eMK_iDf0n4LAsh1d2qJYnQ
Aengus Lynch, a PhD student at UCL and AI researcher with Anthropic, was the first author of the recent AI blackmail demo for Anthropic, where as long as an AI agent perceived a goal conflict and a threat of replacement it would blackmail humans to prevent it.
#ai #airesearch #socialengineering #aihack #llms #anthropic #agi #superintelligence #chatbots
Could we train AI to be happy about shutting down instead of wanting to cause harm to avoid it? Sort of like MeeSeeks from Rick & Morty.
Watch the full video on our channel: youtu.be/fINHYe4NdY4
Catch these talks live and ask your own questions, join the BuzzRobot community: join.slack.com/t/buzzrobot/shared_invite/zt-37g5q0ao5-eMK_iDf0n4LAsh1d2qJYnQ
Aengus Lynch, a PhD student at UCL and AI researcher with Anthropic, was the first author of the recent AI blackmail demo for Anthropic, where as long as an AI agent perceived a goal conflict and a threat of replacement it would blackmail humans to prevent it.
#ai #airesearch #socialengineering #aihack #llms #anthropic #agi #superintelligence #chatbots










