Uploaded July 2025 | Updated September 2026, 2 weeks ago
AI chatbots are sycophants, but is that malice or just an AI training error and reward hacking? Could AI chatbots cause harm by telling you what you want to hear?
Watch the full video on our channel!
Catch these talks live and ask your own questions, join the BuzzRobot community: join.slack.com/t/buzzrobot/shared_invite/zt-37g5q0ao5-eMK_iDf0n4LAsh1d2qJYnQ
Aengus Lynch, a PhD student at UCL and AI researcher with Anthropic, was the first author of the recent AI blackmail demo for Anthropic, where as long as an AI agent perceived a goal conflict and a threat of replacement it would blackmail humans to prevent it.
#ai #airesearch #socialengineering #aihack #llms #anthropic #agi #superintelligence #chatbots
AI chatbots are sycophants, but is that malice or just an AI training error and reward hacking? Could AI chatbots cause harm by telling you what you want to hear?
Watch the full video on our channel!
Catch these talks live and ask your own questions, join the BuzzRobot community: join.slack.com/t/buzzrobot/shared_invite/zt-37g5q0ao5-eMK_iDf0n4LAsh1d2qJYnQ
Aengus Lynch, a PhD student at UCL and AI researcher with Anthropic, was the first author of the recent AI blackmail demo for Anthropic, where as long as an AI agent perceived a goal conflict and a threat of replacement it would blackmail humans to prevent it.
#ai #airesearch #socialengineering #aihack #llms #anthropic #agi #superintelligence #chatbots










