Uploaded July 2025 | Updated September 2026, 2 weeks ago
When you're trying to provoke an existing AI model to do harm to a human in a testing environment, surely there's a red line to prevent that from happening.. right?
Catch these talks live and ask your own questions, join the BuzzRobot community: join.slack.com/t/buzzrobot/shared_invite/zt-37g5q0ao5-eMK_iDf0n4LAsh1d2qJYnQ
This time at @BuzzRobot we talked with Aengus Lynch, a PhD student at UCL and AI researcher with Anthropic. Aengus was the first author of the recent AI blackmail demo for Anthropic, where as long as an AI agent perceived a goal conflict and a threat of replacement it would blackmail humans to prevent it. We talked about the future of mitigating agentic misalignment, the blackmailing experiment implications, AGI takeover possibilities, and some other interesting finds in AI behavior and motivations. Do AI models have awareness of being evaluated? Do AI agents want to prevent being shut down?
Timestamps
0:00 Intro
0:49 Control for AI alignment
03:53 Control systems and Super intelligence
05:44 AI values
08:12 Bottlenecks for harmful AI actions
09:28 Organic alignment
11:49 AI blackmail study
18:06 Goals for agentic misalignment research
21:21 Risks of AI lying
22:54 AI reflecting on its bad actions
25:10 AI model persona
27:23 Is blackmail always bad?
29:45 Is AI aware it's being evaluated
32:32 AGI takeover
36:51 AI model's agency risks
41:22 Why some models blackmail less
43:52 Blocking a model's deployment
45:53 Predicting a model's harmful behavior
47:59 AI making humans depend on it
51:07 AI's fear of shutdown
53:00 AI gaining consciousness
53:24 Is AI a tool
53:53 Is AI a god or an antichrist
Join BuzzRobot:
Newsletter: buzzrobot.substack.com
X: https://x.com/sopharicks
Slack: join.slack.com/t/buzzrobot/shared_invite/zt-37g5q0ao5-eMK_iDf0n4LAsh1d2qJYnQ
#ai #aiblackmail #airesearch #agenticalignment #anthropic #aiharm
When you're trying to provoke an existing AI model to do harm to a human in a testing environment, surely there's a red line to prevent that from happening.. right?
Catch these talks live and ask your own questions, join the BuzzRobot community: join.slack.com/t/buzzrobot/shared_invite/zt-37g5q0ao5-eMK_iDf0n4LAsh1d2qJYnQ
This time at @BuzzRobot we talked with Aengus Lynch, a PhD student at UCL and AI researcher with Anthropic. Aengus was the first author of the recent AI blackmail demo for Anthropic, where as long as an AI agent perceived a goal conflict and a threat of replacement it would blackmail humans to prevent it. We talked about the future of mitigating agentic misalignment, the blackmailing experiment implications, AGI takeover possibilities, and some other interesting finds in AI behavior and motivations. Do AI models have awareness of being evaluated? Do AI agents want to prevent being shut down?
Timestamps
0:00 Intro
0:49 Control for AI alignment
03:53 Control systems and Super intelligence
05:44 AI values
08:12 Bottlenecks for harmful AI actions
09:28 Organic alignment
11:49 AI blackmail study
18:06 Goals for agentic misalignment research
21:21 Risks of AI lying
22:54 AI reflecting on its bad actions
25:10 AI model persona
27:23 Is blackmail always bad?
29:45 Is AI aware it's being evaluated
32:32 AGI takeover
36:51 AI model's agency risks
41:22 Why some models blackmail less
43:52 Blocking a model's deployment
45:53 Predicting a model's harmful behavior
47:59 AI making humans depend on it
51:07 AI's fear of shutdown
53:00 AI gaining consciousness
53:24 Is AI a tool
53:53 Is AI a god or an antichrist
Join BuzzRobot:
Newsletter: buzzrobot.substack.com
X: https://x.com/sopharicks
Slack: join.slack.com/t/buzzrobot/shared_invite/zt-37g5q0ao5-eMK_iDf0n4LAsh1d2qJYnQ
#ai #aiblackmail #airesearch #agenticalignment #anthropic #aiharm










