AI Self-Awareness, Safety, Alignment and Reward Hacking @BuzzRobot
AI Self-Awareness, Safety, Alignment and Reward Hacking  @BuzzRobot
Uploaded May 2025 | Updated September 2026, 2 weeks ago
In this talk, Lawrence Chan from METR.org joins the @BuzzRobot community to share several recent and important research works on #AGI and #AIsafety.

The first paper, co-authored with researchers from @OpenAI and @oxforduniversity, explores critical #AGI alignment issues. The authors argue that AGI could learn to develop its own objectives, pursue conflicting goals, act deceptively to receive higher rewards, and form misaligned internal goals that generalize beyond its training data - potentially leading to power-seeking behavior.

This research expands on emerging evidence for these concerning properties in advanced #AI systems.

Another key paper from METR.org highlights that AI agents’ ability to perform long-term tasks is growing exponentially - doubling every seven months. This suggests we are moving toward AGI at an accelerating pace.

​In this ask me anything session, are covering these studies and explore broader questions around #AIsafety, #alignment, and the future of #AGI.

Research Papers:
The Alignment Problem from a Deep Learning Perspective: arxiv.org/pdf/2209.00626
Measuring AI Ability to Complete Long Tasks: metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks

Timestamps:
0:00 Introduction
3:48 What does "self-awareness" or "situational awareness" in language models mean?
6:54 Why are LLMs now better at understanding themselves?
11:33 Can LLMs develop subjective experience or consciousness?
14:37 Do research labs let AIs express distress - or is it being ignored?
16:20 Is compute the key driver towards AGI, or will we hit a ceiling?
24:45 When will we hit real-world compute limitations?
26:51 The challenge of long-term memory in language models
31:08 Why does it matter if agents use self-related info without being prompted?
33:51 Deceptive alignment: how AIs might fake good behavior during training
36:47 How risky are deceptive agents as they scale?
41:15 What is reward hacking and why is it a major concern?
44:41 AGI persuasion, emotional manipulation, and novel weapon risks - what strategies matter most?
52:12 METR.org's study: agent capability is doubling every 7 months - what does this trend mean for the future?

#ai #agi #aisafety #alignment #rewardsystems #metr #openai #oxforduniversity #artificialintelligence #artificialgeneralintelligence #artificialsuperintelligence #aialignment #aisafety #chatgpt #deeplearning #machinelearning #machineconsciousness #aiconsciousness #largelanguagemodels #llms #aireasoning #airesearch #techtalk #techtalks #aitalks #aitalk #science

Social Links:
Newsletter: buzzrobot.substack.com
X: https://x.com/sopharicks
Slack: join.slack.com/t/buzzrobot/shared_invite/zt-2s067rv7n-guPIMGe62rbp9ncxdnOUfQ
AI Self-Awareness, Safety, Alignment and Reward HackingWant to Build a Robot? Here’s How! #robotics #aiWhat are advanced AI assistants? #aiassistant #aiWhat is AGI? #aiAGI Takeover in 2 Years?How to survive SuperintelligenceHow Google Detects AI-Generated Images?  #ai #googledeepmindAI That Identifies Quantum Computing Errors: AlphaQubit by Google DeepMindWhat is AI consciousness?AI Models Mirror Your Mind - Heres How!How to Train Stable Diffusion for $2,000!The main problem with open sourcing AI models #ai
BuzzRobot |

AI Self-Awareness, Safety, Alignment and Reward Hacking

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER