How Researchers Test AI for Hidden Goals — Apollo Research @MachineLearningStreetTalk
How Researchers Test AI for Hidden Goals — Apollo Research  @MachineLearningStreetTalk
Uploaded July 2026 | Updated September 2026, 1 week ago
Can an AI do the right thing for the wrong reason? Tim Scarfe speaks with Apollo Research’s Alexander Meinke, Axel Højmark and Jérémy Scheurer about Measuring Reward-Seeking via Contrastive Belief Updates, their new research with OpenAI.

The panel asks how models infer what graders reward, why good behaviour can come from the wrong reason, and whether that difference can be measured. The conversation moves through promise-breaking, grader awareness, reward hacking, scheming, opaque reasoning and corrigibility, then turns to a detailed walkthrough of the contrastive-belief method and what its results do and do not show. The o3 results discussed here concern an intermediate checkpoint without safety training.

This episode was made in partnership with Apollo Research. MLST retained full editorial control.

Reference
Apollo Research: apolloresearch.ai

---
TIMESTAMPS:
00:00:00 Cold Open
00:02:12 Right Things, Wrong Reasons
00:12:47 Grader Awareness
00:26:22 Legibility
00:32:35 What To Call It
00:35:58 Intelligence, Agency, Anthropomorphism
00:45:16 Apollo’s Mission
00:48:54 The End of the Exponential
00:55:45 The Paper
01:16:34 Closing Reflection

---
REFERENCES:
tool:
[00:00:08] Claude Fable
anthropic.com/claude/fable
[00:12:50] AlphaGo Zero
https://deepmind.google/blog/alphago-zero-starting-from-scratch/
[00:44:30] AlphaFold 3
https://deepmind.google/science/alphafold/
paper:
[00:01:02] Measuring Reward-Seeking via Contrastive Belief Updates
arxiv.org/abs/2607.18966
[00:16:19] Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
https://transformer-circuits.pub/2026/nla/
[00:26:48] Stress Testing Deliberative Alignment for Anti-Scheming Training
arxiv.org/abs/2509.15541
[00:35:33] Shortcut learning in deep neural networks
arxiv.org/abs/2004.07780
[00:53:49] Measuring AI Ability to Complete Long Software Tasks
arxiv.org/abs/2503.14499
[00:59:52] Modifying LLM Beliefs with Synthetic Document Finetuning
alignment.anthropic.com/2025/modifying-beliefs-via-sdf
[01:10:44] Alignment Faking in Large Language Models
arxiv.org/abs/2412.14093
[01:13:55] Natural Emergent Misalignment from Reward Hacking
anthropic.com/research/emergent-misalignment-reward-hacking
other:
[00:10:14] We Need a Science of Scheming
apolloresearch.ai/science/science-of-scheming
[00:32:56] CoastRunners reward hacking example
https://deepmind.google/blog/specification-gaming-the-flip-side-of-ai-ingenuity/
organization:
[01:06:07] Redwood Research
redwoodresearch.org

---
ReScript:
app.rescript.info/share/718ab68e18cfa3b9b800da6b3290fd42
How Researchers Test AI for Hidden Goals — Apollo ResearchLanguage Models are Modelling The World [Nicholas Carlini]Cohere is not an AGI company - Nick Frosst (co-founder)AI AGENCY ISNT HERE YET... (Dr. Philip Ball)Are We Just Machines? (Mazviita Chirimuuta)Is o1-preview reasoning?Prof Nick Chater on our mysterious brains modelling the worldWill software synthesis replace machine learning?Why does the Chinese Room still haunt AI?Why Every AI Model Is an Impostor — Kenneth StanleyTaming Silicon Valley - Prof. Gary MarcusWhy Program Synthesis Is Next (Kevin Ellis and Zenna Tavares)
Machine Learning Street Talk |

How Researchers Test AI for Hidden Goals — Apollo Research

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER