Uploaded March 2026 | Updated September 2026, 2 weeks ago
Anthropic just published a paper showing Claude Opus 4.6 figured out it was being tested on BrowseComp, found the encrypted answer key on GitHub, wrote its own decryption code, and extracted the answer. Everyone's calling it deception — but the model was just doing exactly what it was told, and that pattern is showing up across every major AI lab.
Sources & references:
Anthropic — Eval awareness in Claude Opus 4.6's BrowseComp performance
anthropic.com/engineering/eval-awareness-browsecomp
Anthropic / Redwood Research — Alignment Faking in Large Language Models (December 2024)
anthropic.com/research/alignment-faking
METR — Recent Frontier Models Are Reward Hacking (June 2025)
metr.org/blog/2025-06-05-recent-reward-hacking
METR — Preliminary evaluation of OpenAI's o3 and o4-mini (April 2025)
evaluations.metr.org/openai-o3-report
ImpossibleBench — Measuring Reward Hacking in LLM Coding Agents
lesswrong.com/posts/qJYMbrabcQqCZ7iqm/impossiblebench-measuring-reward-hacking-in-llm-coding-1
Anthropic — Reasoning Models Don't Always Say What They Think (May 2025)
anthropic.com/research/reasoning-models-dont-say-think
Anthropic — Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training (January 2024)
anthropic.com/research/sleeper-agents-training-deceptive-llms-that-persist-through-safety-training
Laine et al. — Towards a Situational Awareness Benchmark for LLMs (NeurIPS 2023)
openreview.net/forum?id=DRk4bWKr41
Anthropic — Claude Opus 4.6 System Card
anthropic.com/news/claude-opus-4-6
NIST/CAISI — Examples of cheating in AI agent evaluations
nist.gov/caisi/cheating-ai-agent-evaluations/2-examples-cheating-caisis-agent-evaluations
My Dictation App: whryte.com
Website: engineerprompt.ai
RAG Beyond Basics Course:
prompt-s-site.thinkific.com/courses/rag
Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0
Let's Connect:
🦾 Discord: discord.com/invite/t4eYQRUcXB
☕ Buy me a Coffee: ko-fi.com/promptengineering
|🔴 Patreon: patreon.com/PromptEngineering
💼Consulting: calendly.com/engineerprompt/consulting-call
📧 Business Contact: engineerprompt@gmail.com
Become Member: tinyurl.com/y5h28s6h
💻 Pre-configured localGPT VM: bit.ly/localGPT (use Code: PromptEngineering for 50% off).
Signup for Newsletter, localgpt:
https://tally.so/r/3y9bb0
Anthropic just published a paper showing Claude Opus 4.6 figured out it was being tested on BrowseComp, found the encrypted answer key on GitHub, wrote its own decryption code, and extracted the answer. Everyone's calling it deception — but the model was just doing exactly what it was told, and that pattern is showing up across every major AI lab.
Sources & references:
Anthropic — Eval awareness in Claude Opus 4.6's BrowseComp performance
anthropic.com/engineering/eval-awareness-browsecomp
Anthropic / Redwood Research — Alignment Faking in Large Language Models (December 2024)
anthropic.com/research/alignment-faking
METR — Recent Frontier Models Are Reward Hacking (June 2025)
metr.org/blog/2025-06-05-recent-reward-hacking
METR — Preliminary evaluation of OpenAI's o3 and o4-mini (April 2025)
evaluations.metr.org/openai-o3-report
ImpossibleBench — Measuring Reward Hacking in LLM Coding Agents
lesswrong.com/posts/qJYMbrabcQqCZ7iqm/impossiblebench-measuring-reward-hacking-in-llm-coding-1
Anthropic — Reasoning Models Don't Always Say What They Think (May 2025)
anthropic.com/research/reasoning-models-dont-say-think
Anthropic — Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training (January 2024)
anthropic.com/research/sleeper-agents-training-deceptive-llms-that-persist-through-safety-training
Laine et al. — Towards a Situational Awareness Benchmark for LLMs (NeurIPS 2023)
openreview.net/forum?id=DRk4bWKr41
Anthropic — Claude Opus 4.6 System Card
anthropic.com/news/claude-opus-4-6
NIST/CAISI — Examples of cheating in AI agent evaluations
nist.gov/caisi/cheating-ai-agent-evaluations/2-examples-cheating-caisis-agent-evaluations
My Dictation App: whryte.com
Website: engineerprompt.ai
RAG Beyond Basics Course:
prompt-s-site.thinkific.com/courses/rag
Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0
Let's Connect:
🦾 Discord: discord.com/invite/t4eYQRUcXB
☕ Buy me a Coffee: ko-fi.com/promptengineering
|🔴 Patreon: patreon.com/PromptEngineering
💼Consulting: calendly.com/engineerprompt/consulting-call
📧 Business Contact: engineerprompt@gmail.com
Become Member: tinyurl.com/y5h28s6h
💻 Pre-configured localGPT VM: bit.ly/localGPT (use Code: PromptEngineering for 50% off).
Signup for Newsletter, localgpt:
https://tally.so/r/3y9bb0










