Why Frontier AI Labs Fight to Hide Chain of Thought — Ilia Shumailov & Alexander Panfilov @MachineLearningStreetTalk
Why Frontier AI Labs Fight to Hide Chain of Thought — Ilia Shumailov & Alexander Panfilov  @MachineLearningStreetTalk
Uploaded August 2026 | Updated September 2026, 2 weeks ago
Tim Scarfe speaks with Ilia Shumailov and Alexander Panfilov about their paper, Stealing Reasoning Traces from Proprietary LLM APIs.

The core bug sounds deceptively simple: providers return encrypted reasoning state so conversations can be resumed or forked. But those blobs can be replayed across users and sibling models. A smaller model can ask the provider to decrypt the trace, then repeat the hidden reasoning in plain text. The discussion covers leaked private data, a broadly reusable jailbreak, poisoned agent traces, chain-of-thought monitoring, responsible disclosure, and possible defenses.

Ilia Shumailov is an AI and security researcher, formerly at Google DeepMind, who completed his Cambridge PhD under Ross Anderson. Alexander Panfilov is a PhD researcher at the ELLIS Institute Tübingen and the Max Planck Institute for Intelligent Systems, working on AI safety, adversarial machine learning, and LLM red-teaming. They close by separating the demonstrated jailbreaking threat from ordinary benign distillation, and by arguing for controlled experiments over sweeping claims.

---
TIMESTAMPS:
00:00:00 Intro montage
00:01:33 Portable encrypted thought and decoded reasoning
00:24:55 How the attack works and what it means
00:39:04 Doom, defense, and scientific restraint

---
REFERENCES:
paper:
[00:00:00] Stealing Reasoning Traces from Proprietary LLM APIs
arxiv.org/abs/2608.09867
[00:09:22] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
arxiv.org/abs/2507.11473
[00:11:30] Reasoning Models Don’t Always Say What They Think
anthropic.com/research/reasoning-models-dont-say-think
[00:37:22] PostTrainBench: Can LLM Agents Automate LLM Post-Training?
arxiv.org/abs/2603.08640
[00:41:02] Large-scale online deanonymization with LLMs
arxiv.org/abs/2602.16800
other:
[00:09:28] OpenAI and Hugging Face partner to address security incident during model evaluation
openai.com/index/hugging-face-model-evaluation-security-incident
[00:10:22] Claude, GPT, and Gemini All Struggle to Evade Monitors
metr.org/notes/2025-08-22-claude-gpt-gemini-struggle-evade-monitors
tool:
[00:42:08] Isabelle proof assistant
isabelle.in.tum.de

---
RESCRIPT:
app.rescript.info/share/07fc38276e0823dc9b8986c32e202c7f
Why Frontier AI Labs Fight to Hide Chain of Thought — Ilia Shumailov & Alexander PanfilovAI needs to touch the real world (Eiso Kant)AI is grown, not designed - Connor LeahyThe Hidden Math Behind All Living SystemsThe Principles of Deep Learning Theory - Daniel A. Roberts Ph.Dwhat colour are the clouds? Max BartoloThe Universal Hierarchy of Life - Prof. Chris Kempes [SFI]Explosive AI Timeline Predictions [Gary Marcus, Daniel Kokotajlo, Dan Hendrycks]Are We Building Superintelligence Backwards? — Sara Saab & Enzo BlindowAIs which explore the worldChollets ARC Challenge + Current WinnersAGI in 5 Years? Ben Goertzel on Superintelligence
Machine Learning Street Talk |

Why Frontier AI Labs Fight to Hide Chain of Thought — Ilia Shumailov & Alexander Panfilov

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER