Uploaded August 2024 | Updated September 2026, 1 week ago
DeepMind machine learning scientist / MIT scholar Dr. Timothy Nguyen discusses his recent paper on understanding transformers through n-gram statistics. Nguyen explains his approach to analyzing transformer behavior using a kind of "template matching" (N-grams), providing insights into how these models process and predict language.
MLST is sponsored by Brave:
The Brave Search API covers over 20 billion webpages, built from scratch without Big Tech biases or the recent extortionate price hikes on search API access. Perfect for AI model training and retrieval augmentated generation. Try it now - get 2,000 free queries monthly at brave.com/api.
Key points covered include:
A method for describing transformer predictions using n-gram statistics without relying on internal mechanisms.
The discovery of a technique to detect overfitting in large language models without using holdout sets.
Observations on curriculum learning, showing how transformers progress from simpler to more complex rules during training.
Discussion of distance measures used in the analysis, particularly the variational distance.
Exploration of model sizes, training dynamics, and their impact on the results.
We also touch on philosophical aspects of describing versus explaining AI behavior, and the challenges in understanding the abstractions formed by neural networks. Nguyen concludes by discussing potential future research directions, including attempts to convert descriptions of transformer behavior into explanations of internal mechanisms.
Timothy Nguyen's earned his B.S. and Ph.D. in mathematics from Caltech and MIT, respectively. He held positions as Research Assistant Professor at the Simons Center for Geometry and Physics (2011-2014) and Visiting Assistant Professor at Michigan State University (2014-2017). During this time, his research expanded into high-energy physics, focusing on mathematical problems in quantum field theory. His work notably provided a simplified and corrected formulation of perturbative path integrals.
Since 2017, Nguyen has been working in industry, applying his expertise to machine learning. He is currently at DeepMind, where he contributes to both fundamental research and practical applications of deep learning to solve real-world problems.
Refs:
The Cartesian Cafe
youtube.com/@TimothyNguyen
Understanding Transformers via N-Gram Statistics
researchgate.net/publication/382204056_Understanding_Transformers_via_N-Gram_Statistics
TOC
00:00:00 Timothy Nguyen's background
00:02:50 Paper overview: transformers and n-gram statistics
00:04:55 Template matching and hash table approach
00:08:55 Comparing templates to transformer predictions
00:12:01 Describing vs explaining transformer behavior
00:15:36 Detecting overfitting without holdout sets
00:22:47 Curriculum learning in training
00:26:32 Distance measures in analysis
00:28:58 Model sizes and training dynamics
00:30:39 Future research directions
00:32:06 Conclusion and future topics
DeepMind machine learning scientist / MIT scholar Dr. Timothy Nguyen discusses his recent paper on understanding transformers through n-gram statistics. Nguyen explains his approach to analyzing transformer behavior using a kind of "template matching" (N-grams), providing insights into how these models process and predict language.
MLST is sponsored by Brave:
The Brave Search API covers over 20 billion webpages, built from scratch without Big Tech biases or the recent extortionate price hikes on search API access. Perfect for AI model training and retrieval augmentated generation. Try it now - get 2,000 free queries monthly at brave.com/api.
Key points covered include:
A method for describing transformer predictions using n-gram statistics without relying on internal mechanisms.
The discovery of a technique to detect overfitting in large language models without using holdout sets.
Observations on curriculum learning, showing how transformers progress from simpler to more complex rules during training.
Discussion of distance measures used in the analysis, particularly the variational distance.
Exploration of model sizes, training dynamics, and their impact on the results.
We also touch on philosophical aspects of describing versus explaining AI behavior, and the challenges in understanding the abstractions formed by neural networks. Nguyen concludes by discussing potential future research directions, including attempts to convert descriptions of transformer behavior into explanations of internal mechanisms.
Timothy Nguyen's earned his B.S. and Ph.D. in mathematics from Caltech and MIT, respectively. He held positions as Research Assistant Professor at the Simons Center for Geometry and Physics (2011-2014) and Visiting Assistant Professor at Michigan State University (2014-2017). During this time, his research expanded into high-energy physics, focusing on mathematical problems in quantum field theory. His work notably provided a simplified and corrected formulation of perturbative path integrals.
Since 2017, Nguyen has been working in industry, applying his expertise to machine learning. He is currently at DeepMind, where he contributes to both fundamental research and practical applications of deep learning to solve real-world problems.
Refs:
The Cartesian Cafe
youtube.com/@TimothyNguyen
Understanding Transformers via N-Gram Statistics
researchgate.net/publication/382204056_Understanding_Transformers_via_N-Gram_Statistics
TOC
00:00:00 Timothy Nguyen's background
00:02:50 Paper overview: transformers and n-gram statistics
00:04:55 Template matching and hash table approach
00:08:55 Comparing templates to transformer predictions
00:12:01 Describing vs explaining transformer behavior
00:15:36 Detecting overfitting without holdout sets
00:22:47 Curriculum learning in training
00:26:32 Distance measures in analysis
00:28:58 Model sizes and training dynamics
00:30:39 Future research directions
00:32:06 Conclusion and future topics
![This is what happens when you let AIs debate
Akbir Khan, AI researcher and ICML 2024 Best Paper winner, joins Tim Scarfe to discuss his groundbreaking work on using debate between language models to improve AI truthfulness and oversight. Khan explains how pitting two LLMs against each other in structured arguments helps non-expert judges arrive at more accurate answers than simply querying a single model — a result with profound implications for supervising AI systems that may eventually surpass human capabilities.
The conversation moves through the mechanics of scalable oversight and the sandwiching protocol that operationalises the problem of checking entities smarter than their supervisors. Khan describes how debate naturally surfaces the cruxes of disagreements, making complex expert judgments more accessible to laypeople. The discussion broadens into the relationship between intelligence and agency, the risks of deceptive alignment and reward tampering, and whether the current trajectory of AI development constitutes the early stages of a Cambrian explosion in artificial minds.
Khan and Scarfe also explore open-ended AI systems, Kenneth Stanleys arguments against objective-driven optimization, and the philosophical terrain mapped by thinkers like Aaron Sloman and Francois Chollet on the space of possible minds and measuring intelligence. The episode closes with a frank exchange on cultural evolution, memetics, and whether the intelligence that matters most for AI safety is the kind we can measure.
REFERENCES:
person:
[00:00:00] Akbir Khan
https://akbir.dev/
other:
[00:00:00] MLST Episode Shownotes PDF
https://www.dropbox.com/scl/fi/sjekivbg3ok6qugsv2p1u/AkbirKhan.pdf?rlkey=ewiyvq0aq7mjvql4u7os0jos2&st=vblhp7af&dl=0
[00:06:05] OpenAI Superalignment Team
https://openai.com/index/introducing-superalignment/
[00:19:10] DeepMind Responsible AI
https://deepmind.google/about/responsibility-safety/
paper:
[00:00:40] Akbir Khan et al. - Debating with More Persuasive LLMs
https://arxiv.org/html/2402.06782v3
[00:08:10] Sam Bowman - Scalable Oversight in AI Systems
https://arxiv.org/abs/2211.03540
[00:10:35] Sam Bowman - Artificial Sandwiching Protocol
https://www.alignmentforum.org/posts/nekLYqbCEBDEfbLzF/artificial-sandwiching-when-can-we-test-scalable-alignment
[00:14:35] Janus - Simulators
https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulators
[00:21:30] Eliezer Yudkowsky - AI FOOM Debate
https://intelligence.org/files/AIFoomDebate.pdf
[00:21:45] Sammy Martin - Discontinuous AI Progress
https://www.alignmentforum.org/posts/5WECpYABCT62TJrhY/will-ai-undergo-discontinuous-progress
[00:24:35] Nora Belrose - Counting Arguments vs AI Doom
https://www.lesswrong.com/posts/YsFZF3K9tuzbfrLxo/counting-arguments-provide-no-evidence-for-ai-doom
[00:25:35] Evan Hubinger - Deceptive Alignment
https://www.lesswrong.com/posts/zthDPAjh9w6Ytbeks/deceptive-alignment
[00:26:50] Anthropic - Reward Tampering
https://www.anthropic.com/research/reward-tampering
[00:34:58] Ryan Greenblatt et al. - AI Control
https://arxiv.org/pdf/2312.06942
[00:37:20] Aaron Sloman - The Space of Possible Minds
https://www.cs.bham.ac.uk/research/projects/cogaff/sloman-space-of-minds-84.pdf
[00:38:45] Francois Chollet - On the Measure of Intelligence
https://arxiv.org/abs/1911.01547
[00:42:45] Jonathan Cook et al. - Artificial Generational Intelligence
https://arxiv.org/abs/2406.00392
video:
[00:03:28] Yann LeCun on Machine Learning Debates
https://www.youtube.com/watch?v=OKkEdTchsiE
book:
[00:16:35] Thomas Suddendorf - The Gap
https://www.amazon.in/GAP-Science-Separates-Other-Animals/dp/0465030149
[00:32:35] Kenneth Stanley - Why Greatness Cannot Be Planned
https://www.amazon.co.uk/Why-Greatness-Cannot-Planned-Objective/dp/3319155237
[00:42:30] Richard Dawkins - The Selfish Gene
https://www.amazon.co.uk/Selfish-Gene-Richard-Dawkins/dp/0192860925
LINKS:
Full Transcript: https://app.rescript.info/share/c75b45c237ad52154d276851af8f812b
Download PDF transcript: https://app.rescript.info/api/public/sessions/571abed49dab3e7f/pdf
Akbir Khan:
https://x.com/akbirkhan
https://akbir.dev/ This is what happens when you let AIs debate](https://i.ytimg.com/vi/WlWAhjPfROU/mqdefault.jpg)

![What is “reasoning” in modern AI?
Professor Swarat Chaudhuri from the University of Texas at Austin and visiting researcher at Google DeepMind discusses breakthroughs in AI reasoning, theorem proving, and mathematical discovery. Chaudhuri explains his groundbreaking work on COPRA (a GPT-based prover agent), shares insights on neurosymbolic approaches to AI.
Professor Swarat Chaudhuri:
https://www.cs.utexas.edu/~swarat/
SPONSOR MESSAGES:
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments.
https://centml.ai/pricing/
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on ARC and AGI, they just acquired MindsAI - the current winners of the ARC challenge. Are you interested in working on ARC, or getting involved in their events? Goto https://tufalabs.ai/
TOC:
[00:00:00] 0. Introduction / CentML ad, Tufa ad
1. AI Reasoning: From Language Models to Neurosymbolic Approaches
[00:02:27] 1.1 Defining Reasoning in AI
[00:09:51] 1.2 Limitations of Current Language Models
[00:17:22] 1.3 Neuro-symbolic Approaches and Program Synthesis
[00:24:59] 1.4 COPRA and In-Context Learning for Theorem Proving
[00:34:39] 1.5 Symbolic Regression and LLM-Guided Abstraction
2. AI in Mathematics: Theorem Proving and Concept Discovery
[00:43:37] 2.1 AI-Assisted Theorem Proving and Proof Verification
[01:01:37] 2.2 Symbolic Regression and Concept Discovery in Mathematics
[01:11:57] 2.3 Scaling and Modularizing Mathematical Proofs
[01:21:53] 2.4 COPRA: In-Context Learning for Formal Theorem-Proving
[01:28:22] 2.5 AI-driven theorem proving and mathematical discovery
3. Formal Methods and Challenges in AI Mathematics
[01:30:42] 3.1 Formal proofs, empirical predicates, and uncertainty in AI mathematics
[01:34:01] 3.2 Characteristics of good theoretical computer science research
[01:39:16] 3.3 LLMs in theorem generation and proving
[01:42:21] 3.4 Addressing contamination and concept learning in AI systems
REFS:
00:04:58 The Chinese Room Argument, https://plato.stanford.edu/entries/chinese-room/
00:11:42 Software 2.0, https://medium.com/@karpathy/software-2-0-a64152b37c35
00:11:57 Solving Olympiad Geometry Without Human Demonstrations, https://www.nature.com/articles/s41586-023-06747-5
00:13:26 Lean, https://lean-lang.org/
00:15:43 A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play, https://www.science.org/doi/10.1126/science.aar6404
00:19:24 DreamCoder: Bootstrapping Inductive Program Synthesis with Wake-Sleep Library Learning (Ellis et al., PLDI 2021), https://arxiv.org/abs/2006.08381
00:24:37 The Lambda Calculus, https://plato.stanford.edu/entries/lambda-calculus/
00:26:43 Neural Sketch Learning for Conditional Program Generation, https://arxiv.org/pdf/1703.05698
00:28:08 Learning Differentiable Programs With Admissible Neural Heuristics, https://arxiv.org/abs/2007.12101
00:31:03 Symbolic Regression With a Learned Concept Library (Grayeli et al., NeurIPS 2024), https://arxiv.org/abs/2409.09359
00:41:21 Turing Machines, https://plato.stanford.edu/entries/turing-machine/#HaltProb
00:41:30 Formal Verification of Parallel Programs, https://dl.acm.org/doi/10.1145/360248.360251
01:00:08 The Feynman Lectures, https://www.feynmanlectures.caltech.edu/
01:00:37 Training Compute-Optimal Large Language Models, https://arxiv.org/abs/2203.15556
01:12:26 Fermats Last Theorem, https://en.wikipedia.org/wiki/Fermat%27s_Last_Theorem
01:18:19 Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, https://arxiv.org/abs/2201.11903
01:18:42 Draft, Sketch, and Prove: Guiding Formal Theorem Provers With Informal Proofs, https://arxiv.org/abs/2210.12283
01:19:49 Learning Formal Mathematics From Intrinsic Motivation, https://arxiv.org/pdf/2407.00695
01:20:19 An In-Context Learning Agent for Formal Theorem-Proving (Thakur et al., CoLM 2024), https://arxiv.org/pdf/2310.04353
01:23:58 Learning to Prove Theorems via Interacting With Proof Assistants, https://arxiv.org/abs/1905.09381
01:35:50 Algorithmic Game Theory, https://www.amazon.ca/Algorithmic-Game-Theory-Noam-Nisan/dp/0521872820
01:39:58 An In-Context Learning Agent for Formal Theorem-Proving (Thakur et al., CoLM 2024), https://arxiv.org/pdf/2310.04353
01:42:24 Programmatically Interpretable Reinforcement Learning (Verma et al., ICML 2018), https://arxiv.org/abs/1804.02477 What is “reasoning” in modern AI?](https://i.ytimg.com/vi/XFMk0snybAc/mqdefault.jpg)


![NEURAL NETWORKS ARE WEIRD! - Neel Nanda (DeepMind)
SPONSOR MESSAGES:
***
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments.
https://centml.ai/pricing/
Neel Nanda leads the mechanistic interpretability team at Google DeepMind. At 26, hes become one of the most prominent researchers working on the question of whats actually going on inside neural networks systems that can win IMO medals and write complex software, but which nobody actually designed or understands.
This nearly four-hour conversation is a deep technical dive into the field. Nanda explains why machine learning is fundamentally weird: we produce artifacts that do impressive things, but unlike conventional software, no one wrote the code or planned the architecture. His teams goal is reverse-engineering these systems by finding the internal structures and algorithms that emerge during training.
The discussion covers the mechanics of sparse autoencoders at length how they decompose model activations into interpretable feature vectors, the mathematical foundations (ReLU vs TopK activation functions), scaling laws for feature learning, and the engineering challenges of running them at the scale of frontier models. Nanda walks through the Golden Gate Claude experiment (amplifying a single feature to make Claude obsessed with the Golden Gate Bridge), induction heads (the circuits responsible for in-context learning), and activation patching as a causal intervention technique.
On AI safety, Nanda is pragmatic. He argues that mechanistic interpretability gives us genuine empirical evidence about questions that are otherwise stuck in philosophical debate do models have goals? Do they deceive? He also discusses the limitations: sparse autoencoders havent yet demonstrated capabilities beyond what fine-tuning already achieves, and at sufficient model complexity, models could potentially facade interpretability measurements. The conversation covers his path from pure maths at Cambridge through Anthropic to DeepMind, and why he thinks hands-on coding matters more than reading papers for new researchers entering the field.
REFERENCES:
person:
[00:00:00] Neel Nanda - Personal Website
https://www.neelnanda.io/
tool:
[00:35:00] TransformerLens
https://github.com/TransformerLensOrg/TransformerLens
paper:
[01:00:31] A Mathematical Framework for Transformer Circuits
https://transformer-circuits.pub/2021/framework/index.html
[01:01:40] In-context Learning and Induction Heads
https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html
[01:21:06] Scaling Monosemanticity
https://transformer-circuits.pub/2024/scaling-monosemanticity/
[01:33:27] Refusal in Language Models Is Mediated by a Single Direction
https://arxiv.org/abs/2406.11717
LINKS:
Full Transcript: https://app.rescript.info/share/acb415fa59ae2d2909d60d761c8f4ff4
Download PDF transcript: https://app.rescript.info/api/public/sessions/c3a4bf1e32a46ce7/pdf
NEEL NANDA:
https://www.neelnanda.io/
https://scholar.google.com/citations?user=GLnX3MkAAAAJ&hl=en
https://x.com/NeelNanda5 NEURAL NETWORKS ARE WEIRD! - Neel Nanda (DeepMind)](https://i.ytimg.com/vi/YpFaPKOeNME/mqdefault.jpg)
![Biologically-inspired AI and Mortal Computation
MLST is sponsored by Tufa Labs:
Are you interested in working on ARC and cutting-edge AI research with the MindsAI team (current ARC winners)?
Focus: ARC, LLMs, test-time-compute, active inference, system2 reasoning, and more.
Future plans: Expanding to complex environments like Warcraft 2 and Starcraft 2.
Interested? Apply for an ML research position: benjamin@tufa.ai
Professor Alexander Ororbia from the Rochester Institute of Technology takes Tim Scarfe through the case for bio-inspired AI. The central idea is mortal computation: you cannot divorce the software from the hardware that runs it. The brain manages remarkable things on a few watts because its computations are entangled with its physical substrate. GPT-class models, running on von Neumann architectures designed for immortal computation where software and hardware are deliberately decoupled pay a staggering energy penalty for that separation.
Ororbia explains the building blocks: Markov blankets as the formalism for system boundaries, Karl Fristons free energy principle as the optimization target, and the MILLS framework (Mortal Inference, Learning, and Selection) operating across multiple timescales. He then surveys the landscape of alternatives to backpropagation predictive coding, Hebbian learning, contrastive methods, and Geoff Hintons forward-forward algorithm showing how each maps to observations from neuroscience.
The conversation gets practical with Ororbias ngc-learn library for implementing these algorithms, the stability-plasticity dilemma in continual learning, and the current state of neuromorphic hardware from Intel Loihi to IBM TrueNorth. He closes with his neural generative coding work, which showed that predictive coding networks can synthesize data they were never trained on outperforming VAEs and GANs and his vision for bio-inspired AI systems that coexist with humanity rather than replacing it.
TIMESTAMPS:
00:00:00 Introduction to Bio-Inspired AI and Mortal Computation
00:04:50 Principles of Mortal Computation
00:17:41 Markov Blankets and Free Energy Principle
00:24:38 MILLS Framework and Biological Systems
00:31:00 Challenging Backpropagation: Alternative Approaches
00:31:49 Predictive Coding and Free Energy Principle
00:41:52 Biologically Plausible Credit Assignment Methods
00:50:11 Taxonomy of Bio-inspired Learning Algorithms
00:59:30 Forward-Only Learning and ngc-learn Implementation
01:03:25 Stability-Plasticity Dilemma and Continual Learning
01:09:00 Neuromorphic Hardware and Challenges
01:12:58 Neural Generative Coding and Future Directions
REFERENCES:
website:
[00:04:43] The Levin Lab
https://drmichaellevin.org/
[00:18:20] Good Regulator Theorem
https://en.wikipedia.org/wiki/Good_regulator
[00:41:52] Hebbian Theory
https://en.wikipedia.org/wiki/Hebbian_theory
[00:45:00] Hopfield Network
https://en.wikipedia.org/wiki/Hopfield_network
[01:09:00] Intel Loihi 2
https://www.intel.com/content/www/us/en/research/neuromorphic-computing-loihi-2-technology-brief.html
paper:
[00:04:50] Mortal Computation: A Foundation for Biomimetic Intelligence
https://arxiv.org/abs/2311.09589
[00:06:53] The Forward-Forward Algorithm
https://arxiv.org/abs/2212.13345
[00:07:20] Theres Plenty of Room Right Here
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10046700/
[00:17:41] The Free-Energy Principle: A Rough Guide to the Brain
https://www.fil.ion.ucl.ac.uk/~karl/The%20free-energy%20principle%20-%20a%20rough%20guide%20to%20the%20brain.pdf
[00:31:49] Predictive Coding in the Visual Cortex
https://www.nature.com/articles/nn0199_79
[00:41:52] Brain-Inspired Machine Intelligence: Neurobiologically-Plausible Credit Assignment
https://arxiv.org/abs/2312.09257
[00:45:50] A Tutorial on Energy-Based Learning
https://yann.lecun.com/exdb/publis/pdf/lecun-06.pdf
[00:46:40] A Learning Algorithm for Boltzmann Machines
https://www.cs.toronto.edu/~hinton/absps/cogscibm.pdf
[00:50:11] A Review of Neuroscience-Inspired Machine Learning
https://arxiv.org/abs/2403.18929
[00:53:20] NEAT: NeuroEvolution of Augmenting Topologies
https://nn.cs.utexas.edu/downloads/papers/stanley.ec02.pdf
[00:56:40] A Path Towards Autonomous Machine Intelligence
https://openreview.net/pdf?id=BZ5a1r-kVsf
[00:59:30] Test-Time Model Adaptation with Only Forward Passes
https://arxiv.org/abs/2404.01650
[01:03:25] Spiking Neural Predictive Coding for Continual Learning
https://www.sciencedirect.com/science/article/pii/S0925231223004150
[01:10:00] IBM TrueNorth
https://research.ibm.com/publications/truenorth-design-and-tool-flow-of-a-65-mw-1-million-neuron-programmable-neurosynaptic-chip
book:
[00:24:38] Active Inference: The Free Energy Principle in Mind, Brain, and Behavior
https://direct.mit.edu/books/oa-monograph/5299/Active-InferenceThe-Free-Energy-Principle-in-Mind
LINKS:
Full Transcript: https://app.rescript.info/share/cfa14d0f39f86d035d5caf6173d6207f
Download PDF transcript: https://app.rescript.info/api/public/sessions/94d5c32d7c507e6a/pdf Biologically-inspired AI and Mortal Computation](https://i.ytimg.com/vi/ZTE-JVd_QkA/mqdefault.jpg)

![Strange Geometric Shapes Found Inside AIs — Tom McGrath
Tom McGrath is co-founder and Chief Scientist at Goodfire, and a former Google DeepMind researcher. He joins Tim Scarfe to ask what neural networks actually learn, whether their internal representations converge on structures in the world, and whether interpretability can extract new scientific knowledge rather than merely explain model outputs.
Beginning with AlphaZero and learned modularity, the conversation moves into neural geometry: concept manifolds, reusable computation inside Llama, and why activation steering can fail when it pushes a model off-manifold. McGrath then makes the case for intentional design, using interpretability as part of the training loop. They examine controlled generalisation, features as rewards, predictive data debugging, and the uncomfortable fact that a model may recognise a hallucination or reward hack and still produce it.
The discussion closes on grader awareness, oversight and collusion between adaptive agents, then returns to sparse autoencoders. SAEs are useful, McGrath argues, but they may fracture the higher-dimensional structures networks actually use. This episode was made with support from Goodfire.
TIMESTAMPS:
00:00:00 Introduction: Can interpretability speed-run science?
00:02:03 The invisible grader
00:06:51 What AlphaZero learned from the world
00:12:24 Interpretability as a control loop
00:21:54 The forbidden method and safer interventions
00:37:36 Why models catch hallucinations too late
00:46:19 Debug the dataset before training
00:50:44 Why neural networks become modular
00:55:57 Finding the geometry inside a network
01:02:55 Why steering falls off the manifold
01:12:10 A reusable calculator inside Llama
01:17:19 From abstractions to goals
01:25:28 Reward hacking, oversight and collusion
01:37:23 Are sparse autoencoders dead?
REFERENCES:
paper:
[00:05:45] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
https://arxiv.org/abs/2502.17424v7
[00:11:05] Acquisition of Chess Knowledge in AlphaZero
https://arxiv.org/abs/2111.09259
[00:25:30] Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
https://arxiv.org/abs/2507.16795
[00:29:30] Persona Vectors: Monitoring and Controlling Character Traits in Language Models
https://arxiv.org/abs/2507.21509
[00:41:14] Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability
https://arxiv.org/abs/2602.10067
[00:47:03] Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal
https://arxiv.org/abs/2606.12360
[01:00:26] Do Sparse Autoencoders Capture Concept Manifolds?
https://arxiv.org/abs/2604.28119
[01:03:04] Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
https://arxiv.org/abs/2605.05115
[01:14:20] Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
https://arxiv.org/abs/2605.01148
[01:29:35] Measuring Reward-Seeking via Contrastive Belief Updates
https://arxiv.org/abs/2607.18966v1
other:
[00:15:44] Intentional Design
https://www.goodfire.com/blog/intentional-design
[00:56:12] The World Inside Neural Networks
https://www.goodfire.com/research/the-world-inside-neural-networks
[01:37:28] A Pragmatic Vision for Interpretability
https://www.alignmentforum.org/posts/StENzDcD3kpfGJssR/a-pragmatic-vision-for-interpretability
RESCRIPT:
https://app.rescript.info/share/846cfee4131b664fd09209cc3b98018e Strange Geometric Shapes Found Inside AIs — Tom McGrath](https://i.ytimg.com/vi/_egu7OFem-k/mqdefault.jpg)
![AI Agents can write 10,000 lines of hacking code in seconds [Dr. Ilia Shumailov]
Dr. Ilia Shumailov - Former DeepMind AI Security Researcher, now building security tools for AI agents
Ever wondered what happens when AI agents start talking to each other—or worse, when they start breaking things? Ilia Shumailov spent years at DeepMind thinking about exactly these problems, and hes here to explain why securing AI is way harder than you think.
**SPONSOR MESSAGES**
—
Check out notebooklm for your research project, its really powerful
https://notebooklm.google.com/
—
Take the Prolific human data survey - https://www.prolific.com/humandatasurvey?utm_source=mlst and be the first to see the results and benchmark their practices against the wider community!
—
cyber•Fund https://cyber.fund/?utm_source=mlst is a founder-led investment firm accelerating the cybernetic economy
Oct SF conference - https://dagihouse.com/?utm_source=mlst - Joscha Bach keynoting(!) + OAI, Anthropic, NVDA,++
Hiring a SF VC Principal: https://talent.cyber.fund/companies/cyber-fund-2/jobs/57674170-ai-investment-principal#content?utm_source=mlst
Submit investment deck: https://cyber.fund/contact?utm_source=mlst
—
Were racing toward a world where AI agents will handle our emails, manage our finances, and interact with sensitive data 24/7. But there is a problem. These agents are nothing like human employees. They never sleep, they can touch every endpoint in your system simultaneously, and they can generate sophisticated hacking tools in seconds. Traditional security measures designed for humans simply wont work.
Dr. Ilia Shumailov
https://x.com/iliaishacked
https://iliaishacked.github.io/
https://sequrity.ai/
TRANSCRIPT:
https://app.rescript.info/public/share/dVGsk8dz9_V0J7xMlwguByBq1HXRD6i4uC5z5r7EVGM
More from Ilia on our Patreon:
https://www.patreon.com/posts/116142401/ (interview from last year)
https://www.patreon.com/posts/ilia-shumailov-140359158 (extended version of this interview)
TOC:
00:00:00 - Introduction & Trusted Third Parties via ML
00:03:45 - Background & Career Journey
00:06:42 - Safety vs Security Distinction
00:09:45 - Prompt Injection & Model Capability
00:13:00 - Agents as Worst-Case Adversaries
00:15:45 - Personal AI & CAML System Defense
00:19:30 - Agents vs Humans: Threat Modeling
00:22:30 - Calculator Analogy & Agent Behavior
00:25:00 - IMO Math Solutions & Agent Thinking
00:28:15 - Diffusion of Responsibility & Insider Threats
00:31:00 - Open Source Security Concerns
00:34:45 - Supply Chain Attacks & Trust Issues
00:39:45 - Architectural Backdoors
00:44:00 - Academic Incentives & Defense Work
00:48:30 - Semantic Censorship & Halting Problem
00:52:00 - Model Collapse: Theory & Criticism
00:59:30 - Career Advice & Ross Anderson Tribute
REFS:
Lessons from Defending Gemini Against Indirect Prompt Injections
https://arxiv.org/abs/2505.14534
Defeating Prompt Injections by Design. Google, Google DeepMind, and ETH Zurich. (CAML)
Debenedetti, E., Shumailov, I., Fan, T., Hayes, J., Carlini, N., Fabian, D., Kern, C., Shi, C., Terzis, A., & Tramèr, F.
https://arxiv.org/pdf/2503.18813
Agentic Misalignment: How LLMs could be insider threats
https://www.anthropic.com/research/agentic-misalignment
STOP ANTHROPOMORPHIZING INTERMEDIATE TOKENS AS REASONING/THINKING TRACES!
Subbarao Kambhampati et al
https://arxiv.org/pdf/2504.09762
Meiklejohn, S., Blauzvern, H., Maruseac, M., Schrock, S., Simon, L., & Shumailov, I. (2025).
Machine learning models have a supply chain problem.
https://arxiv.org/abs/2505.22778
Gao, Y., Shumailov, I., & Fawaz, K. (2025).
Supply-chain attacks in machine learning frameworks.
In Proceedings of the 8th MLSys Conference.
https://openreview.net/pdf?id=EH5PZW6aCr
Apache Log4j Vulnerability Guidance
https://www.cisa.gov/news-events/news/apache-log4j-vulnerability-guidance
Bober-Irizar, M., Shumailov, I., Zhao, Y., Mullins, R., & Papernot, N. (2023).
Architectural backdoors in neural networks.
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 21163–21173).
Bober-Irizar, M., Shumailov, I., Zhao, Y., Mullins, R., & Papernot, N. (2022).
Architectural backdoors in neural networks. arXiv preprint arXiv:2206.07840.
https://arxiv.org/pdf/2206.07840
Langford, H., Shumailov, I., Zhao, Y., Mullins, R., & Papernot, N. (2024).
Architectural neural backdoors from first principles.
arXiv preprint arXiv:2402.06957.
Küchler, N., Petrov, I., Grobler, C., & Shumailov, I. (2025).
Architectural backdoors for within-batch data stealing and model inference manipulation.
arXiv preprint arXiv:2505.18323.
Position: Fundamental Limitations of LLM Censorship Necessitate New Approaches
David Glukhov, Ilia Shumailov, Yarin Gal, Nicolas Papernot, Vardan Papyan
https://proceedings.mlr.press/v235/glukhov24a.html
AlphaEvolve MLST interview [Matej Balog, Alexander Novikov]
https://www.youtube.com/watch?v=vC9nAosXrJw AI Agents can write 10,000 lines of hacking code in seconds [Dr. Ilia Shumailov]](https://i.ytimg.com/vi/aoX_pGQMbEM/mqdefault.jpg)
