AI Interpretability, Safety, and Meaning - Nora Belrose @MachineLearningStreetTalk
AI Interpretability, Safety, and Meaning - Nora Belrose  @MachineLearningStreetTalk
Uploaded November 2024 | Updated September 2026, 1 week ago
SPONSOR MESSAGES:
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments.
centml.ai/pricing

Nora Belrose, Head of Interpretability Research at EleutherAI, delivers a wide-ranging conversation that moves from the mathematical foundations of concept erasure in neural networks to fundamental questions about consciousness, AI safety, and Buddhist philosophy.

The technical core centers on LEACE (LEAst-squares Concept Erasure), a method Belrose developed for surgically removing targeted information from neural network representations. She explains how LEACE emerged from connecting two prior approaches (RLACE and spectral attribute removal) through a mathematical equivalence proof, and demonstrates its applications in both fairness-oriented debiasing and interpretability research. A key finding: language models remain functional even after erasing part-of-speech information from every layer, suggesting robust reliance on redundant cues.

Belrose then presents her ICML paper on simplicity biases in deep learning, showing that neural networks learn to exploit statistical moments in order -- first means, then covariances, then higher-order statistics. This has implications for understanding when and why concept erasure techniques may backfire against sufficiently deep models.

The second half pivots to AI safety, where Belrose delivers a detailed critique of "counting arguments" used to predict AI misalignment. She argues these arguments rely on the principle of indifference applied to poorly-defined outcome spaces, drawing an analogy to an identical argument structure that would absurdly predict all neural networks must overfit. She connects this to broader questions about goal attribution, agency, and whether instrumental convergence arguments hold up under scrutiny.

The conversation concludes with an exploration of 4E cognition, Evan Thompson's philosophy of mind, Belrose's departure from effective altruism, and her growing interest in Buddhist philosophy as a framework for thinking about meaning in a post-automation world.

---
REFERENCES:
Paper:
[00:00:00] Episode Shownotes
dropbox.com/scl/fi/38fhsv2zh8gnubtjaoq4a/NORA_FINAL.pdf?rlkey=0e5r8rd261821g1em4dgv0k70&st=t5c9ckfb&dl=0
[00:05:00] LEACE Paper
arxiv.org/abs/2306.03819
[00:06:40] RLACE Paper
arxiv.org/abs/2201.12091
[00:08:20] Spectral Attribute Removal
arxiv.org/abs/2012.14424
[00:15:00] Pythia Models
arxiv.org/abs/2304.01373
[00:20:30] LoRA
arxiv.org/abs/2106.09685
[02:00:00] Holden Karnofsky
forum.effectivealtruism.org/posts/T975ydo3mx4YnRv4J/ea-is-about-maximization-and-maximization-is-perilous
Company:
[00:01:37] CentML
centml.ai/pricing
[00:01:37] Tufa AI Labs
tufalabs.ai
[00:02:20] EleutherAI
eleuther.ai
Person:
[00:02:20] Nora Belrose
norabelrose.com
[01:03:00] Evan Thompson
evanthompson.me

---
LINKS:
Full Transcript: app.rescript.info/share/79d69cf24406cc36d8f7e8eee389e3ae
Download PDF transcript: app.rescript.info/api/public/sessions/61e64e1737593802/pdf

Nora Belrose:
norabelrose.com
scholar.google.com/citations?user=p_oBc64AAAAJ&hl=en
https://x.com/norabelrose
AI Interpretability, Safety, and Meaning - Nora BelroseDemocracy as a Model for AI GovernanceIs ChatGPT an N-gram model on steroids?This is what happens when you let AIs debateMicrosofts professor Chris Bishop on the Sparks of AGIWhat is “reasoning” in modern AI?AI Doesnt Need Intelligence. Thats the Whole Point. - Luciano FloridiNoam Chomsky - Science Vs EngineeringNEURAL NETWORKS ARE WEIRD! - Neel Nanda (DeepMind)Biologically-inspired AI and Mortal ComputationA CERN like effort for AI (Jakob Foerster)Strange Geometric Shapes Found Inside AIs — Tom McGrath
Machine Learning Street Talk |

AI Interpretability, Safety, and Meaning - Nora Belrose

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER