Uploaded January 2025 | Updated September 2026, 2 weeks ago
Can we scale #ai interpretability? Researchers use interpretability tools to predict, control, and understand #deeplearning models—but only in limited domains. Now, it’s time to automate and expand these methods for a broader understanding of general-purpose #ai.
@BuzzRobot guest Atticus Geiger, Stanford graduate and head of the Pr(Ai)²R Group, challenges the current approach using sparse autoencoders. Instead, he proposes a new method leveraging interventional data to better control and understand #deeplearning models.
Timestamps:
0:00 Introduction
0:55 Prediction: probes
2:45 Control: fine-tuning, prompt engineering, steering, representation fine-tuning (ReFT)
9:36 Understanding interpretability: Part 1. Computational explanation, causal abstraction
13:50 The modern AI and the ideal experimental subject
16:05 Understanding interpretability: Part 2. Recipe, interchange intervention analysis, distributed alignment search
22:10 Scaling interpretability: sparce autoencoders (SAEs), four goals of scaling interpretability
#ai #AIInterpretability #deeplearning #neuralnetworks #explainableai #aitransparency #aiselfImprovement #aimodel #machinelearning #reinforcementlearning #llm #llms #airesearch #aichallenges #aiprogress #tech #techtalk #techtalks #aitalks #aitalk #science #python #pythonprogramming #ai #programming
Social Links:
Newsletter: buzzrobot.substack.com
X: https://x.com/sopharicks
Slack: join.slack.com/t/buzzrobot/shared_invite/zt-2s067rv7n-guPIMGe62rbp9ncxdnOUfQ
Can we scale #ai interpretability? Researchers use interpretability tools to predict, control, and understand #deeplearning models—but only in limited domains. Now, it’s time to automate and expand these methods for a broader understanding of general-purpose #ai.
@BuzzRobot guest Atticus Geiger, Stanford graduate and head of the Pr(Ai)²R Group, challenges the current approach using sparse autoencoders. Instead, he proposes a new method leveraging interventional data to better control and understand #deeplearning models.
Timestamps:
0:00 Introduction
0:55 Prediction: probes
2:45 Control: fine-tuning, prompt engineering, steering, representation fine-tuning (ReFT)
9:36 Understanding interpretability: Part 1. Computational explanation, causal abstraction
13:50 The modern AI and the ideal experimental subject
16:05 Understanding interpretability: Part 2. Recipe, interchange intervention analysis, distributed alignment search
22:10 Scaling interpretability: sparce autoencoders (SAEs), four goals of scaling interpretability
#ai #AIInterpretability #deeplearning #neuralnetworks #explainableai #aitransparency #aiselfImprovement #aimodel #machinelearning #reinforcementlearning #llm #llms #airesearch #aichallenges #aiprogress #tech #techtalk #techtalks #aitalks #aitalk #science #python #pythonprogramming #ai #programming
Social Links:
Newsletter: buzzrobot.substack.com
X: https://x.com/sopharicks
Slack: join.slack.com/t/buzzrobot/shared_invite/zt-2s067rv7n-guPIMGe62rbp9ncxdnOUfQ










