Scaling AI Interpretability. #artificialintelligance #aiinterpretability #aitalk @BuzzRobot
Scaling AI Interpretability. #artificialintelligance #aiinterpretability #aitalk  @BuzzRobot
Uploaded January 2025 | Updated September 2026, 2 weeks ago
Can we scale #ai interpretability? Researchers use interpretability tools to predict, control, and understand #deeplearning models—but only in limited domains. Now, it’s time to automate and expand these methods for a broader understanding of general-purpose #ai.

@BuzzRobot guest Atticus Geiger, Stanford graduate and head of the Pr(Ai)²R Group, challenges the current approach using sparse autoencoders. Instead, he proposes a new method leveraging interventional data to better control and understand #deeplearning models.

Timestamps:
0:00 Introduction
0:55 Prediction: probes
2:45 Control: fine-tuning, prompt engineering, steering, representation fine-tuning (ReFT)
9:36 Understanding interpretability: Part 1. Computational explanation, causal abstraction
13:50 The modern AI and the ideal experimental subject
16:05 Understanding interpretability: Part 2. Recipe, interchange intervention analysis, distributed alignment search
22:10 Scaling interpretability: sparce autoencoders (SAEs), four goals of scaling interpretability

#ai #AIInterpretability #deeplearning #neuralnetworks #explainableai #aitransparency #aiselfImprovement #aimodel #machinelearning #reinforcementlearning #llm #llms #airesearch #aichallenges #aiprogress #tech #techtalk #techtalks #aitalks #aitalk #science #python #pythonprogramming #ai #programming

Social Links:
Newsletter: buzzrobot.substack.com
X: https://x.com/sopharicks
Slack: join.slack.com/t/buzzrobot/shared_invite/zt-2s067rv7n-guPIMGe62rbp9ncxdnOUfQ
Scaling AI Interpretability. #artificialintelligance #aiinterpretability #aitalkAI Model Self-Improvement: Progress and Challenges. #artificialintelligance #airesearch #aitalkWhat Does It Mean for #AI to Be #Aligned?AI will give humans the freedom to modify themselves #artificialgeneralintelligenceThe future of AI investmentsHow Google’s Med #Gemini Models Are Trained to Trace Their ReasoningAI model meltdownsTraining Robots in 3D Environment! #AI #Robotics #3DWorldsHow AI Can Accelerate ScienceCan AI solve LSAT problems better than humans? #ai #aiscalabilityArtificial General Intelligence by 2027? #agi #artificialintelligenceResolving geopolitical tensions around AI
BuzzRobot |

Scaling AI Interpretability. #artificialintelligance #aiinterpretability #aitalk

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER