Uploaded September 2026 | Updated September 2026, 3 weeks ago
Recorded 03 September 2026. Dmitry Vaintrob of Principles of Intelligence, PIBBSS, presents "Statistical theory of learning sparse structure" at IPAM's Foundations of Interpretability Workshop.
Abstract: LLMs in practice have activation data that is relatively well explained by a sparse linear combination of a dictionary of feature vectors. I discuss some mean field theory-flavored work on solutions found by toy models trained on such data, and show how analogs of such structure appear in real LLM activations.
Learn more online at: https://www.ipam.ucla.edu/programs/workshops/foundations-of-interpretability/?tab=overview
Recorded 03 September 2026. Dmitry Vaintrob of Principles of Intelligence, PIBBSS, presents "Statistical theory of learning sparse structure" at IPAM's Foundations of Interpretability Workshop.
Abstract: LLMs in practice have activation data that is relatively well explained by a sparse linear combination of a dictionary of feature vectors. I discuss some mean field theory-flavored work on solutions found by toy models trained on such data, and show how analogs of such structure appear in real LLM activations.
Learn more online at: https://www.ipam.ucla.edu/programs/workshops/foundations-of-interpretability/?tab=overview










