Uploaded October 2024 | Updated September 2026, 1 week ago
Djuna von Maydell, MIT
Understanding how training data influences model predictions ("data attribution") is an active area of machine learning research. In this tutorial, we will introduce a data attribution method (datamodels: gradientscience.org/datamodels-1/) and explore how it can be applied in the life sciences to identify meaningful subgroups in biomedical datasets, such as disease subtypes. We will begin with a simple example from image classification (CIFAR10), offering a step-by-step guide to demonstrate how the data attribution method works in practice. Since the approach involves training thousands of lightweight classifiers, we will focus on strategies for fast and efficient model training. Next, we will explore its applications in biomedical science, with a focus on single-cell and genetic datasets, highlighting the biological insights gained from applying this computational approach. The tutorial will conclude with an interactive, hands-on session using Google Colab, where participants can apply the techniques themselves and explore the approach further. This session is designed to be accessible to participants of all coding and machine learning experience levels—whether you're new to machine learning or curious about its intersection with biomedical applications.
Slides: drive.google.com/file/d/1qGahNYBUnThba07D2D9gZTviiU_kOedF/view?usp=sharing
Github repository of tutorial code: github.com/djunamay/datamodels_tutorial
Code with outputs: colab.research.google.com/drive/1u2jZzWs7SVT6kj-O8rMsUphHfvyeqnHh?usp=sharing
Code no outputs:
colab.research.google.com/drive/1lwl7-Xsc7lg9bTg97hEEqPt54x-J1qeU?usp=sharing
Djuna von Maydell, MIT
Understanding how training data influences model predictions ("data attribution") is an active area of machine learning research. In this tutorial, we will introduce a data attribution method (datamodels: gradientscience.org/datamodels-1/) and explore how it can be applied in the life sciences to identify meaningful subgroups in biomedical datasets, such as disease subtypes. We will begin with a simple example from image classification (CIFAR10), offering a step-by-step guide to demonstrate how the data attribution method works in practice. Since the approach involves training thousands of lightweight classifiers, we will focus on strategies for fast and efficient model training. Next, we will explore its applications in biomedical science, with a focus on single-cell and genetic datasets, highlighting the biological insights gained from applying this computational approach. The tutorial will conclude with an interactive, hands-on session using Google Colab, where participants can apply the techniques themselves and explore the approach further. This session is designed to be accessible to participants of all coding and machine learning experience levels—whether you're new to machine learning or curious about its intersection with biomedical applications.
Slides: drive.google.com/file/d/1qGahNYBUnThba07D2D9gZTviiU_kOedF/view?usp=sharing
Github repository of tutorial code: github.com/djunamay/datamodels_tutorial
Code with outputs: colab.research.google.com/drive/1u2jZzWs7SVT6kj-O8rMsUphHfvyeqnHh?usp=sharing
Code no outputs:
colab.research.google.com/drive/1lwl7-Xsc7lg9bTg97hEEqPt54x-J1qeU?usp=sharing





![Efficient representation, learning, and planning through abstraction: clustering cognitive spaces...
[full title] Efficient representation, learning, and planning through abstraction: clustering cognitive spaces into submaps
Ila Fiete, MIT
Abstract: Episodic memory involves fragmenting the continuous stream of experience into discrete episodes. Not coincidentally, the hippocampus, which plays a central role in both episodic memory and spatial navigation, represents large spatial environments in a fragmented way even when explored in a continuous trajectory. In non-spatial and non-memory contexts too, humans report sudden contextual re-anchoring or re-orientation when reading garden path sentences (“Time flies like an arrow, fruit flies like a banana.) or watching a movie with viewpoint changes. In this talk, I will describe a theory for the online and real-time generation of fragmented representations and contextual re-anchoring from continuous experience that resemble those obtained by principled but offline and computationally complex information-based algorithms. The resulting fragmentations closely match those observed from neural recordings in animals navigating through complex environments. I will discuss the utility of map fragmentation, as a form of state abstraction that enables representation fidelity, flexible and rapid learning through reuse of existing fragments, and many-fold improvements in the ability to plan and navigate through complex environments relative to more global representations. Efficient representation, learning, and planning through abstraction: clustering cognitive spaces...](https://i.ytimg.com/vi/gfgoLjhrh7k/mqdefault.jpg)




