Uploaded March 2026 | Updated September 2026, 2 weeks ago
Samory Kpotufe (Columbia University)
https://simons.berkeley.edu/talks/samory-kpotufe-columbia-university-2026-02-26
Learning from Heterogeneous Sources
Multitask Learning refers to the problem of aggregating many datasets from separate source distributions to improve performance on a target prediction task. Our aim is to understand (1) sufficient and necessary conditions for speedup in convergence rates over vanilla prediction with just the target data, (2) how such speedup depends on the number of datasets and samples per dataset, and (3) whether such speedup is achievable adaptively, i.e., by procedures with no prior distributional information.
The picture turns out to be mixed as the problem displays sharp gaps between oracle rates and adaptive rates, i.e., there exist situations where no procedure can do better than using the target data alone even though a large subset of datasets are informative about the target task. On the other hand, a bit of information on the relation between sources and target distributions can allow for near optimal adaptive rates. These results lead to many interesting new questions which I'll attempt to properly convey.
The talk is based on various works with collaborators such as S. Hanneke, A. Gretton, M. Z. Li, D. Meunier.
Samory Kpotufe (Columbia University)
https://simons.berkeley.edu/talks/samory-kpotufe-columbia-university-2026-02-26
Learning from Heterogeneous Sources
Multitask Learning refers to the problem of aggregating many datasets from separate source distributions to improve performance on a target prediction task. Our aim is to understand (1) sufficient and necessary conditions for speedup in convergence rates over vanilla prediction with just the target data, (2) how such speedup depends on the number of datasets and samples per dataset, and (3) whether such speedup is achievable adaptively, i.e., by procedures with no prior distributional information.
The picture turns out to be mixed as the problem displays sharp gaps between oracle rates and adaptive rates, i.e., there exist situations where no procedure can do better than using the target data alone even though a large subset of datasets are informative about the target task. On the other hand, a bit of information on the relation between sources and target distributions can allow for near optimal adaptive rates. These results lead to many interesting new questions which I'll attempt to properly convey.
The talk is based on various works with collaborators such as S. Hanneke, A. Gretton, M. Z. Li, D. Meunier.






![Latent Variable models and Subset Smoothing
Ravi Kannan (Simons Institute, UC Berkeley)
https://simons.berkeley.edu/talks/ravi-kannan-2026-05-26
The Role of TCS in Modern Machine Learning
A number of Latent Variable Models in Machine Learning (including Mixture Models, Topic Models, Stochastic block models and Mixed Membership Community Mod els) can be abstracted to the geometric problem of learn ing a latent polytope K given data points, each obtained by randomly perturbing a latent point in K. The challenge is that perturbations are typically much larger than the dimensions of K and so data points lie (far) outside K. To tackle this, we introduce the “Subset Smoothed” polytope K′ which is the convex hull of (n/k) points, each obtained by averaging a k− subset of the n data points. [k is a parameter.] We will observe that K′ ≈ K under reasonable assumptions on data. We will also observe that K′ has a polynomial time optimization oracle. These simple observations are the starting point of our provable algorithm for learning K which the talk will describe.
Joint Work with Chiranjib Bhattacharyya, Amit Kumar Latent Variable models and Subset Smoothing](https://i.ytimg.com/vi/Dm1YnND7Qmo/mqdefault.jpg)



