Uploaded March 2026 | Updated September 2026, 2 weeks ago
Kevin Kuo (Carnegie Mellon University)
https://simons.berkeley.edu/talks/kevin-kuo-carnegie-mellon-university-2026-02-24
Learning from Heterogeneous Sources
To improve trust and adoption of collaborative learning systems, it is important to provide opt-out guarantees which enable removal of unwanted (e.g. harmful or private) information after it has been used to train a model. A popular solution to this problem is approximate unlearning, which efficiently updates an LLM so that it behaves (roughly) as if it was not trained on a subset of data to begin with. However, existing methods are brittle in practice and can easily be attacked to reveal supposedly unlearned information. To alleviate issues with approximate unlearning, we instead propose SIFT-Masks (SIgn-Fixed Tuning-Masks), an exact unlearning method based on model merging. SIFT-Masks addresses two key limitations of standard model merging: (1) merging a large number of tasks can severely harm utility; and (2) methods that boost utility by sharing extra information across tasks make exact unlearning prohibitively expensive. SIFT-Masks solves these issues by (1) applying local masks to recover task-specific performance; and (2) constraining finetuning to align with a global sign vector as a lightweight approach to determine masks independently before merging. Across four settings where we merge up to 500 models, SIFT-Masks improves accuracy by 5-80% over naive merging and uses up to 250x less compute for exact unlearning compared to other merging baselines.
Kevin Kuo (Carnegie Mellon University)
https://simons.berkeley.edu/talks/kevin-kuo-carnegie-mellon-university-2026-02-24
Learning from Heterogeneous Sources
To improve trust and adoption of collaborative learning systems, it is important to provide opt-out guarantees which enable removal of unwanted (e.g. harmful or private) information after it has been used to train a model. A popular solution to this problem is approximate unlearning, which efficiently updates an LLM so that it behaves (roughly) as if it was not trained on a subset of data to begin with. However, existing methods are brittle in practice and can easily be attacked to reveal supposedly unlearned information. To alleviate issues with approximate unlearning, we instead propose SIFT-Masks (SIgn-Fixed Tuning-Masks), an exact unlearning method based on model merging. SIFT-Masks addresses two key limitations of standard model merging: (1) merging a large number of tasks can severely harm utility; and (2) methods that boost utility by sharing extra information across tasks make exact unlearning prohibitively expensive. SIFT-Masks solves these issues by (1) applying local masks to recover task-specific performance; and (2) constraining finetuning to align with a global sign vector as a lightweight approach to determine masks independently before merging. Across four settings where we merge up to 500 models, SIFT-Masks improves accuracy by 5-80% over naive merging and uses up to 250x less compute for exact unlearning compared to other merging baselines.


![Latent Variable models and Subset Smoothing
Ravi Kannan (Simons Institute, UC Berkeley)
https://simons.berkeley.edu/talks/ravi-kannan-2026-05-26
The Role of TCS in Modern Machine Learning
A number of Latent Variable Models in Machine Learning (including Mixture Models, Topic Models, Stochastic block models and Mixed Membership Community Mod els) can be abstracted to the geometric problem of learn ing a latent polytope K given data points, each obtained by randomly perturbing a latent point in K. The challenge is that perturbations are typically much larger than the dimensions of K and so data points lie (far) outside K. To tackle this, we introduce the “Subset Smoothed” polytope K′ which is the convex hull of (n/k) points, each obtained by averaging a k− subset of the n data points. [k is a parameter.] We will observe that K′ ≈ K under reasonable assumptions on data. We will also observe that K′ has a polynomial time optimization oracle. These simple observations are the starting point of our provable algorithm for learning K which the talk will describe.
Joint Work with Chiranjib Bhattacharyya, Amit Kumar Latent Variable models and Subset Smoothing](https://i.ytimg.com/vi/Dm1YnND7Qmo/mqdefault.jpg)







