Uploaded March 2026 | Updated September 2026, 2 weeks ago
Sanmi Koyejo (Stanford University)
https://simons.berkeley.edu/talks/sanmi-koyejo-stanford-university-2026-02-27
Learning from Heterogeneous Sources
Federated and collaborative learning methods proliferate, yet our understanding of when and why they work lags behind empirical results. A central challenge is heterogeneity: how do we characterize it, measure it, and design algorithms that handle it? Drawing on recent work in distribution shift, I argue that the field's treatment of "heterogeneity" as monolithic obscures critical distinctions between interpolation, adaptation, and generalization scenarios—each requiring different theoretical and algorithmic approaches.
I will present a taxonomy of data and algorithmic interventions for distribution shifts, and translate its implications for federated settings: When does model averaging perform safe interpolation vs. risky extrapolation across clients? What is the fundamental tradeoff between personalization and generalization, and are we optimizing for the wrong objective? Do current benchmarks measure worst-case heterogeneity, or just benign shifts? I will close with open problems at the intersection of measurement science and federated learning: how do we design benchmarks with construct validity and adapt evaluation frameworks to match real-world collaborative learning scenarios?
Sanmi Koyejo (Stanford University)
https://simons.berkeley.edu/talks/sanmi-koyejo-stanford-university-2026-02-27
Learning from Heterogeneous Sources
Federated and collaborative learning methods proliferate, yet our understanding of when and why they work lags behind empirical results. A central challenge is heterogeneity: how do we characterize it, measure it, and design algorithms that handle it? Drawing on recent work in distribution shift, I argue that the field's treatment of "heterogeneity" as monolithic obscures critical distinctions between interpolation, adaptation, and generalization scenarios—each requiring different theoretical and algorithmic approaches.
I will present a taxonomy of data and algorithmic interventions for distribution shifts, and translate its implications for federated settings: When does model averaging perform safe interpolation vs. risky extrapolation across clients? What is the fundamental tradeoff between personalization and generalization, and are we optimizing for the wrong objective? Do current benchmarks measure worst-case heterogeneity, or just benign shifts? I will close with open problems at the intersection of measurement science and federated learning: how do we design benchmarks with construct validity and adapt evaluation frameworks to match real-world collaborative learning scenarios?










