Uploaded April 2025 | Updated September 2026, 2 weeks ago
Talk Title: Weak to Strong Generalization in Random Feature Models
Speaker: Nathan Srebro (TTIC)
Date: April 7, 2025
Abstract: Weak-to-Strong Generalization (Burns et al., 2023) is the phenomenon whereby a strong student, say GPT-4, learns a task from a weak teacher, say GPT-2, and ends up significantly outperforming the teacher. We show that this phenomenon does not require a strong and complex learner like GPT-4, nor pre-training. We consider students and teachers that are random feature models, described by two-layer networks with a random and fixed bottom layer and trained top layer. A weak teacher, with a small number of units (i.e. random features), is trained on the population, and a strong student, with a much larger number of units (i.e. random features), is trained only on labels generated by the weak teacher. We demonstrate, prove and understand, how the student can outperform the teacher, even though trained only on data labeled by the teacher, with no pretraining or other knowledge or data advantage over the teacher. We explain how such weak-to-strong generalization is enabled by early stopping. Importantly, we also show the quantitative limits of weak-to-strong generalization in this model. Joint work with Marko Medvedev, Kaifeng Lyu, Dingli Yu, Sanjeev Arora and Zhiyuan Li.
Bio: Nathan (Nati) Srebro is a professor at the Toyota Technological Institute at Chicago and the University of Chicago. He obtained his PhD at the Massachusetts Institute of Technology (MIT) in 2004, and previously was a post-doctoral fellow at the University of Toronto, a Visiting Scientist at IBM, and an Associate Professor at the Technion. He has also held visiting positions in UC Berkeley, EPFL and the Weizmann Institute of Science, and was consulting faculty with Google and Microsoft. Some of Prof. Srebro’s significant contributions include work on learning wider Markov networks; introducing the use of the nuclear norm for machine learning and matrix reconstruction; work on fast optimization techniques for machine learning the optimality of stochastic methods, and on the relationship between learning and optimization more broadly; study of fairness measures for non-discrimination and introduction of equalized odds; and highlighting the importance of implicit optimization bias as a central driving force in deep learning. His work has been recognized by multiple Best Paper awards at COLT (Conference on Learning Theory), ICML (International Conference on Machine Learning), and UAI (Uncertainty in Artificial Intelligence), as well a ten-year Test of Time award at ICML.
This video is closed captioned.
Talk Title: Weak to Strong Generalization in Random Feature Models
Speaker: Nathan Srebro (TTIC)
Date: April 7, 2025
Abstract: Weak-to-Strong Generalization (Burns et al., 2023) is the phenomenon whereby a strong student, say GPT-4, learns a task from a weak teacher, say GPT-2, and ends up significantly outperforming the teacher. We show that this phenomenon does not require a strong and complex learner like GPT-4, nor pre-training. We consider students and teachers that are random feature models, described by two-layer networks with a random and fixed bottom layer and trained top layer. A weak teacher, with a small number of units (i.e. random features), is trained on the population, and a strong student, with a much larger number of units (i.e. random features), is trained only on labels generated by the weak teacher. We demonstrate, prove and understand, how the student can outperform the teacher, even though trained only on data labeled by the teacher, with no pretraining or other knowledge or data advantage over the teacher. We explain how such weak-to-strong generalization is enabled by early stopping. Importantly, we also show the quantitative limits of weak-to-strong generalization in this model. Joint work with Marko Medvedev, Kaifeng Lyu, Dingli Yu, Sanjeev Arora and Zhiyuan Li.
Bio: Nathan (Nati) Srebro is a professor at the Toyota Technological Institute at Chicago and the University of Chicago. He obtained his PhD at the Massachusetts Institute of Technology (MIT) in 2004, and previously was a post-doctoral fellow at the University of Toronto, a Visiting Scientist at IBM, and an Associate Professor at the Technion. He has also held visiting positions in UC Berkeley, EPFL and the Weizmann Institute of Science, and was consulting faculty with Google and Microsoft. Some of Prof. Srebro’s significant contributions include work on learning wider Markov networks; introducing the use of the nuclear norm for machine learning and matrix reconstruction; work on fast optimization techniques for machine learning the optimality of stochastic methods, and on the relationship between learning and optimization more broadly; study of fairness measures for non-discrimination and introduction of equalized odds; and highlighting the importance of implicit optimization bias as a central driving force in deep learning. His work has been recognized by multiple Best Paper awards at COLT (Conference on Learning Theory), ICML (International Conference on Machine Learning), and UAI (Uncertainty in Artificial Intelligence), as well a ten-year Test of Time award at ICML.
This video is closed captioned.


