Uploaded August 2026 | Updated September 2026, 2 weeks ago
Modern Machine Learning Methods: Large Step-Size Optimization, Implicit Bias, and Benign Overfitting
~
The impressive performance of modern machine learning methods seems to arise through different mechanisms from those of classical statistical learning theory, mathematical statistics, and optimization theory. Simple gradient methods find excellent solutions to non-convex optimization problems, and without any explicit effort to control model complexity they exhibit excellent prediction performance in practice. This talk will describe recent progress in statistical learning theory and optimization theory that demonstrates the optimization benefits of step-sizes that are too large to allow gradient methods to be viewed as an accurate time discretization of a gradient flow differential equation, that characterizes the solutions that are favored by gradient optimization methods, and that illustrates when those solutions can overfit training data but still provide good predictive accuracy.
Modern Machine Learning Methods: Large Step-Size Optimization, Implicit Bias, and Benign Overfitting
~
The impressive performance of modern machine learning methods seems to arise through different mechanisms from those of classical statistical learning theory, mathematical statistics, and optimization theory. Simple gradient methods find excellent solutions to non-convex optimization problems, and without any explicit effort to control model complexity they exhibit excellent prediction performance in practice. This talk will describe recent progress in statistical learning theory and optimization theory that demonstrates the optimization benefits of step-sizes that are too large to allow gradient methods to be viewed as an accurate time discretization of a gradient flow differential equation, that characterizes the solutions that are favored by gradient optimization methods, and that illustrates when those solutions can overfit training data but still provide good predictive accuracy.










