Uploaded August 2026 | Updated September 2026, 2 weeks ago
Molei Tao (Georgia Tech)
'https://simons.berkeley.edu/talks/molei-tao-georgia-tech-2026-08-04
Diffusion Generative Modeling: Progress and Next Steps
Diffusion model is a prevailing paradigm for generative AI, and this talk will briefly report two rigorous results that are centered around its performance.
First, I will quantify the generalization capability of the classical diffusion model, i.e. the one for Euclidean data. The question is, when the generative model is not memorizing the training data, what new samples will it generate? This question is not only pertinent to privacy and copyright considerations, but also important for understanding whether/how diffusion model creates new knowledge. The inductive bias of diffusion model’s generation will be examined, leading to a quantification of how diffusion model generalizes. This quantification is purely based on the empirical distribution without considering any population limit.
Then I will switch gear and describe a new test-time scaling method that improves masked diffusion model for discrete data, so that its output can maximize a given reward function, without any fine-tuning of the model weights. The method, MDM-VGB, is a discrete diffusion sampler that augments unmasking generation with principled reward-guided remasking. It is inspired by a recent success, VGB, that leverages the classical Jerrum-Sinclair backtracking Markov chain for the reward-tilted generation of autoregressive model; meanwhile, the Any-Order AutoRegressive nature of MDM-VGB allows more efficient error corrections. Besides strong empirical performance, we also prove that MDM-VGB is robust to process verifier noise, achieving a quadratic complexity, while popular test-time heuristics like best-of-N suffer from an exponential complexity due to accumulative errors.
Molei Tao (Georgia Tech)
'https://simons.berkeley.edu/talks/molei-tao-georgia-tech-2026-08-04
Diffusion Generative Modeling: Progress and Next Steps
Diffusion model is a prevailing paradigm for generative AI, and this talk will briefly report two rigorous results that are centered around its performance.
First, I will quantify the generalization capability of the classical diffusion model, i.e. the one for Euclidean data. The question is, when the generative model is not memorizing the training data, what new samples will it generate? This question is not only pertinent to privacy and copyright considerations, but also important for understanding whether/how diffusion model creates new knowledge. The inductive bias of diffusion model’s generation will be examined, leading to a quantification of how diffusion model generalizes. This quantification is purely based on the empirical distribution without considering any population limit.
Then I will switch gear and describe a new test-time scaling method that improves masked diffusion model for discrete data, so that its output can maximize a given reward function, without any fine-tuning of the model weights. The method, MDM-VGB, is a discrete diffusion sampler that augments unmasking generation with principled reward-guided remasking. It is inspired by a recent success, VGB, that leverages the classical Jerrum-Sinclair backtracking Markov chain for the reward-tilted generation of autoregressive model; meanwhile, the Any-Order AutoRegressive nature of MDM-VGB allows more efficient error corrections. Besides strong empirical performance, we also prove that MDM-VGB is robust to process verifier noise, achieving a quadratic complexity, while popular test-time heuristics like best-of-N suffer from an exponential complexity due to accumulative errors.










