Uploaded August 2026 | Updated September 2026, 2 weeks ago
Zahra Kadkhodaie (MIT)
https://simons.berkeley.edu/talks/zahra-kadkhodaie-mit-2026-08-04
Diffusion Generative Modeling: Progress and Next Steps
Deep neural networks can learn powerful probability models of images, as demonstrated by the high-quality samples produced by score-based diffusion models. Yet it remains unclear how these networks capture complex global image statistics without suffering from the curse of dimensionality. I will present two complementary studies of this question in diffusion U-Nets. First, using a simplified multiscale U-Net, we show that coarse-to-fine conditioning in the decoder makes fine-scale image structure locally predictable: once conditioned on coarser-scale information, fine-scale features can be modeled using stationary local Markov models with small receptive fields. This transforms a high-dimensional modeling problem into a sequence of lower-dimensional conditional problems. Global structure, however, is handled differently at the U-Net bottleneck, where the spatial resolution is small enough for receptive fields to cover the entire image. In the second work, we investigate the representation that emerges in this bottleneck. We find that the middle block represents each image through a sparse subset of active channels, whose spatial averages form a nonlinear representation of the underlying clean image. Euclidean distances in this representation space reflect semantic similarity, even though the model is trained without external conditioning such as text or class labels. We further introduce a self-guided reconstruction method that uses a representation extracted by the unconditional model to condition its own stochastic synthesis, revealing the image features encoded in the bottleneck. Together, these results suggest that U-Nets manage image dimensionality by combining compact global representations with low-dimensional conditional models of local detail.
Zahra Kadkhodaie (MIT)
https://simons.berkeley.edu/talks/zahra-kadkhodaie-mit-2026-08-04
Diffusion Generative Modeling: Progress and Next Steps
Deep neural networks can learn powerful probability models of images, as demonstrated by the high-quality samples produced by score-based diffusion models. Yet it remains unclear how these networks capture complex global image statistics without suffering from the curse of dimensionality. I will present two complementary studies of this question in diffusion U-Nets. First, using a simplified multiscale U-Net, we show that coarse-to-fine conditioning in the decoder makes fine-scale image structure locally predictable: once conditioned on coarser-scale information, fine-scale features can be modeled using stationary local Markov models with small receptive fields. This transforms a high-dimensional modeling problem into a sequence of lower-dimensional conditional problems. Global structure, however, is handled differently at the U-Net bottleneck, where the spatial resolution is small enough for receptive fields to cover the entire image. In the second work, we investigate the representation that emerges in this bottleneck. We find that the middle block represents each image through a sparse subset of active channels, whose spatial averages form a nonlinear representation of the underlying clean image. Euclidean distances in this representation space reflect semantic similarity, even though the model is trained without external conditioning such as text or class labels. We further introduce a self-guided reconstruction method that uses a representation extracted by the unconditional model to condition its own stochastic synthesis, revealing the image features encoded in the bottleneck. Together, these results suggest that U-Nets manage image dimensionality by combining compact global representations with low-dimensional conditional models of local detail.










