Text-to-video models explained @GoogleResearch
Text-to-video models explained  @GoogleResearch
Uploaded February 2023 | Updated September 2026, 2 weeks ago
When it comes to text-to-video models, the way they create clips is very similar to the way text-to-image models create images from simple text prompts. In this episode of Hidden Layers, we take a look at how these models operate under the hood - understanding how it uses Temporal Super Resolution and Spatial Super Resolution (SSR) models to create high resolution videos from frames of images. Moreover, you’ll learn how text-to-video models - like Imagen- are
an orchestration of various models working together to produce high resolution videos from a single image and text prompt.

Resources:
Watch our previous episode → https://goo.gle/3XX0bxv
Check out Imagen → https://goo.gle/3yvFNsS

Chapters:
0:00 - Intro
0:16 - What are text-to-video models?
0:34 - How do text-to-video models create videos?
1:40 - What are the complexities of modeling video?
2:04 - How do we get high-resolution videos from text-to-video models like Imagen?
3:38 - Recap of how Imagen works
3:47 - Leave us questions in the comments!

Watch more episodes of Hidden Layers→ https://goo.gle/HiddenLayers
Subscribe to the Google Research Channel → https://goo.gle/GoogleResearch
Text-to-video models explainedIntroducing Lumiere: A space-time diffusion model for video generationWhat is Responsible AI? | Research BytesResearch@ NYC: PAIR GuidebookGeospatial Reasoning: Unlocking insights with generative AI and multiple foundation modelsImagining with ImagenResearch Retrospectives: an interview with Zoubin Ghahramani, VP of Google ResearchPicturing the world with PartiResearch@ NYC: TextFXBefore Rivers Rise | Using AI to Make Global Flood ForecastsAI for maternal and child healthCan AI help predict cyclones? | Behind the Breakthroughs
Google Research |

Text-to-video models explained

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER