Evolving GenAI Media Infrastructure Deployments | Rushaan Mahajan, Sima Labs @scaleconference
Evolving GenAI Media Infrastructure Deployments | Rushaan Mahajan, Sima Labs  @scaleconference
Uploaded October 2025 | Updated September 2026, 1 week ago
This presentation by Rushaan Mahajan from Sima Labs explores how to optimize video generation infrastructure to meet the growing demand for personalized, high-quality video experiences.

Video generation is a computationally intensive process that requires significantly more resources than text generation, leading to high latency and costs. [00:35]

The key challenges in video inference include the iterative denoising loop, the memory and bandwidth bottleneck in the VAE decoder, and the need to optimize the entire runtime stack to achieve real-time, high-fidelity, and personalized video generation at scale.

Optimizing the balance between the VAE compression ratio and the denoiser complexity is crucial to reducing the overall computational cost of the video generation pipeline. [09:52]

Techniques like latent space compression, sampling optimization, caching, and pruning can significantly improve the runtime efficiency of video diffusion models without compromising quality. [11:39]

A multi-GPU strategy that utilizes different GPU types for different tasks (base generation, super-resolution, personalization) can help scale video generation without proportional cost increases. [13:35]

The goal is to make personalized video generation truly instant, transitioning it from a compute-bound novelty to a mainstream, interactive, and globally relevant platform. [13:57]
Evolving GenAI Media Infrastructure Deployments | Rushaan Mahajan, Sima LabsOur Journey to Safely Unleash Agents at Meta Scale | David Pariag from MetaLive from SCCC: MetaRoCE: Meta’s RDMA Transport | Arvind Srinivasan and Kingshuk MandalInside Instagrams VR Revolution: AI-Powered 3D Content at Scale🚀 Registration is LIVE!Co-Designing Communication for AI Accelerators | Rajeev Nair and Wes Bland from MetaTrack 1 - Live Q&A Session #2 by Shashi GandhamMetaRoCE: Meta’s RDMA Transport | Arvind Srinivasan, Meta and Kingshuk Mandal, KeysightEnhancing Runtime Reliability in LLM Training via Fine-Grained Observability - Live from SCCPhysical Network Design at Scale | Brandon Premo and Richard Cziva from MetaPerformance Optimizations at 100K+ Scale - Live from SCCSecuring Production Debugging at Hyperscale | Shridivya Sharma and Luke Kelly from Microsoft
@Scale |

Evolving GenAI Media Infrastructure Deployments | Rushaan Mahajan, Sima Labs

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER