Uploaded June 2026 | Updated September 2026, 1 week ago
Speaker: Joe Duffy from Pulumi
To handle the scale and velocity of AI-written code, we will have no choice but to let AI manage our infrastructure too. Yet Andrej Karpathy recently described getting an app running in production as “assembling IKEA furniture”: cloud consoles, API keys, copy-pasted config, glue, ... all things that sit outside the code an LLM can reason about. Frontier models are trained on billions of lines of real languages like Python, TypeScript, and Go, and vanishingly little bespoke DSLs and manual procedures. By modeling infrastructure in code space so the LLM can do what it does best — code — we just need an oracle that can map code changes back to infrastructure outcomes. In this talk, I’ll share what we’ve learned at Pulumi working alongside leading AI companies and frontier labs to build for a world where agents manage infrastructure. The platforms that win in this new era will look different, but many of the human-ergonomic benefits of programming languages are what will get us there.
Learn more about the @Scale conferences here: atscaleconference.com
Speaker: Joe Duffy from Pulumi
To handle the scale and velocity of AI-written code, we will have no choice but to let AI manage our infrastructure too. Yet Andrej Karpathy recently described getting an app running in production as “assembling IKEA furniture”: cloud consoles, API keys, copy-pasted config, glue, ... all things that sit outside the code an LLM can reason about. Frontier models are trained on billions of lines of real languages like Python, TypeScript, and Go, and vanishingly little bespoke DSLs and manual procedures. By modeling infrastructure in code space so the LLM can do what it does best — code — we just need an oracle that can map code changes back to infrastructure outcomes. In this talk, I’ll share what we’ve learned at Pulumi working alongside leading AI companies and frontier labs to build for a world where agents manage infrastructure. The platforms that win in this new era will look different, but many of the human-ergonomic benefits of programming languages are what will get us there.
Learn more about the @Scale conferences here: atscaleconference.com


![Evolving GenAI Media Infrastructure Deployments | Rushaan Mahajan, Sima Labs
This presentation by Rushaan Mahajan from Sima Labs explores how to optimize video generation infrastructure to meet the growing demand for personalized, high-quality video experiences.
Video generation is a computationally intensive process that requires significantly more resources than text generation, leading to high latency and costs. [00:35]
The key challenges in video inference include the iterative denoising loop, the memory and bandwidth bottleneck in the VAE decoder, and the need to optimize the entire runtime stack to achieve real-time, high-fidelity, and personalized video generation at scale.
Optimizing the balance between the VAE compression ratio and the denoiser complexity is crucial to reducing the overall computational cost of the video generation pipeline. [09:52]
Techniques like latent space compression, sampling optimization, caching, and pruning can significantly improve the runtime efficiency of video diffusion models without compromising quality. [11:39]
A multi-GPU strategy that utilizes different GPU types for different tasks (base generation, super-resolution, personalization) can help scale video generation without proportional cost increases. [13:35]
The goal is to make personalized video generation truly instant, transitioning it from a compute-bound novelty to a mainstream, interactive, and globally relevant platform. [13:57] Evolving GenAI Media Infrastructure Deployments | Rushaan Mahajan, Sima Labs](https://i.ytimg.com/vi/PrZpl4w1Lxk/mqdefault.jpg)







