Uploaded June 2026 | Updated September 2026, 1 week ago
Architecting Infrastructure for the AI Native Future: Scaling Autonomous Agents on Google TPUs by Sabastian Mugazambi from Google
As the industry pivots toward an "AI Native" paradigm, the bottleneck for innovation has shifted from algorithmic design to the underlying infrastructure's ability to handle unprecedented scale and complexity. This session explores how Google TPU (Tensor Processing Unit) infrastructure serves as the catalyst for this transformation, specifically within the domains of large-scale Recommender Systems, MoEs, LLMs and the emerging era of Autonomous Agents.
We will delve into the architectural innovations of the latest TPU generations, demonstrating how their purpose-built design facilitates the massive throughput required for real-time recommendation engines and the high-speed inference necessary for agentic orchestration.
Want to learn more about the @Scale Conference? Visit our website here: atscaleconference.com
Architecting Infrastructure for the AI Native Future: Scaling Autonomous Agents on Google TPUs by Sabastian Mugazambi from Google
As the industry pivots toward an "AI Native" paradigm, the bottleneck for innovation has shifted from algorithmic design to the underlying infrastructure's ability to handle unprecedented scale and complexity. This session explores how Google TPU (Tensor Processing Unit) infrastructure serves as the catalyst for this transformation, specifically within the domains of large-scale Recommender Systems, MoEs, LLMs and the emerging era of Autonomous Agents.
We will delve into the architectural innovations of the latest TPU generations, demonstrating how their purpose-built design facilitates the massive throughput required for real-time recommendation engines and the high-speed inference necessary for agentic orchestration.
Want to learn more about the @Scale Conference? Visit our website here: atscaleconference.com









![Evolving GenAI Media Infrastructure Deployments | Rushaan Mahajan, Sima Labs
This presentation by Rushaan Mahajan from Sima Labs explores how to optimize video generation infrastructure to meet the growing demand for personalized, high-quality video experiences.
Video generation is a computationally intensive process that requires significantly more resources than text generation, leading to high latency and costs. [00:35]
The key challenges in video inference include the iterative denoising loop, the memory and bandwidth bottleneck in the VAE decoder, and the need to optimize the entire runtime stack to achieve real-time, high-fidelity, and personalized video generation at scale.
Optimizing the balance between the VAE compression ratio and the denoiser complexity is crucial to reducing the overall computational cost of the video generation pipeline. [09:52]
Techniques like latent space compression, sampling optimization, caching, and pruning can significantly improve the runtime efficiency of video diffusion models without compromising quality. [11:39]
A multi-GPU strategy that utilizes different GPU types for different tasks (base generation, super-resolution, personalization) can help scale video generation without proportional cost increases. [13:35]
The goal is to make personalized video generation truly instant, transitioning it from a compute-bound novelty to a mainstream, interactive, and globally relevant platform. [13:57] Evolving GenAI Media Infrastructure Deployments | Rushaan Mahajan, Sima Labs](https://i.ytimg.com/vi/PrZpl4w1Lxk/mqdefault.jpg)
