Uploaded August 2026 | Updated September 2026, 1 week ago
Balakrishnan Raman, AMD and Sandeep Nagaraj, Meta present "MetaRoCE: From Spec to NIC to Open Source" live from the Santa Clara Convention Center.
As AI and distributed workloads push datacenter fabrics to their limits, Meta's answer is MetaRoCE — a multipath, out-of-order, receiver-driven protocol that treats Ethernet as inherently lossy and pushes all intelligence into the NIC. This talk traces MetaRoCE's journey from specification to silicon, showcasing its implementation on AMD Programmable NICs via reference software implementation and demonstrating how the architecture elegantly scales out within a cluster and scales across fabrics to meet diverse deployment needs.
We present real-world performance results spanning multiplane FPF topologies, tail-latency optimization, and long-distance RDMA — proving MetaRoce approach solves AI data center requirements. Finally, we present a first look at Meta's upcoming open-source release of the MetaRoCE specification, reference implementation, and compliance suites at OCP — inviting the industry to build, extend, and innovate on MetaRoCE.
Learn more about @Scale here: atscaleconference.com
Balakrishnan Raman, AMD and Sandeep Nagaraj, Meta present "MetaRoCE: From Spec to NIC to Open Source" live from the Santa Clara Convention Center.
As AI and distributed workloads push datacenter fabrics to their limits, Meta's answer is MetaRoCE — a multipath, out-of-order, receiver-driven protocol that treats Ethernet as inherently lossy and pushes all intelligence into the NIC. This talk traces MetaRoCE's journey from specification to silicon, showcasing its implementation on AMD Programmable NICs via reference software implementation and demonstrating how the architecture elegantly scales out within a cluster and scales across fabrics to meet diverse deployment needs.
We present real-world performance results spanning multiplane FPF topologies, tail-latency optimization, and long-distance RDMA — proving MetaRoce approach solves AI data center requirements. Finally, we present a first look at Meta's upcoming open-source release of the MetaRoCE specification, reference implementation, and compliance suites at OCP — inviting the industry to build, extend, and innovate on MetaRoCE.
Learn more about @Scale here: atscaleconference.com
![Evolving GenAI Media Infrastructure Deployments | Rushaan Mahajan, Sima Labs
This presentation by Rushaan Mahajan from Sima Labs explores how to optimize video generation infrastructure to meet the growing demand for personalized, high-quality video experiences.
Video generation is a computationally intensive process that requires significantly more resources than text generation, leading to high latency and costs. [00:35]
The key challenges in video inference include the iterative denoising loop, the memory and bandwidth bottleneck in the VAE decoder, and the need to optimize the entire runtime stack to achieve real-time, high-fidelity, and personalized video generation at scale.
Optimizing the balance between the VAE compression ratio and the denoiser complexity is crucial to reducing the overall computational cost of the video generation pipeline. [09:52]
Techniques like latent space compression, sampling optimization, caching, and pruning can significantly improve the runtime efficiency of video diffusion models without compromising quality. [11:39]
A multi-GPU strategy that utilizes different GPU types for different tasks (base generation, super-resolution, personalization) can help scale video generation without proportional cost increases. [13:35]
The goal is to make personalized video generation truly instant, transitioning it from a compute-bound novelty to a mainstream, interactive, and globally relevant platform. [13:57] Evolving GenAI Media Infrastructure Deployments | Rushaan Mahajan, Sima Labs](https://i.ytimg.com/vi/PrZpl4w1Lxk/mqdefault.jpg)









