Uploaded June 2025 | Updated September 2026, 1 week ago
Presenter: Aparna Ramani, Vice President of Engineering at Meta
Key Points:
- Meta has over a million GPUs in their fleet
- Models have evolved to "think and reason" (referencing OpenAI's o1 model)
- Reinforcement learning has been crucial for model improvement
Evolution of AI Products
- Moving from chatbots to agents that can perform larger tasks and use tools
- Progressing from image generation to video generation
- Expanding AI across more platforms (WhatsApp, wearables like glasses)
Infrastructure Challenges
- Building for scale at every layer of the stack
- Creating massive interconnected clusters (equivalent to a gigawatt in power)
- Addressing networking challenges with new protocols and communication libraries
- Developing fault tolerance for "a million GPUs that can fail a million ways"
Supporting multiple types of hardware accelerators within the same infrastructure
Reinforcement Learning at Scale
- RL setup is vastly more complex than pre-training
- Requires incorporating inference optimizations into the training loop
- Needs scaled environments (coding sandboxes, web browser simulators) for models to learn in
Learn more about the @Scale conference here: atscaleconference.com/events/scale-data-ai-infra
Presenter: Aparna Ramani, Vice President of Engineering at Meta
Key Points:
- Meta has over a million GPUs in their fleet
- Models have evolved to "think and reason" (referencing OpenAI's o1 model)
- Reinforcement learning has been crucial for model improvement
Evolution of AI Products
- Moving from chatbots to agents that can perform larger tasks and use tools
- Progressing from image generation to video generation
- Expanding AI across more platforms (WhatsApp, wearables like glasses)
Infrastructure Challenges
- Building for scale at every layer of the stack
- Creating massive interconnected clusters (equivalent to a gigawatt in power)
- Addressing networking challenges with new protocols and communication libraries
- Developing fault tolerance for "a million GPUs that can fail a million ways"
Supporting multiple types of hardware accelerators within the same infrastructure
Reinforcement Learning at Scale
- RL setup is vastly more complex than pre-training
- Requires incorporating inference optimizations into the training loop
- Needs scaled environments (coding sandboxes, web browser simulators) for models to learn in
Learn more about the @Scale conference here: atscaleconference.com/events/scale-data-ai-infra










