Scaling Llama4 Training to 100K by Saif Hasan @scaleconference
Scaling Llama4 Training to 100K by Saif Hasan  @scaleconference
Uploaded August 2025 | Updated September 2026, 1 week ago
Llama 4's pre-training scale is growing exponentially, with 100K GPUs used, a 6x increase from its predecessor. Initializing training takes longer, and failure probability increases with larger scale. Training throughput aka Effective Training time degrades significantly as a result.

To address these challenges, researchers are experimenting in parallel for faster initialization of large scale jobs, and fault-tolerant paradigms.

Learn more here: atscaleconference.com
Scaling Llama4 Training to 100K by Saif HasanPyTorch Symmetric Memory: A New Paradigm for Programming Distributed AIThe Next Frontier for AI AgentsLive Panel: Infrastructure in an Agentic World | Moderated by Karthik LakshminarayananOpening Keynote | Surupa Biswas from MetaWhy Have We Not Solved Security of Agents? | Ilia ShumailovHow AI Audio Separation is Revolutionizing Media Production 🎵Live from SCCC: Lightning Talks Introduction - AI for NI | Mohab Gawish, MetaWhat if your network capacity plan could learn from every decision?GenAI Research for Creativity & Productivity | Stefano Corazza, CanvaLive Panel: Building at the Roofline — Planning and Delivering Gigawatt-Scale AI InfrastructureMeta App Quality | Bruce Chhay
@Scale |

Scaling Llama4 Training to 100K by Saif Hasan

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER