Uploaded August 2026 | Updated September 2026, 2 weeks ago
Foundation Models for Avatar Generation | Vittorio Ferrari (Meta) | Realistic AI Avatars, Generative AI & 3D Digital Humans | Computing Conference 2026
How are foundation models transforming the future of realistic AI avatars and digital humans?
In this keynote from the 14th Computing Conference 2026, Vittorio Ferrari, Principal Research Scientist at Meta, explores how the same foundation model paradigm that revolutionized large language models, computer vision, and image generation is now driving the next generation of 2D and 3D avatar generation.
Drawing on his leadership experience at Meta, Synthesia, Google, the University of Edinburgh, and ETH Zurich, Vittorio Ferrari explains how large-scale pretraining enables AI systems to generate highly realistic, expressive, and personalized digital humans with minimal input data. Throughout the presentation, he demonstrates how foundation models are making avatar creation more scalable, efficient, and commercially viable while opening exciting opportunities for immersive communication, virtual collaboration, and AI-powered content creation.
In this keynote, you'll discover:
• What foundation models are and why they have transformed modern artificial intelligence
• How large-scale pretraining enables powerful general-purpose representations
• The evolution of AI-powered avatar generation from traditional capture studios to foundation models
• Meta's research on realistic 3D avatars using Large Avatar Creator (LCA)
• How internet-scale video datasets improve avatar quality and generalization
• The role of transformers in learning realistic human appearance and motion
• Advances in personalized avatars that capture individual expressions and behaviors
• How Synthesia applies generative AI to create photorealistic talking avatars
• Emerging applications in virtual communication, enterprise training, mixed reality, wearable devices, and digital content production
• Future research directions for expressive, interactive, and personalized digital humans
The keynote also highlights practical demonstrations comparing traditional avatar pipelines with foundation-model-based approaches, showing how AI can reconstruct high-quality digital humans from only a handful of images while maintaining realism, animation quality, and real-time performance. Additional examples showcase advances in personalized motion, conversational avatars, and multimodal generative AI systems that combine vision, speech, language, and 3D understanding.
About the Speaker
Vittorio Ferrari is a Principal Research Scientist at Meta, where he works on realistic avatars for next-generation communication on wearable devices. Previously, he served as Director of Science at Synthesia, leading research in generative AI for digital humans.
His distinguished career also includes leadership positions at Google, the University of Edinburgh, and ETH Zurich. Vittorio has authored more than 160 scientific publications, received the prestigious ERC Starting Grant, won the ECCV Best Paper Award, helped create the globally influential Open Images dataset, and has contributed AI technologies deployed across products including Google Photos, Google Lens, and Pixel devices.
His research spans:
• Generative AI
• Computer Vision
• Vision-Language Models
• 3D Deep Learning
• Digital Humans
• Generative Video
• Machine Learning
This keynote was presented at the 14th Computing Conference 2026, bringing together leading researchers, academics, and industry experts to discuss cutting-edge developments in artificial intelligence, machine learning, computer vision, data science, cybersecurity, software engineering, and emerging computing technologies.
If you enjoyed this keynote, don't forget to:
👍 Like the video
💬 Share your thoughts in the comments
🔔 Subscribe for more keynote talks, AI research presentations, and conference sessions from Computing Conference.
#FoundationModels #AvatarGeneration #GenerativeAI #MetaAI #VittorioFerrari #DigitalHumans #3DAvatars #ComputerVision #MachineLearning #ArtificialIntelligence #DeepLearning #Transformers #VisionLanguageModels #Synthesia #AIResearch #VirtualHumans #ComputingConference2026 #AIConference #MixedReality #FutureOfAI
Foundation Models for Avatar Generation | Vittorio Ferrari (Meta) | Realistic AI Avatars, Generative AI & 3D Digital Humans | Computing Conference 2026
How are foundation models transforming the future of realistic AI avatars and digital humans?
In this keynote from the 14th Computing Conference 2026, Vittorio Ferrari, Principal Research Scientist at Meta, explores how the same foundation model paradigm that revolutionized large language models, computer vision, and image generation is now driving the next generation of 2D and 3D avatar generation.
Drawing on his leadership experience at Meta, Synthesia, Google, the University of Edinburgh, and ETH Zurich, Vittorio Ferrari explains how large-scale pretraining enables AI systems to generate highly realistic, expressive, and personalized digital humans with minimal input data. Throughout the presentation, he demonstrates how foundation models are making avatar creation more scalable, efficient, and commercially viable while opening exciting opportunities for immersive communication, virtual collaboration, and AI-powered content creation.
In this keynote, you'll discover:
• What foundation models are and why they have transformed modern artificial intelligence
• How large-scale pretraining enables powerful general-purpose representations
• The evolution of AI-powered avatar generation from traditional capture studios to foundation models
• Meta's research on realistic 3D avatars using Large Avatar Creator (LCA)
• How internet-scale video datasets improve avatar quality and generalization
• The role of transformers in learning realistic human appearance and motion
• Advances in personalized avatars that capture individual expressions and behaviors
• How Synthesia applies generative AI to create photorealistic talking avatars
• Emerging applications in virtual communication, enterprise training, mixed reality, wearable devices, and digital content production
• Future research directions for expressive, interactive, and personalized digital humans
The keynote also highlights practical demonstrations comparing traditional avatar pipelines with foundation-model-based approaches, showing how AI can reconstruct high-quality digital humans from only a handful of images while maintaining realism, animation quality, and real-time performance. Additional examples showcase advances in personalized motion, conversational avatars, and multimodal generative AI systems that combine vision, speech, language, and 3D understanding.
About the Speaker
Vittorio Ferrari is a Principal Research Scientist at Meta, where he works on realistic avatars for next-generation communication on wearable devices. Previously, he served as Director of Science at Synthesia, leading research in generative AI for digital humans.
His distinguished career also includes leadership positions at Google, the University of Edinburgh, and ETH Zurich. Vittorio has authored more than 160 scientific publications, received the prestigious ERC Starting Grant, won the ECCV Best Paper Award, helped create the globally influential Open Images dataset, and has contributed AI technologies deployed across products including Google Photos, Google Lens, and Pixel devices.
His research spans:
• Generative AI
• Computer Vision
• Vision-Language Models
• 3D Deep Learning
• Digital Humans
• Generative Video
• Machine Learning
This keynote was presented at the 14th Computing Conference 2026, bringing together leading researchers, academics, and industry experts to discuss cutting-edge developments in artificial intelligence, machine learning, computer vision, data science, cybersecurity, software engineering, and emerging computing technologies.
If you enjoyed this keynote, don't forget to:
👍 Like the video
💬 Share your thoughts in the comments
🔔 Subscribe for more keynote talks, AI research presentations, and conference sessions from Computing Conference.
#FoundationModels #AvatarGeneration #GenerativeAI #MetaAI #VittorioFerrari #DigitalHumans #3DAvatars #ComputerVision #MachineLearning #ArtificialIntelligence #DeepLearning #Transformers #VisionLanguageModels #Synthesia #AIResearch #VirtualHumans #ComputingConference2026 #AIConference #MixedReality #FutureOfAI






