Uploaded February 2024 | Updated September 2026, 2 weeks ago
AI systems like GPT-4 have made headlines for how well they learn and use human language, but they do that by ingesting astronomical amounts of data from the internet—more text a human would encounter in 100,000 years. Human babies, meanwhile, learn words with much less input, just by absorbing what's in their own environment. What would happen if an AI system had to learn words the way kids do, based only on what a single toddler sees and hears? NYU data science researchers recently Wai Keen Vong and Brenden Lake conducted that exact experiment, using video and audio captured from a camera mounted to a child's head over a period of months to train a multimodal neural network. The results—published in the journal Science—shed light on long-standing debates on language acquisition processes in children, as well as on what it would mean to make AI learning processes more childlike, and potentially more efficient.
AI systems like GPT-4 have made headlines for how well they learn and use human language, but they do that by ingesting astronomical amounts of data from the internet—more text a human would encounter in 100,000 years. Human babies, meanwhile, learn words with much less input, just by absorbing what's in their own environment. What would happen if an AI system had to learn words the way kids do, based only on what a single toddler sees and hears? NYU data science researchers recently Wai Keen Vong and Brenden Lake conducted that exact experiment, using video and audio captured from a camera mounted to a child's head over a period of months to train a multimodal neural network. The results—published in the journal Science—shed light on long-standing debates on language acquisition processes in children, as well as on what it would mean to make AI learning processes more childlike, and potentially more efficient.



![How Naturalistic Learning Algorithm PooDLe Works
Most AI systems trained on carefully curated datasets struggle when faced with the messy reality of naturalistic video. PooDLe [https://arxiv.org/abs/2408.11208], a newly published work, addresses this problem.
“We were interested in building embodied learning algorithms that could learn directly from video streams. Says Mengye Ren, author of the paper and assistant professor of computer science and data science at NYU.
The algorithms name comes from a combination of a pooled loss function and a dense loss function used in PooDLes architecture.
Visit nyu.edu/news for updates on this story and more. How Naturalistic Learning Algorithm PooDLe Works](https://i.ytimg.com/vi/uUotGDIibu8/mqdefault.jpg)






