Grounding Foundation Models for Embodied Intelligence @allenai
Grounding Foundation Models for Embodied Intelligence  @allenai
Uploaded November 2025 | Updated September 2026, 20 hours ago
Foundation models have shown remarkable generalization in language and vision, but making them useful for embodied intelligence requires new approaches to planning, execution, and reasoning. In this talk, I will present a line of work on grounding foundation models for robotics, progressing from leveraging existing models to improving their core capabilities. I begin with LLM-Planner, which investigates whether large language models can serve as high-level planners for embodied agents and how their plans can be grounded in the environment. I then introduce M-Track, a model-agnostic framework for long-horizon task execution that uses explicit, planner generated subgoals to monitor progress and keep agents on track. While these works demonstrate the promise of foundation models as planners, they also reveal a key limitation in their lack of robust spatial understanding. To address this, I present RoboSpatial, a large-scale dataset and benchmark designed to teach 3D spatial reasoning to foundation models for robot manipulation. Together, these works outline a path toward general-purpose robot foundation models that can generalize across tasks, remain reliable over long horizons, and acquire the spatial understanding needed for trustworthy autonomy.


Chan Hee (Luke) Song is a Ph.D. candidate in Computer Science at The Ohio State University, advised by Prof. Yu Su. His research focuses on building multimodal foundation models that unify perception, reasoning, and action for embodied intelligence across physical and virtual environments, with an emphasis on spatial understanding, long-horizon planning, and large-scale dataset and benchmark construction. He has published in CVPR, ICLR, ICCV, NeurIPS, AAAI, and COLM, and received the CVPR 2024 Best Student Paper Award. He currently serves as an Area Chair for ICLR and has spent time at Google Research, NVIDIA Research, and Adobe Research, working on multimodal agents and robot foundation models. More about his work can be found at https://chanh.ee
Grounding Foundation Models for Embodied IntelligenceToward Intelligent Writing Support Beyond Completing SentencesEnabling Scientific Research with Language AgentsDREAM: Improving Situational QA by First Elaborating the SituationObjective Mismatch in Reinforcement Learning from Human FeedbackAvenging Polanyis RevengeA Flexible Framework for Machine Learning | AI2Applied AI in High-Expertise Settings, or Curation as Programming
Ai2 |

Grounding Foundation Models for Embodied Intelligence

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER