Structured and Efficient Representations for Robot Learning @allenai
Structured and Efficient Representations for Robot Learning  @allenai
Uploaded October 2025 | Updated September 2026, 2 hours ago
Biological intelligence achieves remarkable generalization and rapid adaptation through compact, efficient representations. Inspired by insights from neuroscience, Sumeet's work seeks to give embodied artificial agents
similarly powerful abstractions for both behavior and perception. In this talk, he presents work in two main directions: (1) defining representations for behavior that enhance learning from expert data, and (2) developing self-supervised, structured visual representations from RGB data that enable zero-shot generalization for real-world robots. He concludes by outlining future research directions and discussing how deliberate, principled choices of representation can unlock greater generalizability in modern robotic systems.

Sumeet is a fifth-year PhD candidate at the University of Southern California, specializing in machine learning and robotics. His research draws inspiration from neuroscience to design representations that enable generalization and rapid adaptation in embodied agents. His work spans reinforcement learning for robotics, diffusion models as control policies, and self-supervised learning for extracting structured visual representations that support zero-shot generalization on real robots. During his internships with NVIDIA’s autonomous vehicles research team, he developed diffusion-based scenario generation tools and high-throughput simulation pipelines for RL training. His overarching goal is to create adaptable, robust, and intelligent robotic systems capable of operating effectively in complex, real-world environments.
Structured and Efficient Representations for Robot LearningIntroducing MolmoBot | Sim-to-real zero shot transfer for roboticsWelcoming AI as a New Colleague: How Should We Evaluate AI for Science?Understanding and Improving Compositional Generalization | AI2Molmo 2 | A new standard for open video intelligenceMolmo 2 | Reasoning across documents and imagesHow far have we come in giving our NLU systems common sense?🚀 Molmo Robotic Demo: AI in ActionOn the Symbiosis of Generative Models and Representation LearningThe BigScience WorkshopThe Pre-trainers toolkit: From dataset construction to model scalingFrom Compression to Convection: A Latent Variable Perspective
Ai2 |

Structured and Efficient Representations for Robot Learning

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER