Uploaded May 2018 | Updated September 2026, 2 weeks ago
ICRA 2018 Spotlight Video
Interactive Session Tue AM Pod T.6
Authors: Sermanet, Pierre; Lynch, Corey; Chebotar, Yevgen; Hsu, Jasmine; Jang, Eric; Schaal, Stefan; Levine, Sergey
Title: Time-Contrastive Networks: Self-Supervised Learning from Video
Abstract:
We propose a self-supervised approach for learning representations and robotic behaviors entirely from unlabeled videos recorded from multiple viewpoints, and study how this representation can be used in two robotic imitation settings: imitating object interactions from videos of humans, and imitating human poses. Imitation of human behavior requires a viewpoint-invariant representation that captures the relationships between end-effectors (hands or robot grippers) and the environment, object attributes, and body post. We train our representations using a triplet loss, where multiple simultaneous viewpoints of the same observation are attracted in the embedding space, while being repelled from temporal neighbors which are often visually similar but functionally different. This signal causes our model to discover attributes that do not change across viewpoint, but do change across time, while ignoring nuisance variables such as occlusions, motion blur, lighting and background. We demonstrate that this representation can be used to allow a robot to directly mimic human poses without an explicit correspondence, and that it can be used as a reward function within a reinforcement learning algorithm to perform complex tasks such as pouring. Video results, open-source code and dataset are available at sermanet.github.io/imitate
ICRA 2018 Spotlight Video
Interactive Session Tue AM Pod T.6
Authors: Sermanet, Pierre; Lynch, Corey; Chebotar, Yevgen; Hsu, Jasmine; Jang, Eric; Schaal, Stefan; Levine, Sergey
Title: Time-Contrastive Networks: Self-Supervised Learning from Video
Abstract:
We propose a self-supervised approach for learning representations and robotic behaviors entirely from unlabeled videos recorded from multiple viewpoints, and study how this representation can be used in two robotic imitation settings: imitating object interactions from videos of humans, and imitating human poses. Imitation of human behavior requires a viewpoint-invariant representation that captures the relationships between end-effectors (hands or robot grippers) and the environment, object attributes, and body post. We train our representations using a triplet loss, where multiple simultaneous viewpoints of the same observation are attracted in the embedding space, while being repelled from temporal neighbors which are often visually similar but functionally different. This signal causes our model to discover attributes that do not change across viewpoint, but do change across time, while ignoring nuisance variables such as occlusions, motion blur, lighting and background. We demonstrate that this representation can be used to allow a robot to directly mimic human poses without an explicit correspondence, and that it can be used as a reward function within a reinforcement learning algorithm to perform complex tasks such as pouring. Video results, open-source code and dataset are available at sermanet.github.io/imitate










