Uploaded May 2018 | Updated September 2026, 3 weeks ago
ICRA 2018 Spotlight Video
Interactive Session Wed PM Pod E.8
Authors: Ojer de Andrés, Marco; Ghazaei Ardakani, M. Mahdi; Robertsson, Anders
Title: Reinforcement Learning for 4-Finger-Gripper Manipulation
Abstract:
In the framework of robotics, Reinforcement Learning (RL) deals with the learning of a task by the robot itself. This paper presents a hierarchical planning approach in which the robot learns the optimal behavior for different levels. For high-level discrete actions, Q-learning was chosen, whereas for the low level we utilize Policy Improvement with Path Integrals (PI$^2$) algorithm to learn the parameters of policies, represented by rhythmic Dynamic Movement Primitives (DMPs). The paper studies the case of a 4-finger-gripper manipulator, which performs the task of continuously spinning a ball around a desired axis. The results demonstrate the efficacy of the hierarchical planning and the improvement obtained in the performance of the task when PI$^2$ is used in conjunction with rhythmic DMPs in a real environment.
ICRA 2018 Spotlight Video
Interactive Session Wed PM Pod E.8
Authors: Ojer de Andrés, Marco; Ghazaei Ardakani, M. Mahdi; Robertsson, Anders
Title: Reinforcement Learning for 4-Finger-Gripper Manipulation
Abstract:
In the framework of robotics, Reinforcement Learning (RL) deals with the learning of a task by the robot itself. This paper presents a hierarchical planning approach in which the robot learns the optimal behavior for different levels. For high-level discrete actions, Q-learning was chosen, whereas for the low level we utilize Policy Improvement with Path Integrals (PI$^2$) algorithm to learn the parameters of policies, represented by rhythmic Dynamic Movement Primitives (DMPs). The paper studies the case of a 4-finger-gripper manipulator, which performs the task of continuously spinning a ball around a desired axis. The results demonstrate the efficacy of the hierarchical planning and the improvement obtained in the performance of the task when PI$^2$ is used in conjunction with rhythmic DMPs in a real environment.










