Arrow Research search
Back to IROS

IROS 2017

Deep dynamic policy programming for robot control with raw images

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Deep reinforcement learning has drawn much attention in robot control since it enables agents to learn control policies from very high dimensional states such as raw images. On the other hand, its dependency upon the availability of a significant quantity of training samples and its fragility in learning makes it difficult to apply for real world robot tasks. To alleviate these issues we propose Deep Dynamic Policy Programming (DDPP), which combines the sample efficiency and smooth policy updates of dynamic policy programming with the contemporary deep reinforcement learning framework. The effectiveness of the proposed method is first demonstrated in a simulation of the robot arm control problem, with comparison to Deep Q-Networks. As validation on a real robot system, DDPP also successfully learned the flipping of a handkerchief with a NEXTAGE humanoid robot using a reduced number of learning samples, whereas Deep Q-Networks failed to learn the task.

Authors

Keywords

  • Dynamic programming
  • Programming
  • Manipulators
  • Robot control
  • Heuristic algorithms
  • Raw Images
  • Deep Learning
  • High-dimensional
  • Control Problem
  • Sampling Efficiency
  • Deep Reinforcement Learning
  • Humanoid Robot
  • Real Robot
  • Deep Q-network
  • Raw Data
  • Neural Network
  • Convolutional Neural Network
  • Value Function
  • Deep Neural Network
  • Gradient Descent
  • Video Games
  • Kullback-Leibler
  • Weight Vector
  • Fully-connected Network
  • Reward Function
  • Q-function
  • Final Layer
  • Markov Decision Process
  • Local Memory
  • Target Network
  • Optimal Policy
  • Curse Of Dimensionality
  • High-level Features

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
12526768599221194
v2026.09.13