Arrow Research search
Back to IROS

IROS 2022

Graph-Structured Policy Learning for Multi-Goal Manipulation Tasks

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Multi-goal policy learning for robotic manipu-lation is challenging. Prior successes have used state-based representations of the objects or provided demonstration data to facilitate learning. In this paper, by hand-coding a high-level discrete representation of the domain, we show that policies to reach dozens of goals can be learned with a single network using Q-learning from pixels. The agent focuses learning on simpler, local policies which are sequenced together by planning in the abstract space. We compare our method against standard multi-goal RL baselines, as well as other methods that leverage the discrete representation, on a challenging block construction domain. We find that our method can build more than a hundred different block structures, and demonstrate forward transfer to structures with novel objects. Lastly, we deploy the policy learned in simulation on a real robot.

Authors

Keywords

  • Q-learning
  • Planning
  • Task analysis
  • Standards
  • Intelligent robots
  • Policy Learning
  • Single Network
  • Block Structure
  • Robot Manipulator
  • Real Robot
  • Abstract Space
  • Block Construction
  • Standard Method
  • Time Step
  • State Space
  • Global Status
  • Depth Images
  • Building Structures
  • Structure Learning
  • Reward Function
  • Markov Decision Process
  • Transition Function
  • Reward Signal
  • Transition Rules
  • Robot State
  • Abstract States
  • Spatial Space
  • Start Of Episode
  • High-level Planner
  • Discrete State Space

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
1104854068968014232
v2026.09.13