Arrow Research search
Back to IROS

IROS 2020

Sample-Efficient Learning for Industrial Assembly using Qgraph-bounded DDPG

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Recent progress in deep reinforcement learning has enabled agents to autonomously learn complex control strategies from scratch. Model-free approaches like Deep Deterministic Policy Gradients (DDPG) seem promising for applications with intricate dynamics, such as contact-rich manipulation tasks. However, these methods typically require large amounts of training data or meticulous hyperparameter tuning, limiting their usefulness for real-world robotics applications. In this paper, we evaluate and benchmark our recently proposed approach for improving model-free reinforcement learning with DDPG through Qgraph-based bounds in temporal difference learning. We directly apply the algorithm to a challenging real-world industrial insertion task and assess its performance (see https://youtu.be/Z_GcNbCWE-E).Empirical results show that the insertion task can be learned despite significant frictional forces and uncertainty, even in sparse-reward settings. We present an in-depth comparison based on a large number of experiments and demonstrate the advantages and performance of Qgraph-bounded DDPG: the learning process can be significantly sped up, robustified against bad choices of hyperparameters and runs with less memory requirements. Lastly, the presented results extend the current theoretical understanding of the link between data graph structure and soft divergence in DDPG.

Authors

Keywords

  • Shafts
  • Uncertainty
  • Training data
  • Reinforcement learning
  • Task analysis
  • Tuning
  • Convergence
  • Graph Data
  • Deep Reinforcement Learning
  • Model-free Approach
  • Model-free Reinforcement Learning
  • Industrial Tasks
  • Temporal Difference Learning
  • High Force
  • Per Cycle
  • Actor Network
  • Termination Condition
  • Target State
  • Limited Memory
  • Reward Function
  • Markov Decision Process
  • Critic Network
  • Toy Example
  • Static Friction
  • Machine Learning Community
  • Model-free Methods
  • Ball Bearings
  • Limited Memory Capacity
  • Loose Ends
  • Replay Memory
  • Deep Q-learning
  • State-action Pair
  • Random Baseline
  • Scope Of This Paper
  • Autonomous Agents

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
1058724671881613716
v2026.09.13