Arrow Research search
Back to ICRA

ICRA 2023

Off-policy Imitation Learning from Visual Inputs

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Recently, various successful applications utilizing expert states in imitation learning (IL) have been witnessed. However, IL from visual inputs (ILfVI), which has a greater promise to be widely applied by using online visual resources, suffers from low data-efficiency and poor performance resulted from on-policy learning and high-dimensional visual inputs. We propose OPIfVI (Off-Policy Imitation from Visual Inputs), which is composed of an off-policy learning manner, data augmentation, and encoder techniques, to tackle the mentioned challenges, respectively. More specifically, to improve data-efficiency, OPIfVI conducts IL in an off-policy manner, with which sampled data used multiple times. In addition, we enhance the stability of OPIfVI with spectral normalization to mitigate the side effect of off-policy training. The core factor, contributing to the poor performance of ILfVI, that we think is agents could not extract meaningful features from visual inputs. Hence, OPIfVI employs data augmentation from computer vision to help train encoders to better extract features from visual inputs. Besides, a specific structure of gradient backpropagation for the encoder is designed to stabilize the encoder training. At last, we demonstrate that OPIfVI can achieve expert-level performance and outperform existing baselines via extensive experiments using DeepMind Control Suite.

Authors

Keywords

  • Training
  • Backpropagation
  • Visualization
  • Computer vision
  • Automation
  • Computer architecture
  • Feature extraction
  • Visual Input
  • Imitation Learning
  • Data Augmentation
  • High-dimensional Input
  • Gradient Backpropagation
  • Spectral Normalization
  • Convolutional Neural Network
  • Learning Curve
  • Multilayer Perceptron
  • Visual Observation
  • Visual Presentation
  • Reward Function
  • Markov Decision Process
  • Partial Observation
  • Discriminator Loss
  • State-action Pair
  • Replay Buffer
  • Concurrent Work
  • Inverse Reinforcement Learning
  • Training Instability

Context

Venue
IEEE International Conference on Robotics and Automation
Archive span
1984-2025
Indexed papers
30179
Paper id
155330556795930131
v2026.09.13