Arrow Research search
Back to ICRA

ICRA 2019

Learning Action Representations for Self-supervised Visual Exploration

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Learning to efficiently navigate an environment using only an on-board camera is a difficult task for an agent when the final goal is far from the initial state and extrinsic rewards are sparse. To address this problem, we present a self-supervised prediction network to train the agent with intrinsic rewards that relate to achieving the desired final goal. The network learns to predict its future camera view (the future state) from a current state-action pair through an Action Representation Module that decodes input actions as higher dimensional representations. To increase the representational power of the network during exploration we fuse the responses from the Action Representation Module in the transition network, which predicts the future state. Moreover, to enhance the discrimination capability between predictions from different input actions we introduce joint regression and triplet ranking loss functions. We show that, despite the sparse extrinsic rewards, by learning action representations we achieve a faster training convergence than state-of-the-art methods with only a small increase in the number of the model parameters.

Authors

Keywords

  • Training
  • Navigation
  • Task analysis
  • Visualization
  • Robots
  • Predictive models
  • Cameras
  • Action Representation
  • Loss Function
  • Increase In The Number
  • High-dimensional
  • Future Conditions
  • Final Goal
  • Camera View
  • Extrinsic Rewards
  • State-action Pair
  • Intrinsic Rewards
  • High-dimensional Representation
  • Convolutional Neural Network
  • Convolutional Layers
  • Prediction Error
  • Number Of Steps
  • Long Short-term Memory
  • Feature Learning
  • Forward Model
  • Inverse Model
  • Reward Function
  • Inverse Reinforcement Learning
  • Number Of Time Steps
  • Imitation Learning
  • Self-supervised Learning
  • First-person View
  • Expert Demonstrations
  • Discrete Action
  • Expert Supervision
  • Left Turn
  • Action Coding

Context

Venue
IEEE International Conference on Robotics and Automation
Archive span
1984-2025
Indexed papers
30179
Paper id
263354110134533628
v2026.09.27