Arrow Research search
Back to IROS

IROS 2018

Synthesizing Neural Network Controllers with Probabilistic Model-Based Reinforcement Learning

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

We present an algorithm for rapidly learning neural network policies for robotics systems. The algorithm follows the model-based reinforcement learning paradigm and improves upon existing algorithms: PILeO and a sample-based version of PILeo with neural network dynamics (Deep-PILeO). To improve convergence, we propose a model-based algorithm that uses fixed random numbers and clips gradients during optimization. We propose training a neural network dynamics model using variational dropout with truncated Log-Normal noise. These improvements enable data-efficient synthesis of complex neural network policies. We test our approach on a variety of benchmark tasks, demonstrating data-efficiency that is competitive with that of PILeO, while being able to optimize complex neural network controllers. Finally, we assess the performance of the algorithm for learning motor controllers for a six legged autonomous underwater vehicle. This demonstrates the potential of the algorithm for scaling up the dimensionality and dataset sizes, in more complex tasks.

Authors

Keywords

  • Robots
  • Optimization
  • Vehicle dynamics
  • Task analysis
  • Heuristic algorithms
  • Neural networks
  • Stochastic processes
  • Neural Network
  • Neural Control
  • Neural Network Control
  • Model-based Reinforcement Learning
  • Dynamic Model
  • Random Number
  • Robotic System
  • Autonomous Underwater Vehicles
  • Artificial Neural Network
  • State Space
  • Recurrent Neural Network
  • Radial Basis Function
  • Optimal Policy
  • Neural Network Training
  • Stochastic Optimization
  • Policy Evaluation
  • Target System
  • Depth Camera
  • Markov Decision Process
  • Policy Network
  • Bayesian Neural Network
  • Unmanned Underwater Vehicles
  • Multiplicative Noise
  • Gaussian Process Model
  • Backpropagation Through Time
  • Policy Search
  • Stochastic Policy
  • Real Robot
  • Batch Mode
  • Learning Models

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
412504818718641694
v2026.09.13