RLDM 2013
Path Integral Stochastic Optimal Control for Reinforcement Learning
Abstract
Path integral stochastic optimal control based learning methods are among the most efficient and scalable reinforcement learning algorithms. In this work, we present a variation of this idea in which the optimal control policy is approximated through linear regression. This connection allows the use of well- developed linear regression algorithms for learning of the optimal policy, e. g. learning the structural param- eters as well as linear parameters. In path integral reinforcement learning, Policy Improvement with Path Integral (PI2 ) algorithm is one of the most efficient and most similar algorithms to the algorithm we propose here. However, in contrast to the PI2 algorithm that relies on the Dynamic Movement Primitive (DMPs) to become a model free learning algorithm, our proposed method is formulated for an arbitrary parameterized policy represented by a linear combination of nonlinear basis functions. Additionally, as the duration and the goal of the task is part of the optimization in some tasks like shortest-time path optimization problem, our proposed method can directly optimize these quantities instead of assuming them to be given, fixed pa- rameters. Furthermore PI2 needs a batch of rollouts for each parameter update iteration whereas our method can update after just one rollout. The simulation result in this work shows that a simple implementation of our proposed method can at least perform as well as PI2 despite only using ‘out-of-the-box’ regression and a ’naive’ sampling strategy. In this light, the here presented should only be considered as a preliminary step in the development of our new approach which addresses some issues in the derivation of previous algorithms. Basing the development on this improvements, we believe that this work will ultimately lead to more efficient learning algorithms.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- Multidisciplinary Conference on Reinforcement Learning and Decision Making
- Archive span
- 2013-2025
- Indexed papers
- 1004
- Paper id
- 716296154735096442