RLDM 2015
Off-policy learning with linear function approximation based on weighted importance sam- pling
Abstract
An important branch of reinforcement learning is off-policy learning where the agent behaves according to one policy but learns about a different policy. Many modern algorithms for model-free rein- forcement learning have incorporated off-policy learning together with parametric function approximation and sophisticated techniques such as eligibility traces. They all use a Monte Carlo technique known as importance sampling as a core component. The ordinary importance sampling estimator typically has high variance, and consequently, off-policy learning algorithms often exhibit poor performance. In Monte Carlo estimation, this problem is overcome using a variant of importance sampling called weighted importance sampling which often has much lower variance. However, weighted importance sampling has been ne- glected in off-policy learning due to the difficulty of combining it with parametric function approximation and hence not been utilized in modern reinforcement learning algorithms. In this work, we provide the key ideas on how off-policy learning algorithms for linear function approximation can be developed based on weighted importance sampling. We work with two different forms of methods for linear function approx- imation, methods of least squares resulting in an ideal but computationally expensive form and methods based on stochastic gradient descent providing a computationally congenial approximation. We empiri- cally demonstrate that the new algorithms can achieve substantial performance gain over the state-of-the-art off-policy algorithms and hence retain the benefits of weighted importance sampling.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- Multidisciplinary Conference on Reinforcement Learning and Decision Making
- Archive span
- 2013-2025
- Indexed papers
- 1004
- Paper id
- 230548737500958363