EWRL 2018
Stable, Practical and On-line Bootstrapped Conservative Policy Iteration
Abstract
We consider on-line model-free reinforcement learning with discrete actions, and focus on sample-efficiency and exploration quality. In this setting, value-based methods of the Q-Learning family achieve state-of-the-art results, while actor-critic algorithms, that learn an explicit actor policy in addition to their value function, do not achieve empirical results matching their theoretical promises. The majority of actor-critic algorithms combine an on-policy critic with an actor learned with Policy Gradient. In this paper, we propose an alternative to these two components. We base our work on Conservative Policy Iteration (CPI), leading to a new non-parametric actor learning rule, and Dual Policy Iteration (DPI), that motivates our use of aggressive off-policy critics. Our empirical results demonstrate a sample-efficiency and robustness superior to state-of-the-art value-based and actor-critic approaches, in three challenging environments.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- European Workshop on Reinforcement Learning
- Archive span
- 2008-2025
- Indexed papers
- 649
- Paper id
- 121213717456577529