Arrow Research search
Back to EWRL

EWRL 2018

Stable, Practical and On-line Bootstrapped Conservative Policy Iteration

Workshop Paper Accepted Paper Artificial Intelligence · Machine Learning · Reinforcement Learning

Abstract

We consider on-line model-free reinforcement learning with discrete actions, and focus on sample-efficiency and exploration quality. In this setting, value-based methods of the Q-Learning family achieve state-of-the-art results, while actor-critic algorithms, that learn an explicit actor policy in addition to their value function, do not achieve empirical results matching their theoretical promises. The majority of actor-critic algorithms combine an on-policy critic with an actor learned with Policy Gradient. In this paper, we propose an alternative to these two components. We base our work on Conservative Policy Iteration (CPI), leading to a new non-parametric actor learning rule, and Dual Policy Iteration (DPI), that motivates our use of aggressive off-policy critics. Our empirical results demonstrate a sample-efficiency and robustness superior to state-of-the-art value-based and actor-critic approaches, in three challenging environments.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
European Workshop on Reinforcement Learning
Archive span
2008-2025
Indexed papers
649
Paper id
121213717456577529
v2026.09.13