EWRL 2015
Model-Free Preference-based Reinforcement Learning
Abstract
Specifying a numeric reward function for reinforcement learning typically requires a lot of hand-tuning from an human expert. In contrast, preference-based reinforcement learning (PBRL) utilizes only pairwise comparisons between trajectories as a feedback signal, which are often more intuitive to specify. Currently available approaches for PBRL with non-parametric policies require a known or estimated model. In this paper, we integrate preference-based estimation of the reward function into a model-free reinforcement learning (RL) algorithm, resulting in a model-free PBRL algorithm. In a model-free setup, we need to use a random exploration strategy which is implemented by a stochastic policy. We show that, by controlling the greediness of the policy update, a comparable performance to commonly used directed exploration strategies, which require a model, can be achieved. Our new algorithm is based on the Relative Entropy Policy Search algorithm and can learn non-parametric continuous action policies from a small number of preferences.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- European Workshop on Reinforcement Learning
- Archive span
- 2008-2025
- Indexed papers
- 649
- Paper id
- 1125651302656883285