Arrow Research search
Back to EWRL

EWRL 2015

Model-Free Preference-based Reinforcement Learning

Workshop Paper Accepted Paper Artificial Intelligence · Machine Learning · Reinforcement Learning

Abstract

Specifying a numeric reward function for reinforcement learning typically requires a lot of hand-tuning from an human expert. In contrast, preference-based reinforcement learning (PBRL) utilizes only pairwise comparisons between trajectories as a feedback signal, which are often more intuitive to specify. Currently available approaches for PBRL with non-parametric policies require a known or estimated model. In this paper, we integrate preference-based estimation of the reward function into a model-free reinforcement learning (RL) algorithm, resulting in a model-free PBRL algorithm. In a model-free setup, we need to use a random exploration strategy which is implemented by a stochastic policy. We show that, by controlling the greediness of the policy update, a comparable performance to commonly used directed exploration strategies, which require a model, can be achieved. Our new algorithm is based on the Relative Entropy Policy Search algorithm and can learn non-parametric continuous action policies from a small number of preferences.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
European Workshop on Reinforcement Learning
Archive span
2008-2025
Indexed papers
649
Paper id
1125651302656883285
v2026.09.13