PRL 2021
Discount Factor Estimation in a Model-Based Inverse Reinforcement Learning Framework
Abstract
We consider the crucial task of estimating an expert’s discount factor in Inverse Reinforcement Learning (IRL) to facilitate a better synthesis towards the resulting optimal policy. Existing IRL algorithms have significantly overlooked the vital need to estimate the discount factor, experimental studies and theoretical intuitions show variability of the learnt reward function as the discount factor changes. In this work, we adapt the model-based maximum entropy IRL framework and optimize a utility-based softmax likelihood function via a featurebased gradient update to jointly learn the discount factor and reward. To test our approach, we utilize behavioral data from three Markov decision process (MDP) environments, namely, Grid-World, Mountain-Car Driving and Object-World. Experimental and numerical studies show that our approach is viable for the simultaneous estimation of the discount factor and reward function in IRL.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- Bridging the Gap Between AI Planning and Reinforcement Learning
- Archive span
- 2020-2025
- Indexed papers
- 151
- Paper id
- 569754591707216814