RLDM 2019
Contextual Markov Decision Processes using Generalized Linear Models
Abstract
We consider the recently proposed reinforcement learning (RL) framework of Contextual Markov Decision Processes(CMDP), where the agent interacts with infinitely many tabular environments in a se- quence. In this paper, we propose a no regret online RL algorithm in the setting where the MDP parameters are obtained from the context using generalized linear models (GLM). The proposed algorithm GL-ORL is completely online and memory efficient and also improves over the known regret bounds in the linear case. In addition to an Online Newton Step based method, we also extend existing tools to show conversion from any online no-regret algorithm to confidence sets in the multinomial GLM case. A lower bound is also provided for the problem.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- Multidisciplinary Conference on Reinforcement Learning and Decision Making
- Archive span
- 2013-2025
- Indexed papers
- 1004
- Paper id
- 446137975118109881