RLDM 2019
Thompson Sampling for Deep Reinforcement Learning
Abstract
Exploration while learning representations is one of the main challenges of Deep Reinforcement Learning. Popular algorithms like DQN, use simple exploration strategies such as -greedy, which are prov- ably inefficient. The main problem with simple exploration strategies is that they do not use observed data to improve exploration. The Randomized Least Squares Value Iteration (RLSVI) algorithm [Osband et al. , 2016], uses Thompson Sampling for exploration and provides nearly optimal regret. In this work, we extend the DQN algorithm in the spirit of RLSVI: we combine DQN with Thompson Sampling, performed on top of the last layer activations. As this representation is being optimized during learning, a key component to our method is a likelihood matching mechanism, that adapts for the changing representations. We demon- strate that our method outperforms DQN in five Atari benchmarks and shows competitive results with the Rainbow algorithm.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- Multidisciplinary Conference on Reinforcement Learning and Decision Making
- Archive span
- 2013-2025
- Indexed papers
- 1004
- Paper id
- 109677465092307323