Arrow Research search
Back to RLDM

RLDM 2019

Thompson Sampling for Deep Reinforcement Learning

Conference Abstract Accepted abstract Artificial Intelligence · Decision Making · Machine Learning · Reinforcement Learning

Abstract

Exploration while learning representations is one of the main challenges of Deep Reinforcement Learning. Popular algorithms like DQN, use simple exploration strategies such as -greedy, which are prov- ably inefficient. The main problem with simple exploration strategies is that they do not use observed data to improve exploration. The Randomized Least Squares Value Iteration (RLSVI) algorithm [Osband et al. , 2016], uses Thompson Sampling for exploration and provides nearly optimal regret. In this work, we extend the DQN algorithm in the spirit of RLSVI: we combine DQN with Thompson Sampling, performed on top of the last layer activations. As this representation is being optimized during learning, a key component to our method is a likelihood matching mechanism, that adapts for the changing representations. We demon- strate that our method outperforms DQN in five Atari benchmarks and shows competitive results with the Rainbow algorithm.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
Multidisciplinary Conference on Reinforcement Learning and Decision Making
Archive span
2013-2025
Indexed papers
1004
Paper id
109677465092307323
v2026.09.13