RLDM 2019
DeepMellow: Removing the Need for a Target Network in Deep Q-Learning
Abstract
Deep Q-Network (DQN) is a learning algorithm that achieves human-level performance in high- dimensional, complex domains like Atari games. One of the important elements in DQN is its use of target network, which is necessary to stabilize learning. We argue that using a target network is incompatible with online reinforcement learning, and it is possible to achieve faster and more stable learning without a target network, when we use an alternative action selection operator, Mellowmax. We present new mathematical properties of Mellowmax, and propose a new algorithm, DeepMellow, which combines DQN and Mellow- max operator. We empirically show that DeepMellow, which does not use a target network, outperforms DQN with a target network.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- Multidisciplinary Conference on Reinforcement Learning and Decision Making
- Archive span
- 2013-2025
- Indexed papers
- 1004
- Paper id
- 1095340168204184702