Arrow Research search
Back to RLDM

RLDM 2019

DeepMellow: Removing the Need for a Target Network in Deep Q-Learning

Conference Abstract Accepted abstract Artificial Intelligence · Decision Making · Machine Learning · Reinforcement Learning

Abstract

Deep Q-Network (DQN) is a learning algorithm that achieves human-level performance in high- dimensional, complex domains like Atari games. One of the important elements in DQN is its use of target network, which is necessary to stabilize learning. We argue that using a target network is incompatible with online reinforcement learning, and it is possible to achieve faster and more stable learning without a target network, when we use an alternative action selection operator, Mellowmax. We present new mathematical properties of Mellowmax, and propose a new algorithm, DeepMellow, which combines DQN and Mellow- max operator. We empirically show that DeepMellow, which does not use a target network, outperforms DQN with a target network.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
Multidisciplinary Conference on Reinforcement Learning and Decision Making
Archive span
2013-2025
Indexed papers
1004
Paper id
1095340168204184702
v2026.09.13